FAOSTAT crops and livestock products
Use: Country-level production, yield, and price trends for initial baselines and feature exploration.
Watch: Too aggregated for field-level predictions; use it for context, not as ground truth for one farm.
Build the data foundation first
Every input needs a reason, a source, a consistent unit, and a link to a model or cost calculation. The table below is the prototype data contract.
Manual fields should stay short; fetch or prefill where a reliable source exists.
| Category | Input | Why it is needed | Capture method | Expected type / unit |
|---|---|---|---|---|
| Soil | Soil pH | Controls nutrient availability and crop suitability. | Manual soil test or lab dataset | Decimal pH (0–14) |
| Soil | N, P, K and organic matter | Explains yield response and fertilizer need. | Soil lab / SoilGrids proxy | mg/kg, %, or categorical band |
| Soil | Texture and drainage | Affects water holding, disease pressure, and irrigation need. | Soil survey / manual | Category: clay, loam, sand; drainage class |
| Crop | Crop, variety, planting date | Sets the crop-specific yield and disease baseline. | Manual | Category + date (YYYY-MM-DD) |
| Crop | Field area and historic yield | Scales per-hectare estimates and grounds the model in local performance. | Manual farm records | Hectares; tonnes/hectare |
| Crop | Growth stage and planting density | Determines which inputs and disease actions are still practical. | Manual | Stage category; plants/hectare |
| Weather | Rainfall, temperature, humidity | Major drivers of crop growth, water stress, and disease. | Automatic weather API / historical dataset | mm; °C; % RH |
| Weather | Forecast heat, frost, wind, drought | Changes near-term yield loss and spray or irrigation decisions. | Automatic forecast API | Probability or continuous forecast values |
| Disease & pest | Observed symptoms or leaf image | Supports disease or pest classification and targeted action. | Manual scouting / image model | Category, severity 0–100%, JPEG/PNG |
| Disease & pest | Trap counts and prior outbreaks | Improves pest risk beyond a single observation. | Manual farm log / regional advisory data | Count; binary or categorical history |
| Inputs & costs | Fertilizer rate and unit cost | A direct variable cost; rate may also affect expected yield. | Manual invoices / supplier feed | kg/hectare; currency/hectare |
| Inputs & costs | Pesticide and weedicide cost | Direct variable cost; can prevent losses only when an actual risk is present. | Manual invoices / supplier feed | Currency/hectare; product/category |
| Inputs & costs | Labor, fuel, irrigation, equipment | Often determines whether a seemingly high-yield plan is profitable. | Manual farm records | Currency/hectare or hours/hectare |
| Market | Local spot price and buyer/market | Sets the selling-price assumption for revenue. | Automatic market feed + manual buyer quote | Currency/tonne; market ID |
| Market | Expected harvest date and price volatility | Price at harvest can differ from today’s price. | Historical price dataset / manual estimate | Date; % standard deviation or risk band |
| Other | Transport, storage, commissions | Revenue is not fully realized until produce reaches the buyer. | Manual farm records | Currency/tonne or currency/hectare |
| Other | Credit cost, insurance, and farmer risk preference | Important for a realistic decision, but not always a yield driver. | Manual | Currency; low/medium/high |
Use them to create baselines and enrich farm records—not to imply field-level certainty.
Use: Country-level production, yield, and price trends for initial baselines and feature exploration.
Watch: Too aggregated for field-level predictions; use it for context, not as ground truth for one farm.
Use: US crop production, yield, and market-price data where the prototype targets US regions.
Watch: Coverage and commodity definitions vary; align units and geography before joining.
Use: Temperature, rainfall, radiation, humidity, and wind features by date and location.
Watch: Weather-grid resolution may not match a microclimate or a farm weather station.
Use: Global gridded soil properties such as pH, texture fractions, organic carbon, and bulk density.
Watch: Use as a prior when a soil test is unavailable; it does not replace a current lab sample.
Use: Starter image data for a leaf-disease classifier or a disease-risk proof of concept.
Watch: Images are often cleaner than farmer photos, so field accuracy can be much lower.
Use: The most valuable training rows: field, season, crop, input rates, costs, yield, price, and observed losses.
Watch: A hackathon should use a small, consistent sample and clearly label models as prototype estimates.
One row per field-season is enough to start: field ID, crop, planting and harvest dates, soil values, weather summaries, input rates and costs, disease observations, yield, realized price, transport/storage cost, and realized profit. Use a time-based split so future seasons are held out for testing.
Normalize all cost and yield values to the same area and currency unit, record the date each price was known, remove data leakage from post-harvest variables, and keep missing-value flags. A clean baseline table is more valuable than a complicated model with mismatched data.