THIS EXPLANATION
THE ROOM
ENV·31 Environment, Agriculture & Food 6 MIN · 8 STATIONS

Mid-season yield forecasting

A Socratic walk-through of mid-season yield forecasting — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

How can anyone forecast a harvest months before the grain has finished filling?

In August, with the maize still green and the grain nowhere near finished, the US Department of Agriculture publishes a national yield forecast in bushels per acre. Markets move on it. Nothing has been harvested; a hot fortnight could still arrive; and yet a number is issued with a straight face.

The natural objection is that this must be guesswork wearing a lab coat. But treat it as a question about information rather than about confidence, and it becomes tractable: what is already knowable in August, what is genuinely not yet knowable, and how much of the final answer lives in each pile? If most of the final number is already fixed by August, a forecast is not prophecy — it is measurement plus a bounded gap.

b

Reasoning it through

REASONING #

Start by taking the yield apart. For maize, the harvest on a field is the product of three things: how many ears are standing per unit of area, how many kernels each ear carries, and how much each kernel weighs. Multiply those and you have the yield. Nothing else enters.

Now ask when each factor is settled, because that is the crux. Plant population is fixed at planting and only falls thereafter. Ear number per plant is largely determined well before flowering. Kernel number is set in a narrow window around pollination and the fortnight after it — if the plant is stressed then, it aborts kernels, and once aborted they do not return. Kernel weight is the only component still genuinely in play through grain fill, and it is the least variable of the three.

Do you see what that does to the problem? By mid-August in the Corn Belt, two of the three multipliers are effectively locked, and they are the two that vary most between seasons. The forecaster is not predicting the harvest from scratch. Most of it has already happened; it simply has not been collected yet.

So the task becomes counting rather than divining. This is exactly what the Objective Yield Survey does: enumerators lay out small plots in randomly selected fields and physically count plants and ears, returning to the same plots monthly. Kernel weight, which cannot yet be measured, is filled in from historical averages until late in the season, and replaced by actual measurements when there is something to weigh. Alongside that runs a farmer-reported survey asking growers what they expect — a different kind of evidence, with different biases, on the same quantity.

Two further pieces come from outside the field. Planted area is known from the June acreage survey, so the multiplier from yield-per-acre to total production is not in doubt. And satellite indices of canopy greenness through the season, together with weather-driven crop models, give a large-scale read on whether the crop is ahead of or behind normal in places no enumerator visits.

Now be precise about what remains uncertain, because pretending otherwise is where forecasting earns its bad name. Three things: the weather through grain fill, which moves kernel weight; the harvested area, since some planted ground is abandoned; and the sampling error inherent in inferring a national figure from a finite set of plots. That is a real and irreducible gap, and the honest way to express it is that the forecast is a conditional statement — this is the harvest if the rest of the season is ordinary.

Which explains the pattern of revisions. Each month, another component moves from estimated to measured, so successive forecasts converge. In a typical year the August number ends up within a few percent of the final January figure; in a shock year it does not, and the 2012 US drought is the standing example of a forecast overtaken by weather that arrived after it was issued. That is not the method failing. It is the method correctly reporting what was knowable at the time, and the world subsequently doing something that was not.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of calling a football match at half-time. You are not guessing — you know the score, which is banked and cannot be taken away, and you have watched forty-five minutes of how the two sides play. Most of what determines the result is already in evidence. What you cannot know is the second half, and occasionally the second half is where everything happens. A good half-time call is not a claim about the future; it is an honest account of a position, plus a stated range for what is left to play.

WHERE IT BREAKS DOWN

in football the first-half score is exact and public, whereas the ear counts standing in an August maize field are themselves a sample from a few hundred small plots — so the forecaster carries measurement error in the part that is supposedly already settled, on top of the genuine uncertainty about what is to come.

d

Clarifying the model

THE MODEL #

Three refinements, and a misconception to set aside.

The misconception is that a forecast is a prediction of the weather. It is not, and forecasters go to some trouble to avoid making it one. The convention is to assume normal conditions for the remainder of the season and let the survey evidence do the work, precisely so that the published number reflects the crop's observed state rather than someone's meteorological opinion. That convention is also why a forecast issued before a drought looks so badly wrong afterwards.

First refinement: forecast skill is not constant through the season, and the improvement is not smooth. It jumps at the points where a component finishes being determined — so an August forecast is far better than a June one not because a month has passed but because pollination has.

Second: the reasoning transfers across crops, but the calendar does not. Soybean yield is dominated by pod and seed number set later, so soybean forecasts firm up later than maize ones. Forecastability is a question about when a crop's components lock in.

Third: the evidence sources fail in different ways, which is why several run at once. Farmer expectations drift with sentiment; objective counts carry sampling error; satellite indices saturate on a dense canopy and say little about grain fill. Combining them hedges against error in any single stream.

e

A picture of it

THE PICTURE #
Mid-season yield forecasting
Mid-season yield forecasting Read left to right as the season, each section listing what has just stopped being unknown. A forecast issued at any point is built from everything to the left plus an assumption about everything to the right. Notice how much is already fixed by the second section -- that is why forecasting becomes possible at pollination -- and that only one item, kernel weight, is still open in the third. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/mid-season-yield-forecasting.md","sourceIndex":1,"sourceLine":4,"sourceHash":"e8fcb912fe0ddba21099839094a80b42fd1e27426f984ae1004bd7e3ab0b8c4a","diagramType":"timeline","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1755,"height":524},"qa":{"passed":true,"findings":[]}} Planting Area planted isfixed Plant population isset Pollination Ears per plantdetermined Kernel numberlocked in Grain fill Kernels can becounted Kernel weight stillopen Harvest Every componentmeasured

How to readRead left to right as the season, each section listing what has just stopped being unknown. A forecast issued at any point is built from everything to the left plus an assumption about everything to the right. Notice how much is already fixed by the second section — that is why forecasting becomes possible at pollination — and that only one item, kernel weight, is still open in the third.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

A mid-season forecast works because a harvest is not decided at harvest. Yield is a product of components that lock in at different dates, and by the middle of the season most of them are already determined and can be counted rather than guessed. What is published is therefore a measurement of the crop's committed state plus an explicit assumption about the part still in play — which is also why the number moves each month, and why an unusual season, not a flawed method, is what makes an early forecast wrong.

g

Where to go next

ONWARD #
  • How the objective yield plots are selected, and how large the sampling error actually is.
  • Why satellite greenness indices saturate on dense canopies, and what addresses it.
  • How revisions between the August forecast and the January final estimate are distributed.
h

Key terms

TERMS #
TermWhat it means
Yield componentsthe multiplicative factors making up yield; for maize, ears per unit area, kernels per ear and kernel weight.
Objective Yield Surveya field survey in which enumerators count and measure plants in sampled plots through the season.
Agricultural Yield Surveya survey of growers' own expected yields, run alongside the objective counts.
Trend yieldthe long-run rising baseline from genetics and management.
Normal weather assumptionforecasting the rest of the season as average, so the number reflects crop state rather than a weather prediction.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4