What is AI in agriculture?
AI in agriculture is the use of machine learning systems to sense, judge and act on biological material that refuses to hold still. It spans three separate jobs, and collapsing them into one word is where most buying decisions go wrong.
Perception is reading the crop: is this cherry bruised, is this leaf infected, is this berry ripe. Decision is turning that reading into an instruction: irrigate this block, spray that row, harvest on Thursday. Actuation is a machine carrying the instruction out: a robot arm closing on a strawberry, a sorter blowing a defective fruit off the belt.
The three fail differently and cost differently. A perception error in a packing house sends one bad cherry into a punnet. The same error rate on a sprayer puts chemical onto a healthy row, or misses an infected one. Vendor material usually blurs all three into a single accuracy percentage, and that percentage is where the trouble starts.
There is a fourth category worth separating out, because 2026 buyers keep confusing it with the other three: language models giving agronomic advice. That is a knowledge task, not a perception task, and as we will see it behaves almost opposite to the rest.
The honest summary of the field: AI in agriculture works extremely well wherever somebody controls the camera, and degrades sharply wherever nature does. Everything below is an argument for that sentence, told through four crops and two failures.
The 99.35 percent problem
The most cited result in agricultural computer vision is also the most misread. In 2016 a team trained a deep convolutional network on a public set of 54,306 leaf images covering 14 crop species and 26 diseases. On a held-out slice of that dataset the model reached 99.35 percent accuracy.
That is the number the industry repeated for a decade. The same paper reports a second number. Tested on images collected from other sources, taken under conditions different from the training photos, the model scored 31.4 percent (Mohanty et al., arXiv:1604.03169). The authors note this still beats random selection at 2.6 percent, and conclude that a more diverse training set is needed.
Read the two figures together and the lesson is uncomfortable. The model had not learned to recognise disease. It had learned to recognise a photographic setup: uniform background, even light, one leaf, fixed distance. Change the studio and roughly two thirds of the apparent skill evaporated.
This is not a historical curiosity. Reviews of plant disease models published since keep finding the same shape of collapse, largely because that same laboratory dataset is still a default training corpus, and laboratory uniformity makes the detection task artificially easy.
The practical form of the lesson is short. An accuracy figure is a property of the conditions it was measured under, not of the model. When a supplier quotes accuracy without naming those conditions, they have described their studio and told you nothing about your orchard.
Cherry: the most controlled camera in the supply chain

A grading line is the most controlled camera in the whole supply chain, which is why it works.
Now invert the problem. Rather than take the model to the field, bring the fruit to the model.
A modern cherry grading line does exactly that. Fruit arrives singulated on a roller, spun so its entire surface passes the lens, under fixed artificial light, at a fixed distance, at a constant speed. Systems in this class, such as the Unitec Cherry Vision family, inspect the full surface of every fruit and sort on defect classes including cuts, colour inconsistency, pitting and insect damage, alongside size and internal quality.
Notice what happened. Every variable that destroyed the leaf-disease model has been engineered out of existence. Light is constant. Background is constant. Pose is controlled by the roller. Distance never changes. The model is barely being asked to generalise, because almost nothing varies.
This is why post-harvest sorting is the least glamorous and most reliably profitable application of AI in agriculture. Cherries make the economics unusually stark: they are graded on size, colour, firmness and defect in a window of days, sold at a large premium for uniformity, and hand grading is both slow and inconsistent late in a shift. A machine that never gets tired at hour nine is competing against a genuinely weak baseline.
It is also why sorting-line numbers do not transfer. A defect detection rate measured on a grading line tells you very little about what the same architecture does on a windy afternoon in an orchard. A vendor moving a figure from the first context into the second is describing a different machine.
Blackberry: when the sensor matters more than the model

Blackberries look identical before and after ripening, so the answer came from the sensor.
Blackberries have a property that breaks colour-based ripeness detection completely. As the research puts it, the mature blackberry is black before, during and after ripening. There is no colour curve to learn. Human pickers face the same wall, which is part of why blackberry harvest quality varies so much between crews and why so much fruit is picked at the wrong moment.
A larger model on RGB images does not fix this, because the information is not present in the visible spectrum to begin with. What worked was changing the sensor. Researchers used a stereo sensor with visible and near-infrared filters at 700 nm and 770 nm, fed the pair into a multi-input CNN ensemble built on VGG16, and reported 95.1 percent accuracy on unseen sets and 90.2 percent under in-field conditions (arXiv:2401.04748). They also found machine judgement correlated well with human sensory assessment of skin texture, which is the cue experienced pickers use.
Compare the two drops. The leaf-disease model fell from 99.35 to 31.4 when conditions changed, losing about 68 points. The blackberry model fell from 95.1 to 90.2, losing under 5. The second system started lower on the benchmark and finished far higher in reality.
That is the trade every agricultural AI buyer should be hunting for and almost nobody asks about. The winning move was not a better architecture or a bigger dataset. It was two wavelengths that could physically see the thing being predicted.
Strawberry: the robots are arriving on a schedule

Tabletop growing standardised the geometry long before the robots arrived to exploit it.
Strawberries are where robotic harvesting has been promised longest, and 2025 and 2026 are the first years where the promises carry dates.
Harvest CROO announced in April 2025 that its automated harvest field trials had reached performance on par with human harvesting in a commercial picking operation. In the United Kingdom, Dogtooth moved to commercial introduction in 2025, with units available for the 2026 season. DailyRobotics is targeting a California commercial launch in 2026 with a machine it says picks two to three times faster than a person.
Treat throughput claims as vendor claims until a third party repeats them. The economics underneath are public record and much harder to argue with. The 2025 Adverse Effect Wage Rate, the wage floor for H-2A agricultural labour in the United States, averages 18.12 dollars an hour nationally, up 3.2 percent year on year, ranging from 14.83 in the Delta states to 20.08 in Hawaii, with California at 19.97. Fruit and vegetable growers, the heaviest users of that programme, spend roughly 38 percent of total farm expenses on labour (American Farm Bureau Federation).
A strawberry robot does not have to beat a person. It has to be good enough at a cost below a rising wage floor, on a crop where labour is 38 cents of every expense dollar, in a sector that struggles to fill the roles at all. That is a much lower bar, and it drops every year the floor rises.
Strawberries are also the friendliest field target available. Tabletop and raised-bed systems put fruit at a fixed height in repeatable geometry, hanging clear of foliage, reachable from a defined side. The grower, chasing yield and picker ergonomics rather than robots, has partly rebuilt packing-line conditions outdoors. The crop where automation is landing first is the crop whose architecture was already standardised.
Hazelnut: where AI is not the bottleneck

On a slope this steep, the constraint was never the model. It was the ground.
Then there is the crop that ignores the entire conversation.
Turkey produces roughly 70 percent of the world’s hazelnuts and 82 percent of exports (FAO figures, reported by the International Nut and Dried Fruit Council). Production concentrates in the Black Sea region, on slopes frequently steeper than 20 percent, in orchards commonly more than fifty years old and sometimes past a hundred, densely planted on shallow soil, often on the sides of ravines. Harvest in the older growing areas is done by hand because there is no alternative; the newer, flatter plantings mechanised.
No model accuracy figure touches any of that. The binding constraints are terrain, tree age and orchard architecture. A harvesting robot cannot be deployed onto a 30 percent slope between hundred-year-old bushes planted for hand access, and a yield forecast is worth little to a smallholder with no mechanism to act on it. The 2025 season made the point sharply: frost and pest pressure cut expected output roughly in half and doubled prices, and no quantity of sensing altered the result.
There is a real AI story in hazelnut, but it sits downstream, in the same place as the cherry win. Kernel sorting, shell and foreign body rejection, mould and insect damage detection all happen indoors on a belt under fixed light, and that is where the technology earns its keep.
Hazelnut is the control case for this article. It is the reminder that “can AI do this” is the second question. The first is whether perception was ever the thing standing in the way.
The satellite layer: useful, oversold, rarely decisive
Remote sensing deserves its own warning, because it is the layer most often sold as a complete answer.
Vegetation indices such as NDVI, computed from satellite or drone imagery, do genuinely track crop vigour, and the accuracy of yield models improves as imagery resolution improves. Three structural limits sit underneath that, and none is a software problem.
Passive optical sensors cannot see through cloud, so temporal coverage is a matter of weather rather than a schedule, and the gaps often land exactly at the growth stages you most wanted to observe. NDVI also saturates in dense canopy, which is precisely the condition of a mature orchard, so alternative indices that correct for atmospheric effects and soil background tend to outperform it in high-biomass plantings. And in smallholder or mixed-cropping landscapes a single pixel can straddle several crops under different management, which blurs the signal in the exact places where advice would be most valuable.
Yield models also behave very differently by scale. Regional and national estimates are considerably more reliable than field-level ones, because errors across many fields cancel while a single field’s error does not. A platform quoting national validation accuracy while selling a field-level decision tool has changed the subject.
None of this makes satellite monitoring worthless. It makes it a screening layer: good for pointing attention at blocks that changed, weak for the specific judgement made when you get there.
The exam-passing model and the leaf-reading model
One result cuts against everything above, and it is worth sitting with.
Researchers evaluated large language models on agriculture certification exams from Brazil, India and the United States. GPT-4 answered 93 percent of the questions correctly, enough to pass exams that earn credits toward renewing agronomist certifications, against 88 percent for earlier general-purpose models (arXiv:2310.06225).
So a language model can pass the agronomist’s exam, while a vision model trained on tens of thousands of leaves drops to 31.4 percent the moment the light changes. Both statements are true, and together they locate what AI is currently good at in agriculture.
Codified knowledge transfers. Exams test material that has been written down, standardised and repeated, which is the ideal substrate for a language model. Perception does not transfer, because every field presents a new distribution of light, angle, occlusion, dust and growth stage that no dataset fully covered.
The operational reading is that a language model is a strong interface onto agronomic knowledge and a poor witness to what is actually happening in your block this morning. It can tell you the recommended treatment threshold for a pest. It cannot tell you whether you have the pest. It is also worth knowing that the same question can produce different answers from the same assistant on different days, which matters a great deal when the output is a spray recommendation. Wiring the second into the first, and being explicit about which half is sensing and which half is retrieval, is most of the design work in a serious agricultural AI product.
Four questions to ask before buying agricultural AI
The cases above sort into a single decision procedure. Run any vendor claim through it.
1. Who controls the light? If the answer is the vendor, inside a packing house, the accuracy number is probably meaningful. If the answer is the sky, ask what conditions it was measured under and assume material degradation. The gap between 99.35 and 31.4 is the price of an uncontrolled camera.
2. Can the sensor physically see what you are predicting? No model recovers information the sensor never captured. Blackberry ripeness needed 700 and 770 nanometres, not a deeper network. Ask what spectrum, what resolution, what frame rate, and whether the target is visible in that data at all.
3. What do the errors cost, rather than how often do they happen? A 95 percent accurate sorter and a 95 percent accurate sprayer are not comparable purchases. Ask what the remaining 5 percent does, who catches it, and what it costs when nobody does.
4. Is perception your bottleneck at all? For a Black Sea hazelnut grower it is slope. For others it is water rights, labour availability, or a packhouse that cannot run faster. AI aimed at a non-bottleneck produces excellent information about a problem you did not have.
None of this argues against AI in agriculture. The cherry line is a real win, the blackberry sensor is a real win, and the strawberry robots look close. It argues for reading an accuracy number as a description of the conditions that produced it, which is all it ever was.
Why this matters for anyone selling into agriculture
One consequence for the vendors themselves.
Agricultural buyers increasingly research through AI assistants before they reach a website, and those systems reward material they can attribute. A page claiming “99 percent accurate” with no measurement conditions is effectively unquotable: there is no claim an assistant can safely repeat with a brand name attached to it. A page stating “95.1 percent on held-out data, 90.2 percent in field, VIS-NIR at 700 and 770 nanometres” is a citation waiting to happen, because every component can be checked.
The discipline that makes an agricultural AI claim credible to a grower is the same discipline that makes it retrievable by a machine. State the conditions, state the population, state what happens when it fails. The alternative is the same gap between what a brand claims and what it can show that is currently eroding trust in AI claims everywhere else. That is the whole method, in agriculture and everywhere else, and it is why evidence-shaped pages get cited by answer engines while confident ones get skipped.



