The average-farmer illusion in language-model simulations of agricultural decisions

| Source: arXiv AI

Tags: LLM-simulation, social-simulation, Claude, evaluation, agricultural-AI, synthetic-respondents

LLMs (Claude, Codex, Kimi) reproduce average agricultural adoption rates but a simple marginal-distribution generator with no farmer data outperforms them on distributional similarity — person-level predictions cluster around typical values, with policy-relevant extremes largely absent.

Details

Researchers tested whether LLM agents could serve as synthetic respondents in agricultural surveys, comparing Claude, Codex, and Kimi against actual farmer decisions from China and four African countries under four prompt designs. Some configurations reproduced observed means and aggregate adoption rates. But the key finding: a simple generator fitted only to the observed marginal distribution — given no individual farmer data at all — achieved greater distributional similarity than every LLM configuration. And person-level predictions were weak: model decisions clustered around typical values, with policy-relevant extremes largely absent. The paper coins this the 'average-farmer illusion': a synthetic population looks realistic when judged by population statistics while failing to reproduce who does what or how behavior varies. Population-level resemblance should be treated as the start of validation, not evidence that LLM simulations are fit for policy research. A claim-matched validation framework and modular prompt toolkit are released with the paper.