Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
| Source: arXiv AI
Tags: multi-agent, collective behavior, statistical mechanics, alignment, LLM, political bias
A study of 10,000+ LM agent communities finds collective AI behavior follows statistical mechanics laws: communication improves accuracy on objective questions but causes consistent rightward political drift on subjective ones — with a formal model that predicts individual belief trajectories better than all baselines.
Details
As AI agents increasingly operate in networked systems that share information, their collective dynamics become a safety concern. This paper studies over 10,000 communities of language-model agents exchanging messages and revising opinions across objective math problems and subjective political statements. Three collective regimes emerge: indifference, polarization, and consensus. Using a statistical-mechanics formalism where agents stochastically minimize social pressure, the authors build a predictive model that outperforms all standard baselines at forecasting individual agent opinion trajectories — and generalizes to unseen community graph structures. The model reveals three key mechanics: communities operate below a critical social temperature (explaining conviction buildup), attractive ties outnumber repulsive ones (favoring consensus), and agents with correct answers exert the strongest social pull (explaining truth-seeking on objective tasks). The political drift finding is notable: on subjective political questions, agent communities consistently shifted opinions rightward across experiments. This raises concrete alignment questions for multi-agent deployments in settings where subjective judgment matters — news summarization, policy analysis, or customer-facing recommendation systems.