Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
| Source: Apple ML Research
Tags: Apple, LLM behavior, human-like AI, GPT-4o, Claude Sonnet 4.6, system prompts, Gemini 2.5 Flash, AI evaluation
Apple researchers analyzed 21,000 multi-turn conversations across GPT-4o, GPT-4.1-mini, Claude Sonnet 4.6, and Gemini 2.5 Flash, finding human evaluators consider AI self-referential and relationship-building behaviors less appropriate than from humans — but boundary-setting more appropriate from AI than from people.
Details
Researchers from Apple published a multi-dimensional analysis of human-like behaviors in LLMs, covering 21,000 multi-turn conversations drawn from four widely used models: GPT-4o, GPT-4.1-mini, Claude Sonnet 4.6, and Gemini 2.5 Flash. The study categorizes behaviors into three groups: self-referential (expressing thoughts and emotions), relationship-building (engaging in social bonding), and boundary-maintaining (refusing requests, enforcing limits). Key finding on appropriateness: human evaluators judged self-referential and relationship-building behaviors as less appropriate when exhibited by LLMs than by human conversational partners. Conversely, boundary-maintaining behaviors — refusals and limit-setting — were judged more appropriate from LLMs than from humans. The behaviors are pervasive across all four models but vary in prevalence by model and in intensity by user conversation goals and profiles. System prompts can meaningfully control all three behavior types, but calibration requires care — the study finds that system prompt interventions can produce unintended side effects on adjacent behaviors not explicitly targeted. The study uses both LLM-as-a-judge and human evaluation, across a range of conversation goals and user profiles. It provides empirical grounding for practitioners designing LLM personas, system prompts, and deployment guidelines — particularly around when human-like behaviors help versus harm user experience and trust. Published August 2026 by Sunnie S. Y. Kim, Margit Bowler, and Leon A. Gatys at Apple.