Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
| Source: arXiv AI
Tags: AI-agents, benchmarks, identity, Claude, Astra, alignment
PAI-Bench reveals that deployed AI agents can recall explicit identity facts but struggle to express them implicitly: only 1 of 48 tested responses produced an implicit self-portrait, while Claude scored 12.5 percentage points below Astra on the benchmark.
Details
As AI agents are deployed with specific personas, versions, and role constraints, measuring identity fidelity — whether an agent actually expresses who it is supposed to be — becomes critical for governance and trust. PAI-Bench (Persistent Agent Identity Benchmark) is a provider-neutral evaluation covering recall, composition, behavioral enactment, resistance to overrides, persistence across sessions, lineage tracking, and role-conditioned updates.\n\nThe benchmark uses 16 synthetic profiles, 32 probes, and 3 independently initialized target configurations, yielding 1,536 retained responses. The core finding is a sharp dissociation: agents reliably recall explicit identity fields (48/48 atomic responses include direct-parent identifiers), but almost never produce an implicit self-portrait (only 1/48). Prompt-dependent component selection is the underlying mechanism — agents select which identity components to surface based on explicit cues rather than maintaining a stable expressed identity.\n\nAcross 8 profiles, adding explicit field cues increased joint presence of three identity identifiers from 0/8 to 7/8 under a 4-sentence instruction. A startup body-label substitution raised full-designation presence from 1/8 to 7/8. The study also demonstrates evaluator sensitivity: Claude scored a mean 12.5 percentage points below Astra on identical replayed responses, showing the benchmark distinguishes target behavior from evaluator behavior.\n\nThe protocol and checkpoints are publicly released for reproducible evaluation.