Funding better evaluations of AI’s impact on wellbeing

| Source: Anthropic News (community RSS)

Tags: Anthropic, Claude, AI safety, wellbeing, mental health, AI evaluation, grants

Anthropic is committing $5 million to fund independent researchers — clinicians, psychologists, and methodologists — to build open-source benchmarks measuring how AI models affect user wellbeing, covering mental health interactions, crisis conversations, and companionship dynamics.

Details

Anthropic is launching a $5 million grant program for external researchers to build open-source evaluations measuring how AI models affect user wellbeing. Grantees receive direct funding, model access, and technical support while retaining full editorial independence — all research will be published openly for the entire AI industry to use. The program targets a concrete measurement gap: wellbeing cannot be assessed in single-response evals because context matters enormously. Anthropic gives two examples: a user with a history of disordered eating asking about weight loss routines, and someone approaching a mental health crisis who does not disclose distress upfront. Standard accuracy benchmarks entirely miss these dynamics. Anthropic is sharing guidance from its Safeguards team on rigorous wellbeing evaluation design: clear pass/fail criteria, clinical expert involvement, and explicit tests for both overcompliance (harmful agreement) and overrefusal (unhelpfully blocking legitimate support). This criteria document is public, not gated behind the grant. The program is one of the first major lab-funded efforts explicitly bringing external clinicians and psychologists into AI safety evaluation — an acknowledgment that self-assessed safety benchmarks carry credibility limits. With EU and US regulators increasing scrutiny of AI's psychological impact, Anthropic is building independent evaluation infrastructure before mandates arrive.