The twilight of the chatbots
| Source: One Useful Thing (Ethan Mollick)
Tags: Ethan Mollick, Opus 4.7, Claude Fable, METR, autonomous agents, AI benchmarks, agentic AI
Ethan Mollick argues the chatbot era is ending — METR, the UK AI Security Institute, and Epoch all show AI capability growing at better-than-exponential rates, with Opus 4.7 autonomously building software in 14 hours for $251 that would take human engineers 2-17 weeks.
Details
The 'twilight of chatbots' refers to AI systems rapidly graduating from conversation to autonomous work. Mollick synthesizes multiple capability assessments: METR and the UK government's AI Security Institute measure AI in human programmer hours per prompt, while GDPval benchmarks against expert professionals — all show better-than-exponential growth curves.\n\nConcrete numbers ground the analysis: Epoch found Anthropic's Opus 4.7 working autonomously for 14 hours to produce software worth 2-17 weeks of human engineering effort, at a cost of $251 in tokens. Mollick's own experiments with Claude Fable showed 9 hours of autonomous execution on projects a human team would take over a week to complete.\n\nThe analysis distinguishes two tiers: frontier closed-source models from Anthropic, OpenAI, and Google (noting government interventions have blocked access to Claude Fable and GPT-5.6, the two most capable models as of June 2026); and open-weights Chinese models lagging 6-12 months behind but on their own exponential curve. The AA-Briefcase benchmark — simulating a complex multi-week consulting engagement — visualizes the gap.\n\nMollick is honest about limitations: the frontier remains jagged, AIs still fail specific tests, and inference costs vary widely. The directional trend is nonetheless unambiguous.