Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
| Source: Google Research Blog
Tags: Science One, Chain-of-Evidence, Google Research, autonomous research agents, AI hallucination, MLE-Bench, verifiability
Google Research's Science One Framework introduces Chain-of-Evidence (CoE), a verifiability standard for autonomous AI research that eliminates hallucinated references entirely and produces fully reproducible experimental scores, while achieving state-of-the-art on MLE-Bench and Parameter-Golf.
Details
Google Research published the Science One Framework, addressing a critical problem in AI-generated scientific papers: hallucination and verifiability failure. The core contribution is Chain-of-Evidence (CoE) — a framework analogous to database ACID properties — requiring every claim in an AI research artifact to carry a recorded evidence chain (completeness) and for each chain to genuinely support its claim (correctness).\n\nThe problem is real and documented: baseline autonomous research systems hallucinate up to 21% of their references and frequently misalign code descriptions with actual implementations. Systems like Sakana's AI-Scientist, AutoResearchClaw, and DeepScientist can generate complete-looking manuscripts with phantom citations, non-reproducible scores, and method-code mismatches.\n\nScience One demonstrates CoE eliminates phantom references entirely and produces fully verifiable experimental scores. The framework covers all claim types: citations, reported numbers, method descriptions, and conclusions — each must link to traceable evidence such as peer-reviewed papers, experimental log lines, or the code that actually ran. CoE Audit is a companion automated evaluation protocol that measures paper integrity against underlying code and experiments.\n\nCritically, Science One achieves this without sacrificing capability, reaching state-of-the-art on MLE-Bench and Parameter-Golf — two frontier benchmarks for autonomous ML research. The work comes from Google Cloud researchers Rui Meng and Tomas Pfister.