Import AI 469: Science AI; RSI simulator; and Zuck’s technological pessimism

| Source: Import AI (Jack Clark)

Tags: DiG-bench, benchmark, Jack Clark, Import AI, recursive self-improvement, exploration

Import AI #469 highlights DiG-bench, a 70-game benchmark testing AI's ability to discover hidden rules through exploration — current frontier models fail all tiers — alongside commentary on recursive self-improvement simulators and Mark Zuckerberg's skepticism about near-term AI progress.

Details

Import AI #469 leads with DiG-bench (Discovery in Games), a 70-game benchmark developed by researchers from Oxford, Princeton, MIT, Swiss AI Lab, and Inria, with contributions from Juergen Schmidhuber. Unlike static benchmarks, DiG-bench games hide both their rules and objectives — the player must discover them through interaction alone. The games are split into seven difficulty tiers; every game has been beaten by at least one human, but none by today's frontier AI. The benchmark is deliberately contamination-resistant: 49 of 70 games are kept private, and all are handcrafted rather than procedurally generated. Games are text-native, making them directly accessible to language models without vision requirements. The practical goal is measuring 'creative intuition' — a model's ability to update its priors from novel environmental feedback rather than pattern-match training data. The newsletter also touches on RSI (Recursive Self-Improvement) simulator frameworks and Mark Zuckerberg's expressed pessimism about near-term AI progress — a notable contrast given Meta's aggressive model investment in 2026. Jack Clark co-founded Anthropic; Import AI carries significant signal about frontier lab thinking and carries some of the most thoughtful AI research synthesis available in newsletter form.