Top mathematicians say LLMs are strong calculators but poor creative thinkers
| Source: THE DECODER
Tags: mathematical reasoning, LLM limitations, Timothy Gowers, DeepMind, benchmarks
Fields medalist Timothy Gowers and Princeton's Peter Sarnak credit LLMs with serious mathematical ability but identify a hard ceiling: models can combine known techniques across vast search paths but lack the intuition to select productive ones — the core skill behind genuinely new mathematics.
Details
Fields medalist Timothy Gowers and Princeton's Peter Sarnak have publicly assessed the mathematical capabilities of large language models — crediting them with real skills while defining a sharp ceiling. Gowers argues current models are effective at combining established techniques and exploring many possible paths, but the essence of mathematical creativity lies in knowing which of exponentially many paths are worth pursuing. That intuition, he says, LLMs lack. Sarnak's analysis goes further: LLMs can derive results from existing theory but fail to develop the abstractions that underpin major proofs when starting from elementary questions. This maps directly onto the bottleneck identified separately by DeepMind researcher Tom Zahavy in his paper 'LLMs Can't Jump,' which coins 'manipulative abduction' — the ability to invent new foundational assumptions with no prior linguistic representation — as the missing capability. Both assessments feed into ongoing debates about whether benchmark improvements on math tasks reflect genuine mathematical understanding or just better performance on well-defined problem spaces. The two mathematicians's critiques, backed by Zahavy's, suggest the latter. For practitioners, the implication is clear: LLMs are strong assistants for working within known mathematical territory but should not be treated as creative collaborators for frontier research. World models are suggested as a potential architectural direction forward.