Mathematicians want proof OpenAI didn’t use their work

| Source: The Verge AI

Tags: OpenAI, training data, intellectual property, copyright, mathematics, data transparency

Two mathematicians publicly challenge OpenAI over training data transparency: Andreas Thom accuses the company of 'dishonesty' after its AI produced results tightly tied to his non-sofic group techniques, following similar complaints from NYU's Tristan Buckmaster about undisclosed use of his Codex interactions.

Details

The controversy erupted after OpenAI last month announced 10 mathematical breakthroughs, including a result on non-sofic groups — infinite mathematical structures that cannot be approximated by finite ones — that it acknowledged was built heavily on work by Andreas Thom and Gábor Kun. OpenAI quietly amended its writeup only after public criticism for failing to credit them. Thom says he grew suspicious of OpenAI's 'detailed command' of his techniques, which he describes as neither the most obvious nor the most promising approaches at the time. He emailed OpenAI researchers Sébastien Bubeck and Mark Sellke asking whether his ChatGPT conversations entered training data or the model's reasoning process. The response addressed only whether conversations could be 'directly accessed' — not whether they entered training data pools. Thom calls this 'dishonesty to say the least.' This follows Tristan Buckmaster of NYU, who first raised public questions about whether his Codex usage seeded OpenAI's mathematical work. The two cases share a critical structural problem: researchers have no technical means to audit OpenAI's training pipeline. They cannot verify or disprove the company's claims about data provenance, and OpenAI's responses have consistently sidestepped the core question. The pattern is drawing sustained scrutiny: AI labs claiming mathematical breakthroughs that may have been seeded by the researchers now demanding answers. Calls for mandatory training data disclosure or third-party audits are likely to intensify.