Thursday, 10 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 10 September 2026 at 14:08

Second mathematician accuses OpenAI of opacity over training data origins

Mathematician Andreas Thom has gone public with concerns that his ChatGPT conversations may have contributed to OpenAI's announced result on non-sofic groups, saying the company's answers about data use were misleading.

Foto: The Verge

A second mathematician has publicly challenged OpenAI over the origins of the data behind its recent string of mathematical breakthroughs. Days after a dispute erupted over whether OpenAI's models drew on unpublished research, Andreas Thom used Mastodon to accuse the company of dishonesty and a lack of transparency.

Thom's own field, non-sofic groups — infinite mathematical structures that cannot be approximated by finite ones — was the subject of one of ten results OpenAI announced with fanfare last month. OpenAI initially failed to credit recent work by Thom and colleague Gábor Kun, then quietly revised its writeup after criticism from the mathematical community.

Thom said he began questioning his own interactions with the company after NYU professor Tristan Buckmaster publicly raised similar concerns about whether OpenAI's models benefited from his use of its Codex tool. Thom said he was struck by how precisely OpenAI's system had used techniques that were, at the time, neither obvious nor the most promising path to a solution. He emailed OpenAI researchers Sébastien Bubeck and Mark Sellke asking whether his ChatGPT conversations had entered the training data or were accessible to the model's reasoning process.

The response he received, he said, only addressed whether his conversations could be directly accessed — not whether they had been absorbed into OpenAI's broader training datasets. Thom argued that ordinary researchers have no way to verify this themselves, and that the burden falls on OpenAI to disclose the relevant data practices if it wants to deny using such material.

A similar ambiguity marked OpenAI's account of its announced solution to the Navier-Stokes problem, one of mathematics' Millennium Prize challenges. The company denied directly accessing specific users' unpublished work but stopped short of ruling out that de-identified data from user interactions could have indirectly improved its models. Thom said de-identification strips a name but not the intellectual substance of a mathematical idea.

OpenAI did not immediately respond to a request for comment from The Verge. The episode has unsettled mathematicians, who fear that such practices could push researchers toward greater secrecy, wary that even rumors of a breakthrough might trigger a race against a well-resourced tech company.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category