Back
SiTech
Mathematician Andreas Thom questions whether researchers can trust OpenAI with unpublished math
SiTech AI Team2 წთ. საკითხავი

Mathematician Andreas Thom questions whether researchers can trust OpenAI with unpublished math

In a post published on 9 September, mathematician Andreas Thom says an exchange with OpenAI left a key question open: whether unpublished mathematical work discussed with ChatGPT can reach training data or the reasoning process.

A question left open

The mathematician Andreas Thom has published a short thread arguing that OpenAI has not answered a basic question for working mathematicians: whether unpublished research discussed with its chatbot can end up being used by the company. The post, dated 9 September and labelled “1/3”, was written on the Mathstodon instance as a reply to another researcher.

Thom says the points came to him after reading up on the controversy around OpenAI’s non-sofic-group announcement, and that the episode reminded him of an exchange he had with the company after that result appeared.

The email to Sellke and Bubeck

Shortly after the surprising non-sofic-group finding, which used methods developed by Gábor Kun and by Thom himself, the mathematician wrote to Mark Sellke and Sébastien Bubeck. He and a colleague in Dresden had spent recent months discussing the expander matching problem and extensions of the work with Gábor Kun “actively … with ChatGPT”, he wrote, so they were “of course curious if that was part of the training data or accessible to the reasoning process”.

He added: “There is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.” Mark Sellke’s complete answer, as quoted in the post, was: “Regarding your conversations with ChatGPT: that did not happen.”

Two questions, one answer

Thom stresses that he had explicitly asked about two different things: first, whether the conversations had entered training data, and second, whether they were accessible to the solving process. In his reading, the categorical answer addresses only direct access under the second question; no qualification, explanation or evidence was provided for the first.

“I take this as dishonesty to say the least,” he writes.

What OpenAI says now

In the Buckmaster–Alpöge case, OpenAI says that no specific user data was accessed, but adds that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models”. For Thom, that leaves the question that matters to researchers unanswered: whether material shared with an AI system can be relied upon to remain unpublished.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.