HN
Today

More questions about whether researchers can trust OpenAI with unpublished math

A mathematical mystery ignites a firestorm of ethical debate as OpenAI is accused of potentially using researchers' private chats to 'scoop' them on groundbreaking proofs. This story highlights the growing tension between rapid AI advancement and academic integrity, leaving many to question the future of collaborative scientific discovery.

697
Score
640
Comments
#1
Highest Rank
19h
on Front Page
First Seen
Sep 10, 4:00 PM
Last Seen
Sep 11, 10:00 AM
Rank Over Time
8434521224322161918161514

The Lowdown

The world of mathematics is currently embroiled in a heated controversy surrounding OpenAI's alleged use of unpublished research data from prominent mathematicians. Andreas Thom, a mathematician, explicitly questioned OpenAI about whether his conversations with ChatGPT regarding complex mathematical concepts were used to train its models or were accessible to its solving processes.

  • Thom's email to OpenAI staff highlighted a "frankly unacceptable lack of transparency" after OpenAI announced a surprising mathematical finding that seemed related to his ongoing work.
  • OpenAI's initial, categorical denial of direct access to conversations for solving was later qualified, stating they "cannot rule out that de-identified data derived from their usage of our products helped improve our models," which Thom perceived as dishonest.
  • This incident is compounded by a separate but related controversy involving Tristan Buckmaster and Levent Alpöge, who were reportedly nearing a solution to the Navier-Stokes problem when OpenAI learned of their progress and, allegedly, raced to publish its own AI-generated solution.
  • Reports suggest OpenAI then attempted to pressure Buckmaster into omitting his Anthropic-affiliated collaborator, Alpöge, from his publication, further fueling accusations of predatory corporate behavior and a disregard for academic norms.
  • Critics also point to the immense computational resources (estimated at millions of dollars) OpenAI deployed, suggesting a deliberate effort to "scoop" human researchers rather than engage in true collaboration.

This unfolding narrative raises profound ethical questions about intellectual property in the age of large language models, the trustworthiness of AI developers, and the potential chilling effect on academic researchers who might otherwise collaborate with AI tools. It underscores a significant tension between technological acceleration and the established principles of scientific attribution and open inquiry.

The Gossip

Data Dilemmas and Developer Doubts

Many commenters express deep skepticism about OpenAI's data privacy claims, particularly concerning the use of user conversations for model training, even with opt-out options. They highlight that OpenAI's assurances are often vague or seem to contradict observed behavior (e.g., opt-out toggles resetting). There's a strong sentiment that if data isn't used for training, it's either negligence on OpenAI's part or a deliberate act of using "de-identified" data, which still constitutes a breach of trust. This extends to the broader concern that any data fed to cloud AI services is effectively fair game for the provider.

Academic Accusations and Attribution Anxieties

The discussion delves into the ethical implications for academic research, viewing OpenAI's actions as a form of "scooping" or intellectual property laundering. Many argue that OpenAI's race to solve problems that human researchers were actively working on, coupled with alleged attempts to control attribution, undermines the communal process of scientific discovery. Some defend OpenAI, suggesting the models are genuinely capable and that human input is merely a small part of a larger breakthrough. However, a prevailing concern is that AI companies, driven by profit and IPO hype, are exploiting academic labor and diminishing the value of human intellectual pursuits.

The Overhang & Our Own Obsolescence

A significant thread revolves around the concept of the "overhang" in mathematics: a vast body of existing knowledge and low-hanging fruit that AI can exploit to make rapid discoveries by connecting disparate ideas. This leads to concerns that AI might not be truly innovative but merely efficient at synthesizing existing knowledge, thereby 'plucking' problems humans would eventually solve. This prompts existential questions for mathematicians (and by extension, other intellectual professions) about their future relevance and whether human creativity will be stifled or redirected if AI consistently outpaces them in discoverable areas.