What happened?
On 20 November 2025, OpenAI and outside researchers released a paper of short case studies describing how GPT-5 contributed to ongoing research OpenAI blog ↗. It covers mathematics, physics, astronomy, computer science, biology and materials science, and includes four new mathematics results that the human authors checked arXiv paper ↗.
The collaborators came from places including Vanderbilt, UC Berkeley, Columbia, Oxford, Cambridge, Lawrence Livermore National Laboratory and The Jackson Laboratory OpenAI blog ↗. In one biology case, immunologist Derya Unutmaz showed GPT-5 Pro an unpublished chart; it proposed a mechanism for the effect and suggested an experiment whose outcome matched results his lab already had OpenAI blog ↗. In mathematics, GPT-5 proposed an estimate that Mehtaab Sawhney and Mark Sellke turned into a complete proof of an Erdős problem, one of many open questions posed by the mathematician Paul Erdős OpenAI blog ↗.
The paper came a month after a public misstep. In October 2025, an OpenAI manager posted that GPT-5 had found solutions to 10 previously unsolved Erdős problems The Decoder ↗ eWeek ↗. Thomas Bloom, who runs the Erdős problems website, called this "a dramatic misinterpretation": the model had found existing papers that already solved them The Decoder ↗. The November paper separates literature search from new results more carefully OpenAI blog ↗.
What kinds of help are described?
Finding past work
GPT-5 located existing papers that answered questions listed as open OpenAI blog ↗. This is useful, but it is not new discovery The Decoder ↗.
Working alongside experts
In physics, it rebuilt a hidden symmetry in a black hole equation, but only after a simpler warm up problem; it first failed on the full task OpenAI blog ↗.
New results
The paper reports four new mathematics results, verified by the human authors arXiv paper ↗.
Experts in several fields report that a general AI model sped up parts of their real work. Writing down failures as well as successes makes the claims easier to judge.
Leapscope interpretation of the reported result.How did AI help?
Researchers chose the problems, guided the model with prompts and warm up questions, and checked every result OpenAI blog ↗. GPT-5 contributed specific steps: a key idea, a short proof, a counterexample, a calculation check or a pointer to earlier papers OpenAI blog ↗. The authors describe these contributions as modest in scope arXiv paper ↗.
From the arXiv paper's abstract arXiv paper ↗.
OpenAI states that the case studies are curated examples, not a systematic sample, and do not show the full range of failures OpenAI blog ↗. It also says the model can invent citations, mechanisms or proofs that look plausible, can miss subtle points and sometimes fails to credit the sources of ideas OpenAI blog ↗. In one math case, it reproduced an argument from an earlier paper without citing it OpenAI blog ↗.
Which fields could this affect?
The immediate value is help for working researchers; bigger effects are possible later. These connections are our assessment.
Mathematics
GPT-5 helped with proofs and with finding past results OpenAI blog ↗. Experts still had to check each step.
Explore scienceLiterature search
Finding scattered past work was one of its clearest uses The Decoder ↗. Results still need checking because the model can invent citations OpenAI blog ↗.
Explore scienceBiology and medicine research
The immunology case showed a useful suggested mechanism and experiment OpenAI blog ↗. It did not involve any treatment.
Explore healthcareAI led research
OpenAI says the model is not autonomous and does not run projects on its own OpenAI blog ↗. Independent research by AI is not shown here.
What has been checked?
The evidence is a set of case studies written by OpenAI staff and outside collaborators and posted as a preprint, not peer reviewed together. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- Human authors verified four new mathematics results that GPT-5 contributed to arXiv paper ↗.
- GPT-5 Pro suggested an immune cell mechanism and an experiment that matched existing lab data OpenAI blog ↗.
- GPT-5 located published solutions to problems listed as open The Decoder ↗.
Still unknown
- How often GPT-5 helps compared with how often it fails, since the cases were selected OpenAI blog ↗.
- Whether independent researchers see similar benefits in their own work.
- Whether journals will judge the new results significant after peer review.
Evidence status: Case studies. Stage: Claim. Selected case studies from the developer.
From examples to evidence
This is our suggested way to follow the story, not a promised timetable.
Can I use it today?
The paper is free to read on arXiv arXiv paper ↗. GPT-5 and GPT-5 Pro are available through OpenAI's paid products, but results like these needed expert guidance and checking OpenAI blog ↗.
A few things you might be wondering
Did GPT-5 solve 10 famous unsolved math problems?
No. That October 2025 claim was withdrawn: the model had found existing solutions in the literature The Decoder ↗. The later paper reports four new math results checked by humans arXiv paper ↗.
Did GPT-5 do research on its own?
No. Researchers set the problems, guided the model and checked the results OpenAI blog ↗.
Are these results typical?
OpenAI says they are curated examples, not a systematic sample OpenAI blog ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01OpenAI's summary of the case studies, collaborators and stated limits.
The full paper by Bubeck, Gowers, Sellke and others.
Reports the withdrawn Erdős claim and Thomas Bloom's correction.
Reports the October 2025 Erdős claims and the reaction from mathematicians.