AI × MATHEMATICSExplained simply

An AI wrote gold medal level proofs in plain English.
Meet Gemini Deep Think.

Google DeepMind's model solved five of six problems at the 2025 International Mathematical Olympiad, within the contest time, and the IMO's own coordinators graded it. Here's what that means and what it does not.

WHERE THIS STANDS
  1. Claim
  2. Verified
  3. Usable
  4. In use

Graded under official competition rules.

What moves it next: Moves to Usable when code, a model or a tool becomes publicly available. How we decide

THE BREAKTHROUGH35 of 42 points at IMO 2025
THE TEAMGoogle DeepMind
WHERE IT STANDSOfficially graded; research version limited
01 · THE BREAKTHROUGH

What happened?

In July 2025, Google DeepMind announced that an advanced version of its Gemini Deep Think model scored 35 of 42 points at the International Mathematical Olympiad (IMO), a gold medal level score DeepMind blog ↗. It solved five of the six problems perfectly, writing its proofs in ordinary language within the 4.5 hour contest time limit DeepMind blog ↗. IMO coordinators graded the answers using the same criteria as for student solutions DeepMind blog ↗.

The IMO is the world's top maths competition for school students. In 2025, 641 students from 112 countries took part, around 10% reached gold level, and five students scored a perfect 42 AFP report ↗. IMO President Gregor Dolinar said: "We can confirm that Google DeepMind has reached the much-desired milestone" DeepMind blog ↗.

This was a big step from the year before. In 2024, DeepMind's AlphaProof and AlphaGeometry 2 scored 28 points, but needed humans to translate the problems into a formal computer language and took days of computing DeepMind blog ↗. OpenAI also announced a 35 point result for its own experimental model, but that was graded by three former medalists it chose, not by IMO coordinators AFP report ↗ TechRepublic ↗.

THE REASON TO BE EXCITED

A general purpose language model, not a special maths tool, produced complete proofs that official graders accepted under contest conditions.

Leapscope interpretation of the reported result.
02 · AI’S ROLE

How did AI help?

The model read the official problem statements and wrote the proofs itself, end to end in natural language DeepMind blog ↗. DeepMind says it used parallel thinking, meaning it explored several possible solutions at once before settling on one, and was trained with new reinforcement learning techniques, a form of learning by trial and reward DeepMind blog ↗. Humans prepared its setup: it was given a curated collection of high quality maths solutions and general tips on approaching IMO problems DeepMind blog ↗.

35 / 42points scored
5 of 6problems solved perfectly
4.5 hourscontest time limit met

Figures from the Google DeepMind announcement DeepMind blog ↗, confirmed in AFP reporting AFP report ↗.

The IMO checked the answers, not the system. DeepMind's own post says the review confirmed the answers were complete and correct but did not validate its system, processes or model DeepMind blog ↗. News reports added that organizers could not verify how much computing power the models used or whether humans were involved AFP report ↗. Also, the gold level model is not the one most people can use: the public version reaches bronze level on the same benchmark PYMNTS ↗.

03 · THE POSSIBILITIES

Which fields could this affect?

The immediate value is in maths and AI research; wider uses are possible future value, and these connections are our assessment.

Relevant now

AI research

It is a clear, independently graded test of reasoning in a general language model DeepMind blog ↗. It also set a precedent for AI labs submitting results to official graders.

Explore science
Relevant now

Maths learning

A version of Deep Think is available to Google AI Ultra subscribers in the Gemini app, though it is a faster, bronze level variant PYMNTS ↗. Students and teachers can test its explanations, with care.

Possible future use

Mathematical research

DeepMind gave the gold level version to a small group of mathematicians for feedback PYMNTS ↗. Whether it helps with real research questions has not yet been reported in these sources.

Explore science
A more distant possibility

Solving open problems

Contest problems have known solutions written by experts. Doing well on them does not show an ability to solve questions nobody has answered.

04 · THE EVIDENCE

What has been checked?

The evidence is official competition grading by IMO coordinators of the submitted answers, plus independent news reporting. There is no peer reviewed paper in these sources. Leapscope reviewed these sources; we did not repeat the experiments.

Shown so far

Still unknown

  • How much computing power was used; the IMO could not verify it AFP report ↗.
  • Whether the system itself works as DeepMind describes, since the IMO only checked the answers DeepMind blog ↗.
  • How the gold level version performs on research mathematics, rather than contest problems PYMNTS ↗.

Evidence status: Official competition grading. Stage: Verified. Graded under official competition rules.

05 · WHAT COMES NEXT

From contest gold to real research

  1. Hear from the mathematicians.Watch for reports from the trusted testers who received the gold level model.
  2. Compare under shared rules.Look for future contests where several AI labs are graded the same way, with computing use disclosed.
  3. Test on unsolved problems.See whether these models contribute to results that mathematicians did not already know.

This is our suggested way to follow the work, not a promised timetable.

Can I use it today?

Partly. Google AI Ultra subscribers can use a version of Deep Think in the Gemini app, but Google says it reaches bronze, not gold, level on the 2025 IMO problems PYMNTS ↗. The gold level version went only to a small group of mathematicians and academics PYMNTS ↗.

06 · QUICK QUESTIONS

A few things you might be wondering

Did the AI win a real gold medal?

No medal was given. The IMO coordinators graded its answers as gold medal standard, 35 of 42 points DeepMind blog ↗.

Was it the only AI to score gold level?

No. OpenAI also reported 35 points, but its answers were graded by former medalists it chose rather than by IMO coordinators AFP report ↗ TechRepublic ↗.

Can I use the exact same model?

Not the gold level one. The public Deep Think in the Gemini app is a faster version that reaches bronze level PYMNTS ↗.

THE READING LIST

Go straight to the sources

Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.

01
Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical OlympiadGoogle DeepMind · 21 July 2025

The developer's announcement with the score, grading details, IMO President quote and method.

02
Humans beat AI gold-level score at top maths contestAFP via Citizen Digital · 22 July 2025

News report on human results and the IMO's caveats about AI entries.

03
Google DeepMind Achieves Gold-Level Math Olympiad Performance, Matching OpenAITechRepublic · 22 July 2025

News report comparing the Google and OpenAI results and how each was graded.

04
Google Releases Two Versions of Gemini 2.5 Deep ThinkPYMNTS · 1 August 2025

Report on the public bronze level version and the gold level version given to mathematicians.

ONE DISCOVERY LEADS TO ANOTHER

Keep following the possibilities.

AI × MATHEMATICS

Silver medal standard on IMO problems

AI × MATHEMATICS

722 manuscripts. A new scale of AI mathematics.

FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.