What happened?
In January 2024, Google DeepMind described AlphaGeometry, an AI system that writes proofs for geometry problems from the International Mathematical Olympiad (IMO), the world's top maths contest for school students DeepMind blog ↗. On a test set of 30 past Olympiad geometry problems, it solved 25 within the normal contest time limit DeepMind blog ↗. The work was published in the peer reviewed journal Nature Nature paper ↗.
Computers have been proving geometry facts for decades. The best older method, known as Wu's method, solved 10 of the same 30 problems Nature paper ↗. Human medalists on these problems averaged about 19.3 for bronze, 22.9 for silver and 25.9 for gold, so AlphaGeometry came close to an average gold medalist, but only on geometry DeepMind blog ↗. A Scientific American report added that GPT-4, a general chatbot model, solved none of the 30 Scientific American ↗.
What made this new was how the system learned. Instead of copying human proofs, the team generated about 100 million synthetic examples, starting from one billion randomly drawn diagrams, with no human demonstrations DeepMind blog ↗. The paper also says the proofs it produces can be read by people, not just checked by a machine Nature paper ↗.
What are the two halves?
The language model
A neural network, a program that learns patterns from examples, suggests a useful extra point, line or circle to add to the diagram. Mathematicians call these auxiliary constructions, and finding them is often the creative step DeepMind blog ↗.
The symbolic engine
A rule based deduction engine follows strict logical rules to derive every fact it can from the diagram. When it gets stuck, the language model adds one more construction and the loop repeats DeepMind blog ↗.
A system that guesses creatively and then checks every step with strict logic avoids much of the made up reasoning that chatbots are known for.
Leapscope interpretation of the reported result.How did AI help?
The AI did the problem solving itself: the language model proposed constructions and the symbolic engine produced the step by step proof DeepMind blog ↗. Humans built the system, wrote the formal rules of geometry it uses, and translated the test problems into its special formal language Nature paper ↗. Former Olympiad gold medalist and coach Evan Chen reviewed selected solutions and called the output "both verifiable and clean" DeepMind blog ↗.
Figures from the Google DeepMind announcement DeepMind blog ↗ and the Nature paper Nature paper ↗.
There are real limits. It only handles plane geometry, which is usually about two of the six problems in an IMO, so it could not sit a full contest DeepMind blog ↗. The authors note the test set is small, that only about 75% of IMO geometry problems can be written in its formal language, and that the human expert check covered only a few solutions Nature paper ↗. Some proofs are long: one featured solution runs to 109 logical steps DeepMind blog ↗.
Which fields could this affect?
The immediate value is in maths and AI research; wider uses are possible future value, and these connections are our assessment.
AI and automated reasoning research
The code and a trained model were released publicly, so researchers can study and build on the approach GitHub code ↗. It is a clear example of combining a learning system with a strict checker.
Explore scienceMaths education and competitions
Coaches can read its proofs and compare them with human methods DeepMind blog ↗. Scientific American reported it found a more general solution to a 2004 IMO problem Scientific American ↗.
Software and hardware verification
Checking that code or chips behave correctly also relies on step by step logic. Whether this style of system helps there has not been shown.
Explore softwareResearch mathematics
DeepMind presents broader reasoning across maths as a long term goal, not a current ability DeepMind blog ↗. This work solves known contest problems; it did not prove new theorems.
What has been checked?
The main evidence is a peer reviewed Nature paper with public code, plus a review of selected proofs by an Olympiad coach. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- It solved 25 of 30 past Olympiad geometry problems within the contest time limit, against 10 for the best older method Nature paper ↗.
- Its proofs are produced step by step by a logic engine and can be read by people Nature paper ↗.
- The code and model checkpoint are publicly available on GitHub GitHub code ↗.
Still unknown
- How well it would do on a larger or newer set of problems, since 30 is a small test Nature paper ↗.
- Whether the same approach works for other areas such as combinatorics, which the researchers suggest but have not shown here Scientific American ↗.
- How it handles the roughly 25% of IMO geometry problems that its formal language cannot express Nature paper ↗.
Evidence status: Published research. Stage: Usable. The code was released publicly.
From geometry to the whole contest
This is our suggested way to follow the work, not a promised timetable.
Can I use it today?
Yes, if you are comfortable with code. The AlphaGeometry code is on GitHub, with model weights downloadable, though matching the paper's results needs much larger settings and serious hardware GitHub code ↗. It is research code, not an app for everyday use.
A few things you might be wondering
Did AlphaGeometry win an Olympiad medal?
No. It was tested on 30 past geometry problems, not a full live contest, and geometry is only about a third of the IMO DeepMind blog ↗.
Is it just a chatbot like ChatGPT?
No. A language model only suggests constructions; a separate logic engine produces each proof step DeepMind blog ↗. GPT-4 solved none of the same problems in one comparison Scientific American ↗.
Did humans check its proofs?
Partly. Coach Evan Chen reviewed selected solutions, and the paper notes this expert check covered only a few problems DeepMind blog ↗ Nature paper ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01The developer's announcement with results, how the system works and expert comments.
The research paper by Trinh, Wu, Le, He and Luong, including methods and stated limitations.
An independent news report with comparisons to GPT-4 and comments from mathematicians.
The released code, model download script and notes on what is and is not included.