What happened?
In May 2025, Google DeepMind introduced AlphaEvolve, a coding agent that uses Gemini language models to propose changes to computer programs and automated tests to score them DeepMind blog ↗. The best scoring programs are kept and used to prompt the next round, a process modelled on evolution DeepMind blog ↗. DeepMind says it improved parts of Google's computing systems and found new results on open maths problems DeepMind blog ↗.
An algorithm is a step by step recipe a computer follows. AlphaEvolve does not just write one recipe. It keeps a growing database of programs, asks Gemini models to suggest edits, runs each new version through an evaluator (a test that gives it a score), and keeps the best ones DeepMind blog ↗. Its predecessor, FunSearch, evolved short pieces of code, while AlphaEvolve can work on programs hundreds of lines long MIT Technology Review ↗.
DeepMind reported several uses inside Google. A new rule for its data centre scheduling system has run for over a year and recovers on average 0.7% of Google's worldwide computing power DeepMind blog ↗. It sped up one piece of Gemini's training code by 23%, cutting total training time by 1%, and suggested a simplification to a chip circuit that was planned for an upcoming Google TPU DeepMind blog ↗. In maths, it found a way to multiply two 4x4 grids of complex numbers with 48 multiplications, which the team calls the first improvement in this setting in 56 years arXiv paper ↗.
What are the three pieces?
The idea makers
Gemini Flash suggests many quick ideas and Gemini Pro adds fewer, deeper ones DeepMind blog ↗.
The evaluators
Automated tests run each program and score it on a clear measure, such as speed or the size of a result DeepMind blog ↗.
The program database
An evolving collection of past programs decides which ones are used to inspire the next round of ideas DeepMind blog ↗.
When a goal can be scored automatically, a language model plus a test can keep improving code that humans then read and use.
Leapscope interpretation of the reported result.How did AI help?
People pick the problem, write the starting program and build the evaluator that defines what better means DeepMind blog ↗. The AI then proposes and tests large numbers of code changes. On more than 50 open maths problems, DeepMind says it matched the best known answers about 75% of the time and improved on them about 20% of the time DeepMind blog ↗. A later study with outside mathematicians including Terence Tao tested it on 67 problems and found it rediscovered the best known solutions in most cases and improved several Tao et al. preprint ↗.
The first two figures are DeepMind's own DeepMind blog ↗ arXiv paper ↗; the third is from the study with outside mathematicians Tao et al. preprint ↗.
There are real limits. It only works where a computer can score the answer, so it cannot judge things like lab experiments that need human interpretation MIT Technology Review ↗. The mathematician Jakob Moosbauer said it offers little insight into how it reached its answers MIT Technology Review ↗. Records can also fall quickly: its 593 sphere arrangement in 11 dimensions still stands, but a doctoral student at Aalto University reported better bounds than AlphaEvolve in two other dimensions using his own method Popular Science ↗. Most Google infrastructure figures come from DeepMind itself and have not been checked by outsiders DeepMind blog ↗.
Which fields could this affect?
AlphaEvolve already affects computing and maths research, with broader uses possible later; these connections are our assessment.
Computing infrastructure
DeepMind reports that its scheduling and code changes run inside Google's systems. The figures are the company's own.
Explore softwareMathematics research
Mathematicians have used it to search for examples and bounds on open problems. Its outputs can be checked, but it does not explain why they work.
Explore scienceBusiness optimisation
Any task with a clear score, like routing or forecasting, could be a candidate. Results will depend on how well the goal can be tested automatically.
Explore softwareLab science
Problems that need human judgement or slow experiments are outside what it can score today. Linking it to real experiments is not shown in this work.
Explore scienceWhat has been checked?
The evidence is a developer announcement and technical preprint, news reporting with outside experts, and a follow up preprint co-written with independent mathematicians. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- A provably correct way to multiply 4x4 complex matrices with 48 multiplications, described in the technical paper arXiv paper ↗.
- New or matching results on many maths problems, examined in a study co-authored by outside mathematicians Tao et al. preprint ↗.
- Outside experts in matrix multiplication called the matrix result impressive and likely to be useful in practice MIT Technology Review ↗.
Still unknown
- How large the internal Google gains are when measured by someone outside the company DeepMind blog ↗.
- How well it works on problems where a good answer is hard to score automatically MIT Technology Review ↗.
- How long its maths records will last, since some related bounds have already been beaten by humans Popular Science ↗.
Evidence status: Research and deployment report. Stage: Usable. Available to Google Cloud customers, and outside mathematicians have co-published checks of its maths results.
From internal tool to wider use
This is our suggested way to follow the story, not a promised timetable.
Can I use it today?
Can I use it today? You can read the announcement and technical paper DeepMind blog ↗ arXiv paper ↗. Google later made AlphaEvolve generally available to Google Cloud business customers Google Cloud blog ↗, so it is a paid enterprise tool, not something an ordinary person can try for free.
A few things you might be wondering
Is AlphaEvolve just a chatbot writing code?
No. A language model suggests code, but every suggestion is run and scored by an automatic test, and only the best versions survive to the next round DeepMind blog ↗.
Did it solve famous maths problems?
It improved some known bounds and examples, such as the 11 dimensional sphere arrangement, rather than proving big theorems DeepMind blog ↗ Popular Science ↗. Terence Tao and colleagues describe it as a tool for exploring problems at scale Tao et al. preprint ↗.
Are the results checked by anyone outside Google?
Some are. The maths results can be checked, and outside mathematicians co-wrote a study of them Tao et al. preprint ↗. The data centre and training savings are reported by DeepMind itself DeepMind blog ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01The developer's announcement describing how the system works and its reported results.
The technical paper by Novikov and colleagues, including the 48 multiplication result.
News coverage with comments from outside mathematicians Jakob Moosbauer and Manuel Kauers.
Georgiev, Gómez-Serrano, Tao and Wagner test AlphaEvolve on 67 maths problems.
Reports new sphere packing bounds by an Aalto University doctoral student, compared with AlphaEvolve's.
Announces general availability of AlphaEvolve to Google Cloud customers.