What happened?
In December 2023, Google DeepMind described FunSearch, a method in which a large language model writes small computer programs and an automatic evaluator scores them, keeping only the best DeepMind blog ↗. It found new constructions for the cap set problem, a long standing puzzle in mathematics, and better rules for a packing task DeepMind blog ↗. The work was published in the peer reviewed journal Nature Nature paper ↗.
A cap set is a collection of points in a grid where no three points lie on a line, in a precise sense; the question is how large such a set can be Nature paper ↗. In dimension 8, FunSearch found a cap set of size 512, larger than the best previously known Nature paper ↗. NYU computer scientist Ernest Davis notes in a review that the previous record was 496 Davis review ↗. It also nudged a related measure, called the capacity lower bound, from 2.2180 to 2.2184, and later to 2.2202 Nature paper ↗.
The second task was online bin packing: placing items of different sizes into as few bins as possible, without knowing what comes next. FunSearch's programs beat two standard rules, first fit and best fit, on benchmark tests Nature paper ↗. Because FunSearch outputs readable code, mathematician Jordan Ellenberg said: "When I study them, I learn something" DeepMind blog ↗.
What are the three pieces?
The language model
A large language model trained on code (Codey, built on PaLM 2) proposes new versions of a short program Nature paper ↗. It is not told the full problem.
The evaluator
Automatic code runs each program and scores the result, which filters out the mistakes language models are known to make DeepMind blog ↗.
The evolutionary loop
The best programs go back into a pool and are used as starting points for the next round, much like breeding DeepMind blog ↗.
Pairing a creative but unreliable model with a strict automatic checker turned guesses into results anyone can verify.
Leapscope interpretation of the reported result.How did AI help?
Humans chose the problems, wrote the scoring code and a starting program, and decided which small part of the program the AI could change DeepMind blog ↗. The language model then generated very many variations, on the order of a million samples, and the evaluator kept the best Nature paper ↗. The improvement to 2.2202 relied on a symmetry that Ellenberg spotted after reading FunSearch's code, so it was a human and AI effort Davis review ↗.
Figures from the Nature paper Nature paper ↗.
Results were hard to reproduce: only 4 of 140 runs found the size 512 cap set Nature paper ↗. Ernest Davis argues the language model's role is narrow and the gains are modest, noting that a version with no language model also reached a strong result, only much more slowly Davis review ↗. The method only suits problems with a fast, detailed scoring system, so tasks like writing proofs are outside its reach Nature paper ↗.
Which fields could this affect?
The immediate value is in maths and algorithm research; wider uses are possible future value, and these connections are our assessment.
Combinatorics research
The new constructions are published and can be checked by anyone GitHub ↗. Readable code also gave a mathematician new ideas about the problem DeepMind blog ↗.
Explore scienceAlgorithm design
The bin packing programs beat standard rules on benchmarks Nature paper ↗. The discovered programs and an evaluation suite are public GitHub ↗.
Explore softwareLogistics and scheduling
Packing and scheduling problems appear in shipping and computing. Gains on benchmarks have not been shown in real operations.
Solving big open problems
The authors note a large gap remains between the best known lower and upper bounds for cap sets Nature paper ↗. These are improvements to specific constructions, not a full solution.
What has been checked?
The evidence is a peer reviewed Nature paper, with the discovered programs and sets released publicly, plus an outside critical review. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- A cap set of size 512 in dimension 8, larger than previously known Nature paper ↗.
- Bin packing programs that outperformed first fit and best fit on benchmarks Nature paper ↗.
- The discovered programs and sets are on GitHub for anyone to check GitHub ↗.
Still unknown
- How much of the success comes from the language model rather than the search loop, which Davis questions Davis review ↗.
- Whether the general cap set question can be solved; the gap between bounds is still large Nature paper ↗.
- How widely the method applies, since it needs a fast and detailed scoring system Nature paper ↗.
Evidence status: Published research. Stage: Verified. The new mathematical constructions are published and can be checked independently.
From better constructions to real insight
This is our suggested way to follow the work, not a promised timetable.
Can I use it today?
Partly. The discovered programs, bin packing heuristics and a basic version of the FunSearch loop are on GitHub and run in Google Colab GitHub ↗. The release does not include the language model, the safe code runner or the large scale setup used in the paper GitHub ↗.
A few things you might be wondering
Did FunSearch solve the cap set problem?
No. It found larger cap sets in some cases and improved a lower bound, but the general question is still open Nature paper ↗.
Did the AI do this on its own?
No. Humans set up the problem and scoring, and one improvement came from a symmetry a mathematician spotted in its code DeepMind blog ↗ Davis review ↗.
Why does writing code matter?
Code can be run and checked automatically, which catches the errors language models make, and people can read it to understand the idea DeepMind blog ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01The developer's announcement explaining the method, the cap set and bin packing results.
The research paper by Romera-Paredes and colleagues with full results and limitations.
The discovered programs and sets, plus a basic implementation of the search loop.
An outside computer scientist's critical assessment of FunSearch's results and claims.