CLARITY BEFORE COMPARISONS
What are we measuring?
There is no single test for everything an AI model can do. Choose a view to explore a specific kind of progress.
Capability growth over time
The Epoch Capabilities Index combines benchmark results using a statistical model. We reproduce its published estimates and source intervals. We do not run these evaluations ourselves. Scores are index points, not percentages or multiples of intelligence. The scale has no natural zero and differences should not be converted into percentage growth.
Colored lines show each company’s highest score so far among the models in this collection. Each dot keeps its actual source estimate, even when it falls below the line. The main view starts in 2023 because GPT 2, GPT 3 and the 2022 ChatGPT launch have no matching score in the source snapshot. The earlier models remain linked as unscored milestones. Family announcements are not assigned an arbitrary variant score.
ECI plots use the tested variant’s date from Epoch’s table, which can differ from the announcement date on our model story. Intervals can overlap, so small score differences are not decisive. Source data mixes independent and developer evaluations. All values come from the same fitted snapshot; later snapshots may revise older values. Download the unmodified snapshot. Data by Epoch AI, used under CC BY 4.0; we select 47 rows and plot them without rescaling.
Other benchmark views
- Artificial Analysis
- We use one October 8, 2026 snapshot of Intelligence Index v4.3.2. Twelve selected configurations are included; starred partial results and older index versions are excluded. Claude fallback configurations may use other models when safeguards intervene. Original methodology
- Arena
- We use 46 explicitly named variants from the October 8, 2026 Text Arena Overall snapshot. These are current ratings of older and newer models, not their launch day ratings. Preliminary labels, vote counts and reported intervals appear in the evidence panels. Arena leaderboard
- SWE bench Verified
- Eight published submissions using mini-SWE-agent v2.0.0 on the 500 Verified tasks. Effort settings can differ and are shown per point. The latest included model release is February 5, 2026; the latest run is February 26. These submissions are not marked as independently checked by the benchmark team. We do not mix later vendor results or other test suites into this curve. Benchmark source
Our experimental Leapscope Score
Each input is mapped to a common 0 to 100 scale. The three normalized values receive equal weight. If any input is missing, we do not calculate a combined score.
The combined score is not published. Current results do not provide matching model configurations across the three sources. Its tab shows the available components instead. The earlier preview anchors were demonstration fixtures; they are not applied to published results.
This score is not a measure of human intelligence or a scientific claim that one model is twice as capable as another. Overlapping tests and equal weighting can favor certain strengths.
Missing data stays missing
Choose Release timeline to see models and research without benchmark scores. There is no demo mode in the public chart. Models without comparable results appear in the unscored list. A blank value is never replaced with zero.
Reading circles and diamonds
Circles are model announcements. Diamonds are research milestones. Squares are access updates. Both appear in a separate events lane on the release calendar. A diamond never raises a model’s benchmark score. Lines in the timeline connect release dates from one company and do not measure intelligence. Nearby points move vertically within a lane so their dates remain accurate; use zoom or the entry list for crowded periods.
Research status comes from the linked publication or announcement. Published research, official grading, lab experiments and claims under review are different kinds of evidence. Source reviewed means we checked the announcement, not that we reproduced the experiment or proved the theorem. The October 6 OpenAI collection is one release with related result highlights; these highlights are not additional model releases or independently certified solved problems.
Coverage is curated, not exhaustive. Dates are public announcement dates unless otherwise stated. Events after Oct 7, 2026 are outside this snapshot.
Model stories and milestones
Dates may describe an announcement, paper publication, preview or full release. A model family is not a specific tested variant. Product milestones, developer claims and independently confirmed discoveries should be distinguished during source review.
The original spreadsheet notes are preserved. Draft model pages are excluded from search indexing until their content is reviewed. Every model has a permanent page with readable HTML, a distinct title and description. Read the benchmark guides or learn how the stories are prepared.