Early preview
O

OpenAI

GPT-5.2

Announced Dec 11, 2025

Documents, spreadsheets and complex projects

01

What could it do?

OpenAI reported improvements in knowledge work, long context, coding and tool use.

02

What changed?

A new GPT-5 generation focused on professional tasks.

WHY IT MATTERED

Expanded the range of work an assistant could help produce.

03

Where it fell short

This is a model series; Thinking and other configurations must be evaluated separately.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

Developer announcement or release log

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
153.4index points
Tested variant: GPT-5.2Source interval: 151.6 to 155.4Variant date in source: Dec 11, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1437rating points
Tested: gpt-5.2-highText Arena Overall · October 8, 2026Reported interval: 1433 to 144149,685 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
PUBLISHED BENCHMARK RESULT
72.8%% resolved
Tested: GPT 5.2 (high)500 tasks · mini-SWE-agent v2.0.0Run: Feb 17, 2026 · Effort: highPublished submission; not marked as checked by the benchmark team

Only mini-SWE-agent v2.0.0 submissions, one run per model. Reasoning effort varies and is labelled. These are published submissions, not a rerun by Leapscope. Latest included release: February 5, 2026; latest run: February 26, 2026.

Source: SWE bench Verified ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.