Early preview
A

Anthropic

Claude Opus 4.1

Announced Aug 5, 2025

Stronger software agents

01

What could it do?

Anthropic reported improvements in coding, reasoning and agent tasks.

02

What changed?

An update to Opus 4.

WHY IT MATTERED

More capable coding within the existing Claude tools.

03

Where it fell short

Developer reported improvements; this site has not run an independent evaluation.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

Developer announcement or release log

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
144.1index points
Tested variant: Claude Opus 4.1Source interval: 141.7 to 146.0Variant date in source: Aug 5, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1450rating points
Tested: claude-opus-4-1-20250805-thinking-16kText Arena Overall · October 8, 2026Reported interval: 1447 to 145349,555 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.