Early preview
D

DeepSeek

DeepSeek R1

Announced Jan 20, 2025

Reasoning with downloadable weights

01

What could it do?

Work on reasoning tasks such as mathematics and code using a model trained with reinforcement learning.

02

What changed?

DeepSeek reported strong reasoning results and released model weights, plus smaller distilled models.

WHY IT MATTERED

Open reasoning

Downloadable weights gave developers a route to inspect and host a reasoning model themselves.

03

Where it fell short

A smaller distilled model is a different model. Its performance cannot be assumed to match the full R1.

Which release does this page cover?

Original January 20, 2025 release; the paper was first submitted on January 22. Check sources

Access
Downloadable weights and hosted service
Model distinction
Full R1 versus distilled variants
Paper date
January 22, 2025

What this meant in practice

Access to weights changes what developers can build and control. It does not eliminate hardware costs or establish that every deployment will reproduce a published score. Capability and accessibility deserve separate places on a progress timeline.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Compare reasoning approaches

Give two precisely identified models the same logic task and budget. Verify the final answer, record failed attempts and separate the full model from smaller variants.

Common question

Are all models with R1 in their name equivalent?

No. The release includes the main model and distinct distilled variants based on other model families.

THE USEFUL CONTEXT

Understanding DeepSeek R1 and its variants

Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.

What the research actually studied

DeepSeek’s research examines how reinforcement learning can encourage reasoning behavior, including checking and revising an approach. The original release included a main model and distinct smaller models. A family name is therefore not a sufficient description of the system used in a comparison.

A useful analogy is a product range sold under one brand. Shared branding does not make every item interchangeable. If you see an R1 result, look for the exact model name, size or deployment identifier and the evaluation setup. Without those details, the result may not describe the version available to you.

DeepSeek R1 research paper

How to judge a reasoning answer

Consider a small scheduling puzzle with a known solution. Give each system the same constraints and ask for a schedule plus a short explanation. Check every constraint against the final schedule. An answer can sound systematic while assigning a person to two places at once.

Then change one constraint and repeat. Keep track of invalid solutions and correct refusals when no solution exists. This is more informative than choosing a single impressive response. A long explanation should not receive credit merely for being long; what matters is whether the result satisfies the task. This exercise is illustrative, and we have not used it to assign a score to R1.

Downloadable weights and practical usefulness

Having access to weights creates choices about deployment. Those choices still need to be evaluated in context. For a proposed local setup, write down the intended hardware, expected workload and the person responsible for keeping it running. For a hosted setup, record the provider and exact version.

A sensible comparison includes operating effort alongside answer quality. If a task only runs occasionally, the cheapest theoretical per request cost may not be the cheapest arrangement overall. Conversely, a team with an established system may value control differently. This is a framework for asking the right questions, not a pricing claim or a recommendation that one deployment is always better.

FOLLOW THE EVIDENCE

What to watch next

Changes that would make this story worth revisiting:

  • Evidence about the exact R1 variant and configuration someone can actually run.
  • Repeated task evaluations that report incorrect answers, response time and the cost of retries together.

Questions about this milestone

Is a smaller distilled R1 model the full R1?

No. The release distinguishes the main model from smaller distilled models. Keep their scores and deployment requirements separate.

Does more reasoning text prove a better answer?

No. Inspect the final result against explicit criteria. A concise correct result can be more useful than a long explanation with an unnoticed error.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models. The announcement was January 20, 2025; the paper appeared January 22.

DeepSeek: R1 release announcement DeepSeek: R1 research paper

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
139.0index points
Tested variant: DeepSeek-R1Source interval: 136.3 to 140.3Variant date in source: Jan 20, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1398rating points
Tested: deepseek-r1Text Arena Overall · October 8, 2026Reported interval: 1393 to 140318,524 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.