Independent benchmark evidence

The frontier moved. Here is the evidence.

GPT-5.6 Sol (pro, max) currently leads the composite dataset at 162.1 ECI. Individual benchmarks show uneven progress: hard reasoning, autonomous task duration, and evaluation durability remain distinct constraints rather than one finish line.

Current composite frontier
162.1 ECI
Observed leader
GPT-5.6 Sol (pro, max)
Evidence basis
6 public measures
Evidence-weighted, not a countdown. See how the observatory separates measurement from speculation.

Epoch Capabilities Index

Frontier index

162.1 +1.1 ECI frontier change

A stitched general-capability scale built from more than 50 benchmarks.

Frontier index frontier over timeEpoch Capabilities Index rose to 162.1, led by GPT-5.6 Sol (pro, max). Use the time scrubber below for yearly values.20192020202120222023202420252026
2019 0.0 GPT-2 (1.5B)

Useful for comparing the frontier over long periods, but not a probability or countdown to AGI.

Inspect source data

Optional interpretation

Ask GPT-OSS to explain this signal

The model receives the selected public record, not an open web prompt. It cannot alter the score.

01 / Public record

Signals shaping the frontier

Each row is the latest observed leader for one measure. Expand it to see the prior-frontier comparison and source.

Epoch Capabilities Index GPT-5.6 Sol (pro, max) defines the observed frontier index frontier 162.1 on Epoch Capabilities Index. Useful for comparing the frontier over long periods, but not a probability or countdown to AGI. +1.1 ECI High confidence

GPT-5.6 Sol (pro, max), reported by OpenAI, improved the observed frontier by +1.1 ECI relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data
ARC-AGI-2 GPT-5.6 Sol (Max) defines the observed generalization frontier 92.5% on ARC-AGI-2. Strong performance can reveal flexible reasoning, while benchmark saturation can weaken the signal. +2.5 pp Medium confidence

GPT-5.6 Sol (Max), reported by OpenAI, improved the observed frontier by +2.5 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data
FrontierMath GPT-5.5 Pro pre-release (high) defines the observed reasoning frontier 52.4% on FrontierMath. Higher scores show stronger hard-problem solving; they do not establish broad reliability. +0.7 pp Medium confidence

GPT-5.5 Pro pre-release (high), reported by OpenAI, improved the observed frontier by +0.7 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data
SWE-bench Verified Claude Opus 4.7 (max) defines the observed software work frontier 83.5% on SWE-bench Verified. Measures bounded repository tasks, not unattended ownership of production systems. +4.8 pp Medium confidence

Claude Opus 4.7 (max), reported by Anthropic, improved the observed frontier by +4.8 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data
METR task horizon (80%) Gemini 3.1 Pro Preview defines the observed autonomy frontier 1.5 hr on METR task horizon (80%). Longer horizons matter, but a benchmark task is not the same as safe, persistent agency. +20 min Medium confidence

Gemini 3.1 Pro Preview, reported by Google DeepMind, improved the observed frontier by +20 min relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data
Video-MME video-SALMONN 2+ defines the observed multimodal frontier 79.7% on Video-MME. Captures one form of visual-language understanding, not embodied world modelling. +4.7 pp Medium confidence

video-SALMONN 2+, reported by ByteDance, improved the observed frontier by +4.7 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.

Open Epoch AI data

03 / Method

Evidence-weighted, not a countdown.

Road to AGI displays the published Epoch Capabilities Index alongside individual benchmark frontiers. It does not convert those measurements into an AGI probability, arrival date, or hidden proprietary score.

What the index does

  1. 01 Use dated benchmark records and preserve the original unit.
  2. 02 Show the frontier and the previous frontier so movement is inspectable.
  3. 03 Keep capability evidence separate from unresolved reliability questions.
  4. 04 Link every displayed signal to the underlying public source.

Update cycle

The public archive is checked every day at 04:17 UTC. Automation rebuilds the committed evidence snapshot, and a connected Cloudflare Pages project deploys the refreshed result from the production branch.

Where the model stops

Groq-hosted GPT-OSS is optional and only explains already-selected evidence. It does not choose sources, alter scores, or determine the frontier.

Primary references

Road to AGI

A living atlas of measured frontier capability and the evidence gaps that remain.