AI IDE List
Models / Compare

Claude Opus 5.5 vs Claude Opus 5

Opus 5.5 lowers standard input and output token prices by 20%. Use Opus 5 as a baseline and check for regressions on your existing tasks.

Reviewed 2026-09-25 · Official API facts and attributed benchmark results

Model specifications and standard API pricing
SpecificationClaude Opus 5.5Claude Opus 5
Context window1M1M
Max output128K128K
Input / 1M tokens$4$5
Output / 1M tokens$20$25
Cache read / 1M tokens$0.2$0.5
ReasoningAdaptive, always on · medium by defaultAdaptive · high by default

When to evaluate Claude Opus 5.5

A starting point for evaluating coding and agent workloads, with lower token rates than Opus 5 or Fable 5.1.

Benchmark settings differ from the default medium effort. Measure task completion and total tokens on your own workload.

Standard API rates. Batch input/output is 50% lower; fast mode is $8 / $40. Cache writes: $5 (5 min), $8 (1 hour) per million tokens.

Official specifications

When to evaluate Claude Opus 5

Useful as a regression baseline for applications already tuned to Opus 5.

Anthropic marks this as a legacy model and recommends considering Opus 5.5. Retest prompts, tool calls and costs before migrating.

Standard API rates. Batch input/output is 50% lower. Cache writes: $6.25 (5 min), $10 (1 hour) per million tokens.

Official specifications

Reported benchmark results

Anthropic launch benchmarks, September 22, 2026
BenchmarkClaude Opus 5.5Claude Opus 5
Terminal-Bench 4.0Opus 5.5: xhigh effort66.4%52.3%
FrontierCode v1.1Opus 5.5: max effort54.4%48%
CursorBench 4.0Opus 5.5: max effort57.8%46.6%
GDPval-AA v2.1Opus 5.5: max effort1846 Elo1708 Elo
AutomationBenchOpus 5.5: max effort40%26.9%
Humanity's Last ExamWith tools · Opus 5.5: max effort67.7%63.6%

Vendor-reported launch results, September 22, 2026. Opus 5.5 uses max effort except Terminal-Bench (xhigh). Other models' effort, harness and tool settings can differ; see the original methodology. These are not AI IDE List tests and are separate from Anthropic's medium-effort cost charts.

Benchmark source and methodology

Try the work that matters to you

Use the same task, inputs, tools and success criteria. Record the model version, effort, total cost, latency and any human fixes. Repeat tasks before choosing a model; a single demo does not establish a general winner.