Book a demo →
← All insights AI News

Claude Opus 5.5: The Frontier Gets Cheaper, Not Just Better

Manuel RiveiroSEO & GEO consultant · Truffle development partner

Claude Opus 5.5: The Frontier Gets Cheaper, Not Just Better

Anthropic released Claude Opus 5.5 on 22 September 2026. The obvious question is whether it beats GPT-6 Astra. The more useful one is what happens to the price.

Cheaper and faster

According to Anthropic, 22 September 2026, Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Anthropic’s own framing: “40% less than Opus 5 on typical workloads,” alongside 30% faster output. Cache reads drop from $0.50 to $0.20, cache writes from $6.25 to $5.

That’s not a footnote. When a frontier model gets better and cheaper at the same time, the math around using it changes: more AI-assisted work per dollar, more queries, more use cases that used to cost too much to run. Anthropic also lists a fast-mode variant at $8 input and $40 output, a separate option for teams that want to trade some of that cost saving for raw speed.

Pricing vs Opus 5 (per 1M tokens, according to Anthropic, 22 Sep 2026)

InputOutputCache readCache write
Opus 5.5$4$20$0.20$5
Opus 5$5$25$0.50$6.25
Opus 5.5 fast mode$8$40

Where it performs better

On the benchmarks Anthropic itself published (22 September 2026), Opus 5.5 leads in most categories. Terminal-Bench 4.0, an agentic coding test: 66.4% against 52.3% for Opus 5, 57.9% for GPT-6 Astra, and 55.8% for Fable 5.1. Humanity’s Last Exam with tool use: 67.7%, up from 63.6%. GDPval-AA v2.1, a knowledge-work benchmark: 1846 Elo against 1708.

On OSWorld 2.0, a test of unscripted computer use (opening applications, navigating interfaces, finishing tasks with no step-by-step script), Opus 5.5 reaches 81.8%, up from 74.0% for Opus 5.

Anthropic backs this with two examples: a 680,000-line code migration completed in a day instead of weeks, and a C-to-Rust port of HAProxy finished in 9.5 hours at 51% lower cost than Fable 5.1. Those are single cases, not an average: they show what kind of task benefits, not how much any given task will.

Benchmark comparison of Claude Opus 5.5, Opus 5, GPT-6 Astra and Fable 5.1 across Terminal-Bench 4.0, Terminal-Bench-Science 0.1, Humanity's Last Exam, OSWorld 2.0 and AutomationBench, with GPT-6 Astra leading Terminal-Bench-Science and AutomationBench Source: Anthropic, 22 September 2026. “n/a” = not published by Anthropic.

Benchmark comparison (according to Anthropic, 22 Sep 2026)

BenchmarkOpus 5.5Opus 5GPT-6 AstraFable 5.1
Terminal-Bench 4.066.4%52.3%57.9%55.8%
Terminal-Bench-Science 0.158.7%29.0%64.6%52.6%
Humanity’s Last Exam (tools)67.7%63.6%57.2%65.6%
OSWorld 2.081.8%74.0%n/a80.7%
AutomationBench40.0%26.9%41.4%31.4%
GDPval-AA v2.1 (Elo)1846170815421735

The part a vendor would rather skip

GPT-6 Astra leads on two of these benchmarks. On AutomationBench, a business-process test, it scores 41.4% against 40.0% for Opus 5.5, and Anthropic notes Opus 5.5 ran that one without fallback models. On Terminal-Bench-Science 0.1 the gap is wider: 64.6% for Astra against 58.7% for Opus 5.5. Anthropic publishes both numbers itself, and they belong in any honest account of the release: no frontier model wins every category anymore, and the gap between the leading models keeps narrowing overall.

Safety, the underrated part

Anthropic reports Opus 5.5 as its best-performing model yet in an automated behavioral audit spanning roughly 2,000 scenarios. Attempts to cross defined boundaries occur 85% less often than with Opus 5 or Mythos 5.1. On prompt-injection resistance, it matches Fable 5.1, the lowest success rate recorded so far. External testing came from METR and Frontier Design, an EU AI Act-compliant watermark is included, and thinking mode can’t be turned off. None of that shows up in a benchmark table, but it’s the part that determines whether a business can actually put the model into a workflow that touches customers.

Our own test

Beyond reading Anthropic’s benchmarks, we ran our own: on 24 September 2026 we queried Opus 5.5 through our OpenRouter access with a typical business calculation, pricing a €89 jacket at a 30% margin and adjusting against a competitor priced 15% lower. The model answered correctly in 5.8 seconds. That’s one sample, not a benchmark, but it’s the kind of task we actually run these models on.

Early reactions

Since the 22 September 2026 release, discussion on Hacker News (800+ comments) and Reddit’s r/ClaudeCode has read noticeably warmer than it did for Opus 5. The recurring theme isn’t raw performance, it’s tone: several commenters describe Opus 5.5 as more direct and less wordy than its predecessor, to the point that some now call it usable as a daily driver, exactly the complaint that dogged Opus 5.

That’s forum sentiment from the first 48 hours, not a test, and we’re reporting it as exactly that: an early read, not a measurement.

What this means for visibility

A better, cheaper frontier model means more queries going to AI systems, not fewer, and more trust placed in their answers. It also means more of those queries touch categories that were too expensive to run through a model before, which widens the set of questions a brand needs to worry about being answered for. For a brand, “which model wins” is the wrong question. The right one is whether your brand gets named in the answers these models give, whatever happens to be leading that week.

That’s what shifts with every new model generation, and it’s also what Truffle tracks: not one engine’s opinion of your brand, but whether you show up across ChatGPT, Gemini, Claude, and Perplexity at once. The models will keep trading places at the top. Whether your brand appears in the answer stays the same question, no matter which one is answering.

Frequently asked questions

What’s different about Claude Opus 5.5 compared to Opus 5? It costs 40% less on typical workloads, responds 30% faster, and improves on nearly every benchmark Anthropic has published (22 September 2026), from agentic coding to autonomous computer use.

How much cheaper is Opus 5.5? According to Anthropic, input costs $4 per million tokens (down from $5) and output $20 (down from $25). Cache reads drop from $0.50 to $0.20.

Is Opus 5.5 better than GPT-6 Astra? On most of the benchmarks Anthropic has shown, yes. Astra leads on two: AutomationBench, at 41.4% against 40.0% (a run Anthropic notes Opus 5.5 made without fallback models), and Terminal-Bench-Science 0.1, at 64.6% against 58.7%.

What does a cheaper frontier model mean for my brand’s visibility? More queries and more trust placed in what AI systems answer. The question that matters stops being which model wins and becomes whether your brand shows up in the answers they give, whichever one is leading.

Can I use Opus 5.5 today? Yes, it’s available across Anthropic’s main platforms as of 22 September 2026.

Spanish version of this article: zds.es/insights/claude-opus-5-5.

Source: Anthropic, official Claude Opus 5.5 announcement, 22 September 2026 (https://www.anthropic.com/claude-opus-5-5).

Book a demo

Book a demo →