Claude Haiku 5.5 vs Sonnet 5.5 vs Haiku 4.5: price, benchmarks and a debate

DHSeaDev — Chrome Extensions, Windows Tools, & Idle Games

Claude Haiku 5.5 is Anthropic’s newest small model, released on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. This post lines it up against Claude Sonnet 5.5 and the previous Haiku, Haiku 4.5, using the numbers Anthropic published today. I have not run Haiku 5.5 on any of my own projects yet, so this is a read of the published figures, not a test result, and I end with the questions I would like you to argue about in the comments.

What is Claude Haiku 5.5?

It is the small, fast tier of the 5.5 family, below Sonnet 5.5. The API ID is claude-haiku-5-5, and per Anthropic’s announcement it is available on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic’s models overview lists the same 1M-token context window, 128K maximum output, adaptive thinking and June 2026 knowledge cutoff for Haiku 5.5 and Sonnet 5.5. The differences it lists are the default effort (medium for Haiku 5.5, high for Sonnet 5.5) and the latency rating (fastest versus fast). Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, with levels from low up to max. Before launch I wrote about what was confirmed and what was rumor; today’s release replaces that guesswork with Anthropic’s own page.

How does Haiku 5.5 pricing compare with Sonnet 5.5 and Haiku 4.5?

Per million tokensHaiku 5.5 (up to 100K prompt)Haiku 5.5 (over 100K prompt)Haiku 4.5Sonnet 5.5
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache read$0.01$0.05$0.10$0.10
Cache write$0.125$0.625$1.25$2.50
List prices as published by Anthropic on October 7, 2026.

Anthropic says Haiku 5.5 is 90% cheaper than Haiku 4.5 for prompts up to 100K tokens and 50% cheaper above that, which it estimates at about 75% lower cost on average because roughly 90% of Haiku 4.5 requests fall in the shorter tier. It also says that estimate allows for a new tokenizer that uses slightly more tokens. Against Sonnet 5.5, the table works out to 20 times cheaper on input and output for short prompts, and 4 times cheaper above 100K. Sonnet 5.5’s cache read price also dropped from $0.20 to $0.10 on the same day, which is why this table disagrees with the $0.20 figure in my Sonnet post. That post was accurate when I wrote it; the price has since changed.

How do the benchmarks compare?

Benchmark (vendor-reported)Haiku 5.5Haiku 4.5Sonnet 5.5
GDPval-AA v2.116207351840
AA-Briefcase v1.115786141824
OSWorld 2.1 (offline subset)72.4%15.7%83.9%
Humanity’s Last Exam, no tools45.9%10.2%56.9%
Humanity’s Last Exam, with tools57.4%18.7%64.5%
Terminal-Bench 4.039.2%0.0%70.6%
FrontierCode 1.1 (Main)46.4%not published52.1%
Chartography, no tools46.4%6.4%61.6%
Source: Anthropic’s published table, October 7, 2026. Anthropic’s table also lists GPT-6 Luna, which I left out to keep this to the three models in the title.

Two readings sit side by side. Against Haiku 4.5, Haiku 5.5 is far ahead on every row where Haiku 4.5 has a published score, including Terminal-Bench 4.0, where Haiku 4.5 scored 0.0%. Against Sonnet 5.5, Haiku 5.5 is behind on every row Anthropic published, from a gap of 5.7 points on FrontierCode to 31.4 points on Terminal-Bench. Anthropic’s own guidance is that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding, and that Haiku 5.5 suits high-volume, cost-sensitive work such as summarization, classification and database queries, as a subagent, and for speed-sensitive jobs like live support and browser use.

Two details change how I read the coding number. VentureBeat reports that the 39.2% Terminal-Bench score is at maximum effort and that Haiku 5.5 scores about 20% at medium, which is the default. I could not find that medium figure on Anthropic’s page, so treat it as one outlet’s report. And the Sonnet 5.5 FrontierCode figure is labelled Xhigh effort on Anthropic’s page but High effort in VentureBeat’s write-up, so I have not tried to say which is right.

How much should you trust these numbers?

  • Every benchmark and the 75% cost claim above are Anthropic’s own figures. I have not re-run any of them.
  • The customer results Anthropic cites (Asana, HubSpot, AlphaSense, Box and others) are relayed by Anthropic and I have not checked them with those companies.
  • Anthropic’s page does not state an ASL safety level, and I did not read the full system card, so I make no claim about it here.
  • The models overview page lists almost nothing about Haiku 4.5, so the old-versus-new comparison rests on the pricing and benchmark tables only.

Which model would you pick for a job? Tell me in the comments

I do not have a verdict, and I would rather hear from people who have run it. The published numbers set up a few real disagreements, and I would like to see where you land on each:

  • Is 20 times cheaper worth being behind on every benchmark? For a job that runs thousands of times a day, does the Sonnet 5.5 gap matter, or does the price decide it?
  • Does the default effort setting decide it? If medium effort really lands near half the max-effort Terminal-Bench score, is Haiku 5.5 still a coding tool, or only a subagent?
  • Is Haiku 4.5 worth keeping anywhere? I can see no row where Haiku 4.5 is ahead, but your workload may differ from the benchmarks.
  • Do vendor benchmarks settle anything? If you have your own test results on Haiku 5.5, Sonnet 5.5 or Haiku 4.5, say what you ran and what you measured, so others can judge them.

If you share numbers, include the task, the effort setting and how many runs you did. A single run of anything is an anecdote, mine included, and when I do test Haiku 5.5 on my own projects I will write it up as a separate post with what I measured and what I did not.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *