Claude Sonnet 5.5 is Anthropic’s mid-sized model, released on September 28, 2026 at $2 per million input tokens and $10 per million output tokens. I used it to build Spirebound, my free idle wizard-tower game, and to wire that game into my browser arcade. This post covers what Anthropic says the model does, then what I can and cannot say about how it behaved on a real project.
What is Claude Sonnet 5.5?
It is the follow-up to Sonnet 5 and sits below Opus 5.5 in Anthropic’s lineup. The model ID is claude-sonnet-5-5. In its launch announcement, Anthropic reports 70.6% on Terminal-Bench 4.0, output about 30% faster than Sonnet 5, and up to 30% lower cost per task. Those are Anthropic’s own figures, read on October 4, 2026, and I have not re-run any of them. If you want the family context, I wrote up how Opus 5.5 compares with Opus 5, Fable 5.1 and Sonnet 5 on launch day, and what was confirmed and what was rumor about Haiku 5.5 a few days later. In that second post I quoted Anthropic saying Sonnet 5.5 would follow in the coming weeks. It arrived three days after I published it.
How much does Claude Sonnet 5.5 cost?
The API price is $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 per million. I do not pay that rate myself. I use Claude through a subscription, so I cannot give you a dollar cost for building Spirebound, and I did not keep a per-task token count. If you need a real cost figure for your own work, measure it on your own tasks rather than trusting my build.
What did I build with Sonnet 5.5?
Spirebound merges eight of my earlier game prototypes into one idle game that runs from a single HTML file. As of October 3, 2026 it is live at spirebound.dhseadev.online. I counted what is in it from the game itself: 6 rooms, 22 research nodes, 6 familiars, 5 zones, 5 gods, 57 achievements and 7 colour themes. The Spirebound project page has the full description, and the launch post explains the design choices, including why offline progress is capped at two hours.
Where does the arcade come in?
The DHSeaDev arcade is the collection of free browser games I host at play.dhseadev.online. As of today it lists ten games and one app, and Spirebound appears under its own heading, “Hosted on its own site”, rather than among them. That was deliberate. Spirebound’s optional cloud save keeps a sign-in token in browser storage, and storage is shared by every page on the same address, so I gave the game its own subdomain instead of putting a login on the shared arcade address. The arcade’s earlier games were built before Sonnet 5.5 existed, so I am not claiming it for those. How the arcade takes in a new game, with one folder and one row in a table that every shared check runs against, is described in How the arcade works.
How did I check the work?
A new model does not change my rule that code is not trusted until a test has been shown to fail. The Spirebound cloud code has 30 tests. To confirm they could fail, I broke the code 11 different ways and checked that each break was caught. I also had subagent reviewers read the work as a stand-in for a human panel, and they found 12 issues, which I fixed. What none of that proved is a signed-in save on one device loading on another, so that stays on the open list. The Firebase side of the project, including the database rules and what I still have not verified, is in What I learned putting Firebase sign-in behind an idle game.
Is Sonnet 5.5 the reason it worked?
I cannot say, and I would be wary of anyone who claims to. I did one build, with one model, inside a process of tests and review that I would have run on any model. There was no side-by-side run against Sonnet 5 or Opus 5.5, no timing, and no count of retries. What I can say is narrower: the project shipped inside the first week the model existed, the checks above passed, and the open items are written down. If you are choosing between models for your own game or extension, treat this as one dated example, not a benchmark. The games index lists everything else I have built if you want to judge the output yourself.
