Haiku 5.5 beats Sonnet 5: time to rewrite your agent files
The small model that overtook the mid-size one
6 min read
Claude Haiku 5.5 shipped yesterday, October 7, almost a year after 4.5. Opus went through five versions in that time; Haiku went through none. The 5.5 family is now nearly complete. Opus 5.5 costs less than Opus 5 and does more, Sonnet 5.5 kept Sonnet 5's price and jumped a tier, and Haiku 5.5 closes the ladder from below. Fable is the one missing. People are saying next week. Anthropic hasn't said anything.
Here's the part that matters for your repo. The model choices sitting in your subagents, skills and CLAUDE.md were made back when Haiku couldn't handle much. That assumption is gone, and the setup it produced is burning tokens on work a ten-cent model now does fine. I haven't reopened my own files yet. It came out yesterday. This is the list I'll use when I do.
A year without a Haiku, and now it beats Sonnet 5
Put the Haiku 5.5 and Sonnet 5.5 launch pages side by side. On every benchmark you can line up, the new small model beats the old mid-size one:
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5 | Sonnet 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 10.3% | 70.6% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1449 | 1844 |
| Humanity's Last Exam, tools | 57.4% | 18.7% | 54.9% | 64.5% |
| Chartography, no tools | 46.4% | 6.4% | 15.6% | 61.6% |
| FrontierCode 1.1 | 46.4% | n/a | 42.4% | 52.1% (Xhigh) |
And the pricing, per million tokens:
| Per 1M tokens | Haiku 5.5 (up to 100k) | Haiku 5.5 (above) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1 | $2 |
| Output | $0.50 | $2.50 | $5 | $10 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| Cache write | $0.125 | $0.625 | $1.25 | $2.50 |
Under 100k input tokens that's 90% cheaper than Haiku 4.5, and Anthropic says 90% of Haiku 4.5 requests fell in that range. The new tokenizer spends a few more tokens per task. The figure already includes that.
Your Haiku 4.5-era model choices have expired
The right moment is the next time you open .claude/agents to add a subagent.
Before you add it, search for model: in the ones already there. Every
sonnet in that folder was picked against a Haiku that scored zero on
Terminal-Bench.
What surprised me is something else. Claude Code's Explore subagent doesn't run on Haiku. It uses the conversation's model. If you work in Opus, Opus does every codebase search. To change that, add a project subagent with the same name. It replaces the built-in one:
---
name: Explore
description: Read-only codebase search. Use it to find files, symbols and usages.
model: haiku
effort: low
omitClaudeMd: true
tools: Read, Grep, Glob
---
Find what you're asked for and answer with paths and line numbers.
Don't modify anything.Skills take the same model field, plus context: fork to run in a separate
subagent. A skill that writes a changelog or reshapes a JSON file doesn't need
the model you're using to reason about architecture.
Effort is the new dial, subagents included
Haiku 5.5 is the first Haiku with effort levels, from low to max, and
Claude Code defaults to medium. The effort field works per agent and per
skill, and it overrides the session setting.
This one needs some care. Anthropic's prompting guide says that with low
effort and a long system prompt, Haiku 5.5 sometimes stops early and hands
the task back unfinished. Moving to medium roughly halves early stopping, but
more than doubles output tokens. My rule of thumb:
lowfor lookups, classification and code search, with a short prompt;mediumfor anything that edits files;- no
xhighormaxon Haiku. At that point, try Sonnet 5.5 instead.
A long CLAUDE.md gets billed in every subagent
This is the least obvious reason to reopen your files. CLAUDE.md loads into every custom subagent (the built-in Explore and Plan skip it). The Claude Code docs say to keep it under 200 lines, because longer files eat context and lower adherence.
With Haiku the cost isn't just tokens. A long system prompt is exactly the
condition under which the small model quits early and skips searches. A
400-line CLAUDE.md written for Opus is noise to a subagent that only needs to
find three files. So either trim it, moving domain rules into files that load
only when needed, or set omitClaudeMd: true on the agents that don't need it.
A prompt written for Opus isn't enough for Haiku
With Opus, a vague delegation prompt usually holds up, because it fills the gaps itself. Haiku 5.5 is far better than 4.5, but your delegation prompt has to stand on its own. The official guide is practical about it: when you see a specific failure, add an explicit instruction, and it hands you the text. This is the one I'd put in every subagent that touches code:
When you change code that can be run, built, or type-checked, run a real check
that exercises the change before reporting it done: the project's tests,
type-checker, or build, or the changed command itself.Two things to know before you move everything over:
- Haiku 5.5 adds refusals Haiku 4.5 didn't have, with categories like
cyber. A security review subagent moved to Haiku may stop where it used to keep going. I'd keep that one on Sonnet. - Thinking is on by default and counts toward
max_tokens. If you have scripts callingclaude-haiku-4-5with a tightmax_tokens, swapping only the model ID can cut responses off.
Where to start tomorrow morning
grep -r "model:" .claude/andgrep -r "claude-haiku-4-5"across the repo.- Override Explore with
model: haiku. - An explicit
efforton every agent and skill. - CLAUDE.md under 200 lines, or
omitClaudeMdwhere it isn't needed. - The verification line in subagents that edit code.
What comes out is the setup Cognition described at launch: Opus 5.5 as lead and Haiku 5.5 as sidekick hold a top-tier FrontierCode score of 66.2, at lower cost and latency. It's the idea from You are paying an LLM to do an if, one step up: the big model where reasoning happens, the small one where work repeats. I went through the Opus pricing in Claude Opus 5.5 beats Fable 5.1 and costs less.
A month ago I didn't want to set a budget for the small nodes yet. Now I can. It's a tenth of what it was.
Got a project in mind?
I have been building software for companies and startups since 2018. If your product needs a hand, write to me. Worst case, you walk away with a free opinion.
Let's talkRead next
You are paying an LLM to do an if
Code where there is one right answer, reasoning where there is not: the diagram nobody draws anymore
7 min read
Shipping fast is the easy part
Why the illusion of merging 10 PRs a day is destroying code quality
6 min read