Dev

Claude Haiku 5.5 as a Claude Code subagent: effort, the 100K tier and the cost per task

How to route a Claude Code subagent to Haiku 5.5 on purpose, which effort level to pin it to, how to keep it under the long-prompt rate with /autocompact, and what the system card says to check before you trust it with a tool loop.

Earlier today I compared the price cards of Claude Haiku 5.5 and GPT-6 Luna and found them identical up to 100,000 tokens: $0.10 per million input, $0.50 output. Since then the first real workloads have come in, and they are about the bill, not the card. Artificial Analysis counts about 162,000 output tokens per task for Haiku 5.5 at max effort, roughly three times Luna's. On Vals, a run of its index costs $2.99 per test against Luna's $0.43, both at max.

Same price per token, then, and not the same price per task. For the use Anthropic is promoting, a cheap subagent under Opus 5.5 or Sonnet 5.5, three settings decide which side of that gap you land on: the effort level, how fast the subagent's context grows, and where it compacts. None of them is on the price card, and the second one behaves differently from Haiku 4.5 by default.

Which Claude Code calls actually run on Claude Haiku 5.5?

Fewer than the "Opus plans, Haiku executes" demos suggest. As the earlier post noted, haiku means Haiku 5.5 only on the Anthropic API. On Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, the model configuration docs still map it to Haiku 4.5.

The built-in Explore agent no longer defaults to Haiku. Since v2.1.198 the changelog says Explore "now inherits the main session's model (capped at opus) instead of running on haiku," so under an Opus session every codebase search is an Opus call. The built-in claude-code-guide agent and some background jobs, such as auto-titles, still use the haiku alias. Everything else runs on Haiku only where you put it:

  • a custom subagent with model: haiku in its frontmatter;
  • CLAUDE_CODE_SUBAGENT_MODEL=haiku as the default for subagents, or with CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1 to override what each agent asks for (v2.1.257 and later);
  • a project or user agent named Explore with model: haiku, which the subagent docs suggest for running exploration on a cheaper model.

To check, /tasks names the model on each subagent's row, and adds the effort level when the definition sets one. Subagent transcripts are kept under ~/.claude/projects/<project>/<session-id>/subagents/, and each assistant line records its model. Two more details from the docs: a session saved on Haiku 4.5 and resumed with haiku comes back on 5.5, and in Claude Code thinking can't be turned off on Haiku 5.5. A saved alwaysThinkingEnabled: false or MAX_THINKING_TOKENS=0 "has no effect there."

How much does the effort level change Claude Haiku 5.5?

A lot. Haiku 5.5 defaults to medium on the API and in Claude Code, and every number in Anthropic's launch table is at max. The system card says so: "all Haiku 5.5 results use the following standard configuration: adaptive thinking at max effort." It also gives medium figures for two of them. GDPval-AA drops from 1620 to 1277, "while using about a tenth of the output tokens." AA-Briefcase drops from 1578 to 1372.

Artificial Analysis ran every level against Luna:

Artificial Analysis, Oct 7Haiku 5.5 maxHaiku 5.5 highHaiku 5.5 medium (default)Luna maxLuna medium
Intelligence Index4338343830
Time to first answer token323 s26 s12.6 s109 snot listed

At matched scores the token gap mostly closes. Artificial Analysis says Haiku 5.5 at high "scores 38 with ~55k tokens per task against 38 with ~50k for Luna (max)," and that moving from xhigh to max "adds 2 points for ~1.8x the tokens." Its time-to-first-token figure includes thinking, which is why the model Anthropic calls its fastest takes about five minutes to start answering at max.

Haiku 5.5 is also more effort-sensitive than its larger siblings on the tasks Anthropic reports. On PhysicianBench it scores 17.8% at low, 25.2% at medium and 43.0% at max, and time per task rises "from 48 seconds at low … to about 11 minutes at max." On HealthBench Professional, where it moves from 57.9% at low to 64.8% at max, the card notes that Sonnet, Opus and Fable "each varied by less than 2 points across the five levels." On this model, the effort setting decides how much capability you are paying for.

The prompting guide lists what goes wrong at the cheap end. At low, "the model is more likely to skip a search, stop early, or skip a check." At low and medium it "sometimes reports a code change as done without running a check," and Anthropic supplies a paragraph of instructions to counter it. Telling it to answer directly "didn't stop it from thinking." And changing the top-level effort value between requests invalidates the cache for the conversation's messages, one of the quieter ways to break a prompt cache. A per-message effort beta avoids that.

In Claude Code, pin it per subagent with effort: in the frontmatter. It overrides the session level but not CLAUDE_CODE_EFFORT_LEVEL, which beats everything except a maxEffortLevel cap. Without it, the subagent runs at whatever level the session uses.

Why does a Claude Haiku 5.5 subagent cross 100K tokens so fast?

Mostly because it now thinks at every step. The migration guide puts it plainly: "Adaptive thinking is on by default," and it advises choosing a lower effort level "where Claude Haiku 4.5 ran without thinking." Inside a tool loop, the thinking blocks travel back with each tool result, so a subagent that reasons before every call adds that reasoning to its context at every step. In Claude Code you can't turn it off, only turn it down.

Across turns, Haiku 5.5 also keeps what Haiku 4.5 dropped. The extended thinking docs put Haiku 5.5 in the "Keep all prior turns" group and Haiku 4.5 in "Keep the last turn only," and say "retained thinking blocks count as input like any other conversation history." That matters for a subagent that gets several instructions in one session, or a conversation you resume. On most models it is a fair trade for cache hits. On the one model priced by prompt length, it adds to the total. On the API, the clear_thinking_20251015 context-editing strategy can change it, but only with an explicit keep setting, since its default follows the model.

The tokenizer adds the rest. Anthropic puts the change from Haiku 4.5 at about 30% more tokens for the same text, and Simon Willison measured about 1.25x on one long prompt. A Hacker News commenter put the ratio against Luna at 1.5x. That is a claim, not a published measurement.

Then the harness. A Reddit user counted a fresh Claude Code main session on Haiku at about 66,000 tokens before any work, most of it tool and MCP definitions. A subagent gets its own, shorter base instructions, but the tool definitions it loads come along. Another user tracked one Haiku subagent under Opus: it crossed 100K around its twentieth call, and 145 of its 164 calls ran above the line. These are reports, and they fit the mechanism. Anthropic says requests under 100K were "around 90%" of its previous Haiku's traffic. Some commenters on Hacker News argue that is because few people ran Haiku 4.5 as an agent at all.

The earlier post left open whether cached tokens count toward the 100K. The pricing page has a separate cache-read rate for "prompts over 100,000 tokens," $0.05 per million against $0.01, and the reports above read as if cached context counts. The docs still don't define prompt length. And by default Claude Code compacts a Haiku 5.5 session only at about 967K tokens, so without a change most of a long session runs at the upper rate.

How do you keep a Claude Haiku 5.5 subagent under the 100K tier?

Anthropic's Lydia Hallie posted the fix on launch day:

text
If you're on API billing, you can set Haiku 5.5's autocompact window to 100K
so you stay in the cheaper token pricing tier! It's saved per model so this
only applies to Haiku (incl. subagents)
 
> /model haiku
> /autocompact 100k

/model haiku also switches your main session, so switch back to your usual model afterwards. The window is saved under Haiku in modelSettings, so Opus and Sonnet keep their own. CLAUDE_CODE_AUTO_COMPACT_WINDOW, if set, overrides it for every model, which is worth checking in CI.

The command accepts windows "from 100K to 1M tokens," so the lowest setting sits right at the price line. Whether compaction fires at the window or a little below it isn't documented. For more headroom, CLAUDE_AUTOCOMPACT_PCT_OVERRIDE triggers compaction at a lower percentage of the window and, per the docs, applies "to both main conversations and subagents." A small window also leaves little room once tool definitions take their share. Claude Code stops with an error if context refills right after three compactions in a row, so a subagent with a large tool list can hit that instead of finishing.

Beyond compaction:

  • Set effort by job. For classification, routing and extraction, low is what the price card was built for. One user reported $0.00005 per routing call at low, and, on a different job (summaries), 64% of output going to thinking at the default. For agent loops, the prompting guide starts at medium. If you find yourself at xhigh or max, the guide suggests running the same evals on Sonnet 5.5 and comparing cost.
  • Trim what the subagent loads. Loaded tool definitions are input on every call. Claude Code defers MCP tool schemas by default and loads them on demand, but a subagent definition with a short tool list still starts further from the line.
  • On the API, decide about thinking. Outside Claude Code, Haiku 5.5 accepts thinking: {"type": "disabled"} at low, medium and high. The prompting guide warns that with thinking off it "might skip a tool call it needs" when you also request structured JSON output.

What does the Claude Haiku 5.5 system card say about it as a subagent?

The system card has a few things worth knowing before you hand Haiku 5.5 a tool loop.

It has no fallback model. When its safety classifiers block a request, the API returns stop_reason: "refusal", and per the prompting guide, "Sending the same request to Claude Haiku 5.5 again usually returns another refusal." The same guide says these refusals are new for anyone moving from Haiku 4.5. The card adds that Haiku 5.5 "over-refused more than any other model we tested in our automated behavioral audit." Artificial Analysis saw the same thing independently: a pre-release refusal issue held its AutomationBench score down, and it plans a re-run. An orchestrator should treat a subagent refusal as a routing decision, not a reason to retry.

Two weaknesses are in Anthropic's own summary. One is a regression: "Haiku 5.5 used a leaked answer without telling the user 17% of the time, a regression from Claude Haiku 4.5, at 2%." The other isn't new: it "hallucinated more than other recent models, and about as much as Claude Haiku 4.5." Together with the prompting guide's note about reporting work as done without a check, both argue for the parent verifying anything Haiku says it finished.

Prompt-injection resistance moved the other way, sharply. On the Gray Swan benchmark at 15 attempts, the attack success rate fell from 83.2% on Haiku 4.5 to 7.1%. The weak spot is GUI computer use at 24.4%. This has one side effect for orchestrators: the prompting guide warns that a user message relayed inside a tool_result block may be treated "as untrusted text" and ignored. If your parent agent forwards user instructions through tool results, send them as user turns instead.

When is Claude Haiku 5.5 the right subagent?

So far the evidence splits by task shape. For short, single-turn work at volume (classification, routing, extraction, a summary that fits in one call), Haiku 5.5 does what the price card says, and the effort setting is the main thing to get right. For long tool loops the trade is quality against cost. On Bug Hunt Bench, Haiku 5.5 at max fixed 21.5 of 105 planted bugs on average, at an estimated $5.83, against 18.3 at $0.52 for Luna and 36 at $35 for Opus 5.5 at xhigh. Those are list-price estimates from different harnesses, and Luna's figure leaves out its own long-context surcharge, so read the cost gap as an order of magnitude, not a measured ratio.

Before you move a subagent to Haiku 5.5, pin its effort level, set /autocompact 100k, and price one real task end to end. If the bill matches the card, good. If it doesn't, the transcript will show which of the three settings caused it.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles