Claude Haiku 5.5 and GPT-6 Luna cost the same until the prompt gets long
Identical price cards, two different cliffs. On long prompts the models part ways: one reprices the whole request above 100,000 tokens, the other above 272,000. For a subagent that reads files, that line decides the bill more than the rate does.
Anthropic released Claude Haiku 5.5 on October 7, and its short-prompt token rates match OpenAI's GPT-6 Luna to the cent: $0.10 per million input tokens, $0.50 output, $0.01 for a cache read, $0.125 for a cache write. The same day, Claude Code v2.1.293 moved its haiku alias to the new model on the Anthropic API.
Same card, then. What differs is where each one stops applying, and for a subagent that reads files, that line ends up being the price.
Where does Claude Haiku 5.5's price change?
Above 100,000 prompt tokens, and the change applies to the whole request. Anthropic's price table lists two rows for Haiku 5.5, "for prompts up to 100,000 tokens" and "for prompts over 100,000 tokens," and every column switches: input goes to $0.50, output to $2.50, cache reads to $0.05, cache writes to $0.625. That's five times the short rate, output included.
GPT-6 Luna has a line too, above 272,000 input tokens. Past it, OpenAI charges 2x on input and cache and 1.5x on output, "for the full request": $0.20 in, $0.75 out. Even the matching cache write isn't quite the same product: Anthropic's $0.125 buys five minutes, while OpenAI keeps a cached prefix for at least 30 minutes on GPT-5.6 and later.
What do Claude Haiku 5.5 and GPT-6 Luna cost per request?
Same up to 100k, five times apart between 100k and 272k, roughly two and a half times apart above that. Uncached prompt, 2,000 output tokens (thinking included):
| Prompt | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| 90,000 tokens | $0.010 | $0.010 |
| 110,000 tokens | $0.060 | $0.012 |
| 300,000 tokens | $0.155 | $0.062 |
The edge is the part I'd watch. A 99,000-token prompt costs about a cent of input; at 101,000 it costs five. File contents, logs and tool definitions can push a subagent's prompt across that line. Big tool catalogs already spend context before the agent reads the request; on Haiku 5.5 they can also pick your price tier.
What does Anthropic's "75% cheaper" for Claude Haiku 5.5 assume?
A traffic mix that looks like Haiku 4.5's. The footnote says Haiku 5.5 is 90% cheaper than Haiku 4.5 up to 100k and 50% cheaper above, that about 90% of Haiku 4.5 requests were under 100k, and that the average accounts for the new tokenizer.
My read is that the mix is the weak point. Haiku 4.5 had a 200k context window, so its "over 100k" bucket could only hold prompts between 100k and 200k. Haiku 5.5 has 1M. The average was computed on traffic that couldn't contain the long-context jobs the new window invites. Fair for short classification and routing; for agent work that reads a repository, compute your own.
Read the line the other way too. The 100k is counted in the new tokenizer, which produces about 30% more tokens for the same text, so it sits near 77k of Haiku 4.5 tokens. A prompt that ran at 80k on Haiku 4.5 crosses on 5.5 with nothing changed on your side. The footnote doesn't say which tokenizer its "90% under 100k" was counted in.
Holding the output at 2,000 tokens, a 110k prompt still costs less than the same text on Haiku 4.5 (about 85k tokens there): $0.060 against roughly $0.095. Hidden thinking can change that per request.
What don't the Claude Haiku 5.5 and GPT-6 Luna docs settle?
Two things. Above the line, cached tokens are billed at the higher rate; that row is published. Whether they count toward crossing the line isn't stated for Haiku 5.5. Anthropic's caching docs define total input as uncached plus cache reads plus cache writes, which points to yes. If so, an agent with a large cached prefix of instructions and tools pays the high tier on every call. Check usage on a known prompt before you build a router on it.
And "same price" is per token, not per task. Anthropic documents the 30% against Haiku 4.5; OpenAI's model page doesn't say how Luna's tokenizer compares. Count your real prompts on both before trusting the matching card:
n = client.messages.count_tokens(
model="claude-haiku-5-5",
system=system_prompt,
tools=tools,
messages=messages,
).input_tokensThat works for requests with client tools; most server tools and the MCP connector aren't supported, and the count is an estimate, so near 100k trust the billed usage instead.
Which Claude Haiku 4.5 code breaks on Claude Haiku 5.5?
The loud failures return 400: setting temperature, top_p or top_k (omit all three; a classifier pinned to temperature 0 fails on its first call), an assistant prefill, budget_tokens, computer_20250124 on the Claude API and Google Cloud, and sending thinking back after editing earlier turns. Those announce themselves, as they did on Opus 5.5.
The quiet ones come back as 200. Adaptive thinking is on by default at medium effort, the thinking counts toward max_tokens, and a small limit tuned for Haiku 4.5 can stop after a thinking block with no text at all. On the Anthropic API, a Claude Code subagent that follows the default haiku alias changes models without anyone editing its definition. Thinking blocks work only in the producing account or a linked one, the same rule that drops Sonnet 5.5's reasoning on a switch. Priority Tier isn't supported on Haiku 5.5. Refusals get no server-side fallback either.
Which is the cheaper subagent, Claude Haiku 5.5 or GPT-6 Luna?
Up to 100k, neither: they cost the same per token, and on Anthropic's own table Haiku 5.5 leads Luna on every benchmark both appear in, so there's no price reason to leave it there. Between 100k and 272k, Luna is five times cheaper per token.
So pick by prompt length first and vendor second, and measure it in tokens, not requests. Ninety calls at 50k and ten at 150k: the ten are a quarter of the input tokens and over 60% of the input bill. If that share is small, a second vendor buys you a second bill and a second cache to keep warm. If it's large, trim the context or move those jobs to Luna. Sonnet 5.5 has no long-context tier, and with its cache reads cut to $0.10 the same day, a long cached prefix there reads at twice Haiku's over-the-line rate, which is a fair price if the job needed the stronger model anyway. Count your tokens on both tokenizers and find the share past 100k. That number picks the vendor; the matching card doesn't.