Claude Haiku 5.5 API: $0.10/$0.50 Pricing and Setup
Claude Haiku 5.5 is $0.10/$0.50 per 1M tokens for prompts up to 100K, and $0.50/$2.50 above. Not on RouterPlex yet. Sign up and use live models.

Claude Haiku 5.5 is Anthropic's small model for high-volume work, released 7 October 2026. The Claude API id is claude-haiku-5-5. Claude Haiku 5.5 API pricing is $0.10 per 1M input tokens and $0.50 per 1M output tokens for prompts up to 100,000 tokens. Prompts over that line are $0.50 input and $2.50 output per 1M.
Haiku 5.5 is not on RouterPlex yet. Create a RouterPlex account and use the marketplace that is live today, starting with Claude Haiku 4.5 at $1 / $5 per 1M. The new id works here once the catalog lists it. One prepaid key, 0% markup on vendor list price, a $0 top-up fee, and a hard spend limit that stops at $0.

Sources: Anthropic's Claude Haiku 5.5 announcement (7 October 2026), the Haiku 5.5 overview, the migration guide, the pricing page, and the system card. Checked 7 October 2026. Benchmark figures and customer quotes below are Anthropic's. Rates change. Confirm the vendor page before you lock a budget. The images are Anthropic's launch card and the charts on that page.
Claude Haiku 5.5 API pricing #
USD per 1M tokens. The announcement's cache-write cell of $0.125 / $0.625 is the 5-minute rate. The overview and the pricing page also publish a 1-hour cache write.
| Line | Up to 100K prompt tokens | Over 100K prompt tokens | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| 5-minute cache write | $0.125 | $0.625 | $1.25 |
| 1-hour cache write | $0.20 | $1.00 | $2.00 |
The higher rate is the rate for a prompt over 100,000 tokens, the wording on both the announcement and the pricing page. A 40,000-token input and 2,000-token output, with no cache, is $0.004 + $0.001 = $0.005. The same token counts on Haiku 4.5 are $0.04 + $0.01 = $0.05. A 120,000-token input and 2,000-token output, priced entirely at the over-100K rate, is $0.060 + $0.005 = $0.065.
Those examples hold the token count fixed. Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later models. Anthropic says the same text counts as about 30% more tokens than on Haiku 4.5. Recount with model set to claude-haiku-5-5 before you treat $0.005 as the bill for a paragraph you already counted on 4.5.
Anthropic's Batch API is 50% off input and output on that platform. That discount is an Anthropic bill. This model is not listed here, so it is not a RouterPlex bill.
What the 75% claim means #
The announcement says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average. Footnote 2 is more precise: the price is 90% lower for requests up to 100,000 tokens and 50% lower over that line. On Haiku 4.5, about 90% of requests were in the smaller bucket. The average also accounts for the tokenizer change. The sticker and the blended cost are different numbers. Quote the one you mean.
Footnote 1 is the speed claim: Haiku 5.5 is the fastest Claude at each model's standard speed, and it runs less quickly than Opus models in Fast Mode.
What Haiku 5.5 is #
| Claude Haiku 5.5 | |
|---|---|
| Claude API id | claude-haiku-5-5 (no date suffix) |
| Bedrock id | anthropic.claude-haiku-5-5 |
| Google Cloud, Foundry, Claude Platform on AWS | claude-haiku-5-5 |
| Released | 7 October 2026 |
| Context / max output | 1M / 128K (Batch beta 300K with output-300k-2026-03-24) |
| Latency | Fastest at standard speed |
| Thinking | Adaptive, on by default. Default effort medium |
| Knowledge cutoff | June 2026 |
| Tokenizer | About 30% more tokens than Haiku 4.5 for the same text |
| On RouterPlex | When listed |
Anthropic positions it for summaries, compaction, classification, database queries, live support, browser use, and as a subagent next to Sonnet 5.5 and Opus 5.5. The announcement says the larger models stay the better choice for complex agentic coding. Haiku 5.5 is the model for narrow tasks that used to be too expensive to run often.
Swapping the id onto a Haiku 4.5 client returns errors. The migration guide lists the breaks: thinking: {"type": "enabled", "budget_tokens": N} is rejected, so use adaptive thinking and the effort control. If you send temperature, it must be 1. If you send top_p, it must be 0.99. Any top_k, any other sampling value, or sending both temperature and top_p, returns HTTP 400. An assistant prefill returns 400. Thinking blocks work in the account that produced them. Priority Tier on Haiku 4.5 is not supported on Haiku 5.5. There is no server-side fallback when a safeguard returns stop_reason: "refusal".
Vendor benchmarks, 7 October 2026 #
These scores are the table on Anthropic's announcement. RouterPlex has not re-run them. Sonnet 5.5's FrontierCode cell is labeled Xhigh. The other Haiku 5.5 cells are not labeled with an effort on that table. GPT-6 Luna has no HLE number on the page.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| HLE, no tools | 45.9% | 10.2% | not reported | 56.9% |
| HLE, with tools | 57.4% | 18.7% | not reported | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main | 46.4% | not reported | 42.4% | 52.1% Xhigh |
| Chartography, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
Haiku 5.5 is the first Haiku with an adjustable effort setting. The announcement charts plot accuracy against cost at Low, Med, High, Xhigh, and Max. The chart does not print a score on each point. The table above is the set of numbers Anthropic printed.

What Anthropic's customers said #
These are quotes Anthropic published on the launch page. They are not RouterPlex measurements.
Aaron Vinh, Staff Software Engineer at Asana, said an AI Teammates eval (bug triage, project setup, and portfolio search) showed over a 30% reduction in task-completion latency and up to 2.5x faster inference per agent turn versus the model they use today.
Ze'ev Klapow, Distinguished Software Engineer at HubSpot, said Haiku 5.5 scored 92.8% averaged over three runs on a simulated-portal CRM suite, the best score they had seen there, and that it was the fastest model they tested on a stale-record audit, with the highest hit rate and the lowest false positive rate.
Walden Yan, Co-Founder and CPO at Cognition, said Haiku 5.5 as the sidekick in Devin Fusion holds a FrontierCode score of 66.2 while cutting cost and latency, with Opus 5.5 as the lead. That 66.2 is a Fusion score, not the 46.4% Haiku-only FrontierCode cell in the table above.
Daniel Campos at AlphaSense reported 0.84 versus 0.76 against Haiku 4.5 on 400 Ask-in-Document queries. Yashodha Bhavnani at Box reported 11 points higher than Haiku 4.5 at about half the latency. Alex Wang at Rogo described Haiku 5.5 as the subagent that pulls a line from a 10-K while a larger model builds the deck.
Safeguards #
Anthropic says alignment evals improved versus Haiku 4.5, with fewer instances of misaligned behavior and a lower willingness to cooperate with misuse. Cyber safeguards are more restrictive than Haiku 4.5 and less restrictive than Sonnet 5.5: they allow a wider set of defensive tasks than Sonnet 5.5 and still block penetration testing. Biology safeguards match Sonnet 5, Sonnet 5.5, and Opus 5. Details are in the system card.
Sonnet 5.5 cache reads dropped to $0.10 #
The same announcement cuts Claude Sonnet 5.5 cache reads from $0.20 to $0.10 per 1M tokens. Anthropic says that is about 20% cheaper on most agentic work, because cache reads are a large share of those tokens. Input stays $2 and output stays $10. The 5-minute cache write stays $2.50 and the 1-hour write stays $4. That cut is Anthropic's price. Sonnet 5.5 is not on RouterPlex either.

The same update adds a monthly Claude Platform credit for Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. Those credits are Anthropic's. A RouterPlex balance is prepaid by you, and it does not include that credit. Claude's Python and TypeScript SDKs also gain a beta for computer use and browser use. Haiku 5.5 supports the browser toolset browser_toolset_20260801 on the Claude API and Google Cloud. Haiku 4.5 does not.
Call a live model today #
claude-haiku-4-5 is the live RouterPlex small Claude route at $1 / $5 per 1M. It is the previous Haiku, not 5.5.
curl https://api.routerplex.com/v1/chat/completions \-H "Authorization: Bearer $ROUTERPLEX_API_KEY" \-H "Content-Type: application/json" \-d '{"model": "claude-haiku-4-5","messages": [{"role": "user", "content": "Classify this ticket as billing, access, or other."}]}'
GPT-6 Luna is also live, at $0.10 / $0.50 per 1M for prompts at or below 272K. That sticker matches Haiku 5.5's under-100K rate. It is a different model, with a different long-context rule. The same key reaches the rest of the catalog through the OpenAI-compatible API. Compare gateways on the OpenRouter comparison.
On RouterPlex, once Haiku 5.5 is listed #
When the row exists, the contract is vendor list price, 0% markup, prepaid balance, and a hard per-key budget. The request stops at $0. This page will name the RouterPlex route that day. Until then, paste claude-haiku-4-5, not claude-haiku-5-5.
Create a RouterPlex account if you want the key ready, and use the models already on the marketplace while you wait.
Common questions
Frequently asked questions
How much does the Claude Haiku 5.5 API cost?
Anthropic lists Claude Haiku 5.5 at $0.10 per 1M input tokens and $0.50 per 1M output tokens for prompts up to 100,000 tokens. Prompts over 100,000 tokens are $0.50 input and $2.50 output per 1M. Cache reads are $0.01 or $0.05 per 1M on the same split. Those are Anthropic platform rates, checked 7 October 2026. RouterPlex does not bill Haiku 5.5 yet.
What is the Claude Haiku 5.5 model ID?
On the Claude API the ID is claude-haiku-5-5. It has no date suffix. Amazon Bedrock uses anthropic.claude-haiku-5-5. Google Cloud, Microsoft Foundry, and Claude Platform on AWS use claude-haiku-5-5. There is no RouterPlex model ID until the catalog lists it.
Is Claude Haiku 5.5 on RouterPlex?
Not yet. Create an account at routerplex.com/sign-up and use the live marketplace, including claude-haiku-4-5 at $1/$5 per 1M and gpt-6-luna at $0.10/$0.50 per 1M for prompts at or below 272K. Haiku 5.5 can be called here once it is listed.
What is the Claude Haiku 5.5 context window?
The model overview documents a 1M-token context window, up to 128K output tokens, and a June 2026 knowledge cutoff. The Message Batches API beta allows up to 300K output tokens with the output-300k-2026-03-24 header. Default effort is medium.
How does Haiku 5.5 compare with Haiku 4.5 on price?
The sticker for prompts up to 100,000 tokens is 90% lower: $0.10/$0.50 versus $1/$5 per 1M. Prompts over 100,000 tokens are 50% lower, at $0.50/$2.50. Anthropic says about 90% of Haiku 4.5 requests were in the smaller bucket, and that the same text counts as about 30% more tokens on Haiku 5.5. Their blended claim is about 75% less to run.
Can I paste a Haiku 4.5 request into claude-haiku-5-5?
The migration guide says no. Recount tokens, switch thinking to adaptive, and stop ending messages with an assistant prefill. If you send temperature it must be 1. If you send top_p it must be 0.99. Any top_k, any other sampling value, or both temperature and top_p, returns HTTP 400. Priority Tier on Haiku 4.5 does not carry over.
Start from a model that is listed.
This release is not on RouterPlex yet. Open a live catalog row, or browse every listed model and its price.


