Anthropic has launched Claude Haiku 5.5, and the numbers behind the release are aggressive. Token prices drop by up to 90 percent, while the model’s score on the OSWorld computer use benchmark jumps from 15.7 to 72.4 percent. The company positions it as its cheapest and fastest small model, available across AWS, Google Cloud, Microsoft Azure, and GitHub Copilot.
TL;DR: Anthropic has launched Claude Haiku 5.5, its fastest and cheapest small model, replacing Haiku 4.5 after roughly a year. Token prices drop by up to 90 percent (The Decoder), the OSWorld computer use score jumps from 15.7 to 72.4 percent, and the model is available on AWS, Google Cloud, Azure, and GitHub Copilot, with new effort controls and updated safety safeguards.
What Is Claude Haiku 5.5 and Who Is It For?
Claude Haiku 5.5 is Anthropic’s new small model, and it replaces Haiku 4.5 almost exactly a year after that model arrived. Techzine describes Haiku 4.5 as a compact, affordable, and fast alternative to Sonnet and Opus, and Haiku 5.5 continues that role with improved performance and lower costs. It is the latest entry in the Claude 5.5 family.
The target audience is teams running high-volume, latency-sensitive workloads. GitHub’s changelog describes the model as designed for fast, high-volume work like subagents, quick edits, and terminal tasks inside Copilot. That profile fits customer-facing chatbots, background automation, and any pipeline where thousands of small calls per hour would make a larger model uneconomical.
Anthropic also ships updated safety safeguards with the release, according to Firstpost coverage of the launch. So this is not only a cost story. The company is trying to keep the small-model tier current with both capability and safety work happening at the frontier.
Who should actually pay attention? Developers who skipped the Haiku tier because the benchmark gap felt too large now have a different decision to make. The scores below tell that part of the story.
How Much Cheaper Is Haiku 5.5 Than Its Predecessor?
The headline figure is a price reduction of up to 90 percent per million tokens. The Decoder reports that the cuts are huge across the board: cache reads and writes, input costs, and output costs all drop substantially compared with the previous generation. Intelligentliving puts the effective reduction at 75 percent cheaper than Haiku 4.5 for typical usage.
Why does the cache pricing matter so much? Applications that repeatedly send the same context — system prompts, documentation, codebases — lean heavily on cached tokens. Cutting cache read and write prices by similar magnitudes changes the economics for exactly the workloads Haiku targets.
There is one caveat. The Decoder notes that a new tokenizer accompanies the release, which means raw per-token comparisons need some care. Fewer tokens for the same text can offset nominal price changes, or the other way around depending on your content.
VentureBeat’s framing is blunt: a 90 percent API price reduction that matches GPT-6 Luna. Sources describe this as evidence that the AI pricing war is intensifying rather than cooling off. Anthropic took the wraps off the model on a Wednesday, and the tiered pricing structure went live across all major cloud platforms at launch, per a company statement carried by StreetInsider.
How Big Is the Benchmark Jump Over Haiku 4.5?
The single most striking number in the launch materials is OSWorld. Haiku 5.5 scores 72.4 percent on the OSWorld computer use benchmark, up from 15.7 percent for Haiku 4.5. That is a jump of more than 56 percentage points in a single generation.
OSWorld measures how well a model operates a real computer environment — reading screens, taking actions, completing multi-step tasks. A score under 16 percent means the previous Haiku largely failed at these tasks. Above 72 percent, the model becomes genuinely usable for PC automation workflows.
The improvements are not limited to computer use. XDA reports gains for Haiku 5.5 over Haiku 4.4/4.5 in knowledge work, PC tasks, visual reasoning, and coding. That breadth matters, because a small model that only wins on one benchmark is a niche tool; one that improves across four categories is a drop-in upgrade.
Firstpost summarizes the release the same way: faster responses, more affordable pricing, improved performance, adjustable effort settings, and updated safety safeguards. The pattern across sources is consistent. This is a generational replacement, not an incremental refresh.
How Does Haiku 5.5 Stack Up Against GPT-6 Luna?
On price, the models now match. VentureBeat’s headline says it directly: Haiku 5.5 arrives with a 90 percent API price reduction, matching GPT-6 Luna. For buyers comparing small models from the two biggest labs, cost is no longer the differentiator.
On capability, third-party testing suggests Haiku 5.5 holds its own. Intelligentliving reports that the model beats GPT-6 Luna on key tests. The Kingy AI comparison went further, running both models on 160 responses with strict scoring, client-side speed measurement, and zero added spend — and documented the scoring limitations that come with that methodology.
What does Anthropic itself claim about speed? The company calls Haiku its fastest model at standard speeds, while acknowledging that Opus in Fast Mode runs faster. Notably, VentureBeat reports that Anthropic does not supply a tokens-per-second figure in the materials reviewed, so buyers cannot yet compare raw throughput numbers on paper.
The competitive picture is straightforward. Anthropic priced its small model at parity with OpenAI’s small model and claims benchmark wins where it counts. Coverage from Yahoo Finance frames the whole release as a sign the AI pricing war is intensifying, with each lab forcing the other’s hand on cost.
What Are the New Effort Controls?
Haiku 5.5 introduces adjustable effort settings, giving developers a dial between speed and thoroughness on a per-request basis. Firstpost lists these effort controls alongside faster responses and affordable pricing as the core of the launch, and the Kingy AI spec comparison covers effort settings as a first-class feature alongside benchmarks, pricing, and safety evaluations.
The idea will be familiar to anyone who has tuned inference settings before: not every prompt deserves maximum reasoning. A quick edit or a routing decision can run at low effort, while a multi-step terminal task can justify the top setting. On a model priced for high-volume work, the effort dial directly shapes your monthly bill.
How does this interact with the tiered pricing structure Anthropic announced? The company statement describes tiered pricing across all major cloud platforms — AWS, Google Cloud, and Microsoft Azure — so effort levels and pricing tiers together give deployment teams more granular control over cost than the previous flat model.
GitHub’s Copilot integration shows the effort concept in practice. The changelog positions Haiku 5.5 for subagents, quick edits, and terminal tasks, with generally available availability inside Copilot. Those are precisely the workloads where low-effort, high-speed inference pays for itself.
Where Can You Run Haiku 5.5?
Claude Haiku 5.5 is available across all major cloud platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure, according to a company statement. AWS has published its own announcement confirming the model’s arrival on its platform, so teams already running Claude workloads on Amazon infrastructure can adopt the new model without changing providers.
That multi-cloud coverage matters for enterprise buyers. Why lock yourself into one vendor? A development team can prototype on one platform and deploy on another while keeping the same model and the same API surface.
The AWS blog post introducing Claude Haiku 5.5 was authored by engineers on the AWS Machine Learning blog, which typically signals documented setup guidance, supported regions, and integration paths for Bedrock customers. For organizations already paying for AWS commitments, running a cheap, fast small model like Haiku 5.5 in place of larger models for routine tasks is an obvious cost lever. Coverage indicates the model is positioned as a drop-in successor to Haiku 4.5, which Techzine notes was released almost exactly a year earlier as Anthropic’s compact, affordable, fast alternative to Sonnet and Opus.
How Does It Fit Into GitHub Copilot?
Claude Haiku 5.5 is now generally available in GitHub Copilot, according to the GitHub Changelog entry dated October 7, 2026. GitHub positions it as a model designed for fast, high-volume work: subagents, quick edits, and terminal tasks.
That role makes sense given the model’s profile. Copilot’s architecture increasingly delegates small, parallel jobs to lightweight models — spawning subagents to explore a codebase, renaming a variable across files, or running shell commands — while reserving bigger models for architectural reasoning. A cheap, fast model is exactly what those sub-tasks need.
The changelog also notes that the model showed strength in early use, though the available snippet cuts off before detailing the specific early results. What is confirmed is the timing: the general availability in Copilot followed closely on the model’s launch, part of the same wave of announcements that brought Haiku 5.5 to AWS, Google Cloud, and Azure.
For developers, the practical takeaway is simple. If your Copilot plan includes model selection, Haiku 5.5 is now an option for the highest-frequency interactions — the ones where latency and cost accumulate fastest.
How Fast Is Haiku 5.5 in Practice?
Anthropic calls Haiku its fastest model at standard speeds, while acknowledging that Opus in Fast Mode runs faster. Notably, Anthropic does not supply a tokens-per-second figure in the launch materials reviewed for this article, so buyers comparing raw throughput numbers will have to rely on independent testing rather than official specs.
What the company does provide is positioning: Haiku 5.5 is the small, quick member of the Claude 5.5 family, launched with what coverage describes as faster responses, more affordable pricing, and adjustable effort settings. Those effort controls let developers trade response depth against speed and cost on a per-request basis — useful when the same application serves both trivial lookups and harder reasoning tasks.
The benchmark picture supports the speed story. On OSWorld, the computer-use test, Haiku 5.5 jumped from its predecessor’s 15.7 percent to 72.4 percent — a nearly fivefold improvement. Faster and dramatically more capable at the same time is an unusual combination for a budget-tier model.
If you need hard latency numbers, watch for independent measurements rather than launch-day claims. The absence of an official tokens-per-second figure is itself informative about how Anthropic is framing this release.
What Changed on the Safety Side?
Anthropic launched Haiku 5.5 with updated safety safeguards alongside its performance and pricing improvements, according to coverage of the release. The company paired the new effort controls — the adjustable settings that govern how much compute the model applies to a given request — with these updated protections.
This matters more than it might seem for a small model. Haiku-class models are precisely the ones deployed at high volume in agentic setups: running terminal commands, editing files, browsing, and executing multi-step computer-use workflows. The jump to 72.4 percent on OSWorld means the model is now far more capable of actually acting in a computing environment, which raises the stakes on guardrails.
Anthropic’s announcement materials include safety evaluations alongside the benchmark tables and pricing details, according to third-party comparisons of the release. The pattern mirrors what the company has done with previous family members: ship capability improvements and safety work together rather than treating the small, cheap tier as an afterthought.
For teams building agents, the practical advice is unchanged: pair the model’s safeguards with your own sandboxing, permissioning, and human review on consequential actions. A better safety eval score on a launch slide is not a substitute for application-level controls.
What Does This Mean for the AI Pricing War?
Haiku 5.5 arrives with token prices dropping by up to 90 percent, and independent coverage describes it as costing 75 percent less than Haiku 4.5. VentureBeat frames the cut as matching GPT-6 Luna on price, and the launch has been read as evidence that the AI pricing arms race is far from over.
The cuts are not uniform across the board. Sources report huge reductions for cache reads and cache writes, in addition to lower input and output costs — meaning applications that aggressively reuse context through prompt caching benefit disproportionately. A new tokenizer accompanies the release, which coverage flags as a caveat when comparing costs against previous models.
Three consequences follow. First, high-volume workloads — classification, extraction, subagent orchestration — get dramatically cheaper overnight. Second, the competitive floor keeps falling: when a small model that scores 72.4 percent on OSWorld costs a fraction of what its predecessor did a year ago, every vendor’s budget tier comes under pressure. Third, pricing is becoming tiered and effort-based rather than a single per-token number, which complicates simple cost comparisons.
Cheap inference is no longer a differentiator. It is table stakes.
Frequently Asked Questions
How much cheaper is Claude Haiku 5.5 than Haiku 4.5?
Token prices drop by up to 90 percent compared with the previous generation, and one analysis puts the overall cost reduction at 75 percent versus Haiku 4.5. The largest cuts apply to cache reads and cache writes, alongside lower input and output prices.
Did Haiku 5.5 improve on computer use benchmarks?
Yes. On the OSWorld computer use test, Haiku 5.5 jumped from 15.7 percent — its predecessor’s score — to 72.4 percent, a gain of nearly 57 percentage points. That makes the small model dramatically more capable at PC-style agentic tasks.
Can I use Claude Haiku 5.5 on AWS, Google Cloud, and Azure?
Yes. According to a company statement, Claude Haiku 5.5 is available across all major cloud platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. AWS has also published a dedicated blog post introducing the model on its platform.
Is Haiku 5.5 available in GitHub Copilot?
Yes. Claude Haiku 5.5 is generally available in GitHub Copilot as of the October 7, 2026 changelog entry. GitHub describes it as designed for fast, high-volume work such as subagents, quick edits, and terminal tasks.
Summary
Claude Haiku 5.5 is a price-and-performance reset for the small-model tier. Key takeaways:
- Prices drop by up to 90 percent, with the steepest cuts on cache reads and writes; independent analysis puts the total cost reduction versus Haiku 4.5 at 75 percent.
- Computer-use capability improved from 15.7 to 72.4 percent on OSWorld — a nearly fivefold jump.
- The model is available on AWS, Google Cloud, and Azure, and is generally available in GitHub Copilot for subagents, quick edits, and terminal tasks.
- Anthropic calls Haiku its fastest model at standard speeds, though it published no tokens-per-second figure; effort settings and updated safety safeguards ship alongside.
- The release matches GPT-6 Luna on price, which coverage reads as confirmation that the AI pricing war continues.
If you run high-volume Claude workloads, revisit your model routing today — the cached-context economics alone justify the exercise.