Claude Haiku 5.5 makes a strong case for using a small model on work that repeats thousands of times: extracting facts, classifying requests, preparing summaries, and handling short agent tasks. Anthropic launched it on October 7, with short-prompt input and output rates matching GPT-6 Luna. Its launch results put it ahead of Luna on several evaluations.
The comparison with Sol requires more care. GPT-6 Sol and the newer GPT-6.1 Sol occupy a more expensive tier, and their coding results can justify that expense. Haiku also trails Sonnet 5.5 on the demanding Terminal-Bench launch evaluation. For a production application, the useful question is how much of your workload Haiku can complete correctly before a larger model or a person needs to take over.
Two details deserve attention before the score tables. Haiku’s cheapest rates apply to prompts of at most 100,000 tokens, despite its million-token context window. And the headline 72.4% computer-use score measures partial credit: the same evaluation reports a 37.1% rate for completing every checkpoint. Both affect how much useful work you can expect for your money.
Claude Haiku 5.5 specifications
The model specification lists text and image input, text output, adaptive thinking, a one-million-token context window, and 128,000 output tokens. Those last two limits are substantially larger than Haiku 4.5’s 200,000-token context and 64,000-token output limit.
| Specification | Claude Haiku 5.5 |
|---|---|
| Claude API model ID | claude-haiku-5-5 |
| Context window | 1 million tokens |
| Standard maximum output | 128,000 tokens |
| Batch output extension | Up to 300,000 tokens in beta with output-300k-2026-03-24 |
| Input / output modalities | Text and images / text |
| Knowledge and training-data cutoffs | June 2026 |
| Thinking / default effort | Adaptive / medium |
| Release / earliest retirement commitment | October 7, 2026 / not sooner than October 7, 2027 |
| Platforms listed | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
A context limit describes how many tokens an API accepts. It does not promise that the model will find every relevant fact, maintain every instruction, or reason equally well across the whole window. A long transcript with repeated boilerplate presents a different problem from a dense technical document with many conflicting details. Evaluate retrieval and reasoning separately.
The published specifications do not disclose Haiku’s parameter count, active parameters, layer count, or training compute. Calling it a “small” model describes its product tier; it does not establish a particular architecture or hardware requirement. There is no basis here for estimating the GPU memory needed to run Haiku yourself.
Text output also matters. Image understanding can support chart questions and screenshot-driven agents, but it does not make the model a native image, audio, or video generator. Compare the actual input and output contract when considering alternatives.
Pricing: the 100,000-token boundary matters
Haiku 5.5 has two prompt-length bands. The Anthropic pricing documentation gives the following rates. A prompt above 100,000 tokens moves to the higher input, output, and caching rates.
| Token category | Prompt ≤100,000 tokens | Prompt >100,000 tokens |
|---|---|---|
| Uncached input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache read | $0.01 | $0.05 |
| 5-minute cache write | $0.125 | $0.625 |
| 1-hour cache write | $0.20 | $1.00 |
The higher band is five times the lower band for each category. That makes prompt growth a budget issue. A long-lived agent can start cheaply and later enter a different price band as messages, tool results, screenshots, and retrieved documents accumulate.
Anthropic’s launch explanation distinguishes a 90% token-price reduction versus Haiku 4.5 for shorter prompts, a 50% reduction for longer prompts, and an estimated average saving of about 75% after workload and tokenization changes. These describe different comparisons. Your application’s saving will depend on its prompt distribution and generated tokens.
Compare the token rates with Sol and Luna
For the OpenAI entries below, “base” means Standard processing with no more than 272,000 input tokens. OpenAI’s price list applies higher full-request rates above that boundary. Haiku’s cheaper band ends much earlier.
| Model | Input | Output | Cache read | Band shown |
|---|---|---|---|---|
| Haiku 5.5 | $0.10 | $0.50 | $0.01 | Prompt ≤100K |
| Haiku 5.5, longer prompt | $0.50 | $2.50 | $0.05 | Prompt >100K |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | Input ≤272K |
| GPT-6 Sol | $2.00 | $10.00 | $0.20 | Input ≤272K |
| GPT-6.1 Sol | $2.00 | $10.00 | $0.10 | Input ≤272K |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.10 | Current standard rates |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | Current standard rates |
| GPT-6 Astra | $10.00 | $50.00 | $1.00 | Input ≤272K |
Claude rates: price documentation. Sonnet’s $0.10 cache-read rate is confirmed in the October 7 price-cut announcement; some overview tables still displayed the earlier $0.20 rate when checked. OpenAI rates: Standard price table. These exclude tool charges, infrastructure, taxes, and regional premiums.
Three worked examples
These are arithmetic examples using fixed token counts, not measured task costs. Assume uncached input, Standard processing, and no separate tools. Real models tokenize the same text differently and may generate different amounts of reasoning.
| Request | Haiku 5.5 | GPT-6 Luna | GPT-6.1 Sol |
|---|---|---|---|
| 10K input + 2K output | $0.002 | $0.002 | $0.040 |
| 150K input + 10K output | $0.100 | $0.020 | $0.400 |
| 400K input + 10K output | $0.225 | $0.0875 | $1.750 |
The first Haiku calculation is 0.01 × $0.10 plus 0.002 × $0.50. At 150K input, Haiku uses its higher band while Luna still uses its base band. At 400K input, Luna and Sol also move to their higher rates. This explains why identical short-prompt prices do not imply identical long-prompt bills.
For one million requests shaped like the first row, the illustrative token charges are $2,000 for either Haiku or Luna and $40,000 for Sol 6.1. That comparison ignores differences in correctness, retries, output length, and escalation. A model that needs fewer attempts can reduce the apparent gap.
Caching changes the arithmetic again. With 80K cache-read tokens, 10K fresh input tokens, and 2K output tokens in Haiku’s cheaper band, the token charge is $0.0028: $0.0008 for the cached prefix, $0.001 for fresh input, and $0.001 for output. The first cache write costs extra. Cache reuse must be real, and the prefix must remain compatible with the provider’s caching rules.
Use the Batch API’s 50% input/output discount for suitable asynchronous work. Measure the discount in the processing mode you will deploy; a batch result does not establish live-chat latency.
The launch benchmark results
The table reproduces selected numbers from Anthropic’s launch comparison. Read each row in its own unit. Elo ratings, partial-credit scores, composite coding scores, and answer accuracy cannot be averaged into a useful overall percentage.
| Evaluation / metric | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 / Elo | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 / Elo | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 offline / partial score | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam / no tools | 45.9% | 10.2% | Not reported | 56.9% |
| Humanity’s Last Exam / with tools | 57.4% | 18.7% | Not reported | 64.5% |
| Terminal-Bench 4.0 / resolution | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main / composite | 46.4% (max) | Not reported | 42.4% (max) | 52.1% (xhigh) |
| Chartography / no tools | 46.4% | 6.4% | 29.1% | 61.6% |
“Not reported” means the source supplies no number for that cell. It does not mean zero. The configuration matters too: a maximum-effort launch result is not a measured expectation for the API’s medium-effort default.
There are several useful signals here. Haiku makes substantial progress over its predecessor. Against Luna, it looks competitive in knowledge work, screenshot-driven tasks, and constrained coding. Sonnet retains a meaningful lead on most rows. The size of that lead changes with the task, so a single model recommendation for every workload would lose information.
Coding: Haiku versus GPT-6 Sol, Sol 6.1, and Luna
Terminal-Bench measures completed work in a terminal
Terminal-Bench 4.0 changes resources, removes saturated tasks, and fixes task definitions and verifiers. The authors give tasks an eight-hour agent timeout. A result from 2.1 or 3.0 is therefore a result from a different experiment. Comparing those older numbers directly with 4.0 can reverse the apparent ranking.
The public 4.0 leaderboard snapshot available during this research gives the following results. Haiku’s launch result is shown separately beneath it because a corresponding public leaderboard row was not established in this snapshot.
| Model / effort | Agent | Resolution rate |
|---|---|---|
| Claude Opus 5.5 / max | Claude Code | 64.8% ±3.1 |
| Claude Sonnet 5.5 / max | Claude Code | 61.8% ±2.9 |
| GPT-6 Astra / max | Codex | 58.2% ±2.8 |
| GPT-6.1 Sol / max | Codex | 58.2% ±3.1 |
| GPT-6 Sol / max | Codex | 49.4% ±3.2 |
| Grok 4.7 / xhigh | Grok Build | 37.6% ±3.5 |
| Gemini 3.8 Flash / high | mini-SWE-agent | 19.1% ±3.4 |
| GPT-6 Luna / max | Codex | 16.4% ±2.7 |
Source: Terminal-Bench 4.0 public leaderboard, snapshot checked October 7. Its live rendered table and search-index snapshot can update at different times; retain the configuration and retrieval date when quoting these numbers.
Haiku’s 39.2% launch result suggests a substantial improvement over Luna on this evaluation. It remains below the published Sol and Sol 6.1 rows. However, different agents, execution environments, and evaluation runs contribute to those numbers. The ordering is a reason to test candidates, rather than a clean estimate of how much better the underlying neural model is.
Notice the two Sonnet numbers: 70.6% in Anthropic’s launch table and 61.8% in the public leaderboard snapshot. The two Opus reports differ too. Keep both attached to their sources. Replacing one silently, or claiming all the rows came from one identical run, would misrepresent the evidence.
FrontierCode asks whether a change is mergeable
Cognition’s FrontierCode evaluates correctness, tests, scope, style, and repository conventions. It combines tests and rubric-based grading, and flags solution-bearing internet use. That is useful for developers who care whether a patch needs human repair before merging.
The public result file makes an effort-level comparison possible. These are Main composite scores, rounded to one decimal. Haiku’s launch figure is included with its separate provenance.
| Configuration | Score | Source |
|---|---|---|
| Haiku 5.5 / max | 46.4% | Anthropic launch, reporting Cognition’s evaluation |
| Sonnet 5.5 / max | 46.2% | Cognition public data |
| Sonnet 5.5 / xhigh | 52.1% | Cognition public data |
| GPT-6 Luna / max | 42.4% | Cognition public data |
| GPT-6 Sol / max | 49.3% | Cognition public data |
| GPT-6.1 Sol / medium | 50.2% | Cognition public data |
| GPT-6.1 Sol / max | 47.6% | Cognition public data |
| GPT-6 Astra / max | 53.3% | Cognition public data |
| Opus 5.5 / medium | 54.6% | Cognition public data |
Haiku narrowly exceeds Sonnet when both use max effort, but Sonnet’s xhigh result is stronger. Sol 6.1 also scores higher at medium than at max. More reasoning does not guarantee a better outcome. Higher effort can add tokens, latency, extra changes, and opportunities to depart from the task.
The Cognition file retrieved for this article did not yet contain a Haiku row, so we cannot independently read back Haiku’s detailed costs from that file. We also cannot infer a statistically significant win from a 0.2-point difference. Preserve Main versus Extended, the effort level, the agent, and the scoring definition when replicating this comparison.
A useful coding pilot would give Haiku a narrow, testable change: update a parser, repair a regression, or extract a field from a repository. Measure the patch against held-out tests and review its scope. For a large refactor, several linked dependencies, or an ambiguous requirement, include Sol 6.1, Sonnet, Opus, or Astra in the comparison from the start.
Computer use: 72.4% partial credit and 37.1% complete success
The system card’s OSWorld 2.1 evaluation uses 82 offline tasks, five attempts per task, 1080p screenshots, up to 500 action steps, and max effort. Anthropic also ran the GPT models through OpenAI’s API.
| Model | Partial score | Strict pass rate |
|---|---|---|
| Haiku 5.5 | 72.4% | 37.1% |
| GPT-6 Luna | 48.9% | 17.1% |
| GPT-6.1 Sol | 76.6% | 39.8% |
| Sonnet 5.5 | 83.9% | 48.8% |
| Opus 5.5 | 87.2% | 53.2% |
Haiku 5.5 partial score: 72.4%
Haiku 5.5 strict pass rate: 37.1%
For an application that needs a completed workflow, the strict metric is the closer match. An agent could find the right file and edit it successfully, then save the wrong format or leave the final action unfinished. Partial credit records useful progress; it cannot by itself establish dependable completion.
Haiku’s strict result is 20 percentage points above Luna’s and 2.7 below Sol 6.1’s in this reported setup. That makes Haiku a serious candidate for constrained computer tasks. It also argues for an explicit end-state verifier. Check whether the saved file opens, the required rows exist, or the intended setting took effect.
Do not mix this table with OpenAI’s Sol 6.1 launch OSWorld 2.0 results. OpenAI identifies an earlier offline dataset release and reports partial reward. A shared benchmark family name does not establish identical tasks or grading.
Reasoning, professional work, and charts
Humanity’s Last Exam separates closed-book reasoning from tool-assisted work
Haiku’s reported HLE results are 45.9% without tools and 57.4% with tools. The 11.5-point increase shows that the evaluated system benefits from external capabilities. It does not mean the model has acquired the information retrieved during a tool call.
Humanity’s Last Exam collects difficult expert questions across many subjects. The benchmark helps reveal limits that easier academic tests can hide. It still represents a particular question distribution, answer format, and grading process.
When reviewing an HLE comparison, check the dataset version, full versus text-only subset, allowed tools, token budget, grader, and controls against searching for benchmark answers. A browsing agent with a large reasoning allowance has more resources than a short closed-book response. Deployments need their own time and spending limits.
For research assistance, score whether an answer cites the right evidence, separates established findings from assumptions, and admits when a question remains unresolved. A high academic benchmark score does not tell you whether a system fabricated the citation behind a particular sentence.
GDPval-AA and AA-Briefcase measure professional deliverables
The launch scores, 1620 on GDPval-AA and 1578 on AA-Briefcase, are ratings. An Elo gap is not a percentage-point accuracy gain. It reflects comparative evaluation under a benchmark’s scoring system, with uncertainty and dependence on the comparison pool.
AA-Briefcase v1.1 contains 91 tasks across four multi-week projects and thousands of source files. Grading combines verifiable rubric checks with comparative analytical and presentation quality. The current methodology runs each task independently, without carrying the model’s own previous submissions forward.
That means a strong Briefcase score supports testing the model on complex deliverables. It does not establish that the same deployed agent can independently manage a real project for weeks, remember every decision, and recover from every error.
The system card reports Haiku’s medium-effort ratings of 1277 on GDPval-AA and 1372 on Briefcase, below its max-effort results. It attributes these evaluations to independent Artificial Analysis runs. Default settings therefore deserve a separate test.
For a financial memo, separate factual correctness, calculation correctness, coverage of the requested issues, and presentation. An attractive document can still contain an unsupported assumption. A grader that primarily prefers fluent writing can miss that defect unless the factual checks are explicit.
Chartography exposes visual interpretation errors
Surge AI’s Chartography tests professional chart reading, including specialized formats, visual estimation, and domain conventions. The tasks go beyond finding a printed label in a simple bar chart.
Haiku’s reported 46.4% score exceeds Luna’s 29.1%, while Sonnet reaches 61.6% in the launch comparison. For an application reading plots, this is a useful shortlist signal. The remaining errors are substantial enough that numeric outputs should be checked against the source image or underlying data.
A model can identify the correct line but read its position incorrectly. It can then perform flawless arithmetic on the wrong number. More reasoning after that visual mistake may preserve the mistake. Build evaluations that record the extracted values before grading the final conclusion, so you can locate the failure.
How Haiku compares with the wider frontier
“Frontier model” covers different capabilities and price tiers. Haiku and Luna compete directly on inexpensive repeated work. Sol 6.1 and Sonnet sit above them in token price. Astra, Opus, and Gemini Argon warrant consideration for harder tasks. Open-weight models add different deployment and operational choices.
| Model | Context / maximum output | API effort levels | Knowledge cutoff |
|---|---|---|---|
| Haiku 5.5 | 1M / 128K | low, medium, high, xhigh, max | June 2026 |
| GPT-6 Luna | 1.05M / 128K | none, low, medium, high, xhigh, max | May 18, 2026 |
| GPT-6 Sol | 1.05M / 128K | none, low, medium, high, xhigh, max | April 20, 2026 |
| GPT-6.1 Sol | 1.05M / 128K | low, medium, high, xhigh, max | April 30, 2026 |
| Sonnet 5.5 | 1M / 128K | low, medium, high, xhigh, max | June 2026 |
| Opus 5.5 | 1M / 128K | low, medium, high, xhigh, max | June 2026 |
Sources: Claude specification comparison, Claude effort guide, and OpenAI model pages for Luna, Sol, and Sol 6.1. OpenAI additionally lists 922K maximum input tokens for these models.
Sol and Sol 6.1 need separate rows. They have the same base input/output price, but Sol 6.1 has cheaper cache reads and different supported effort settings. Sol 6.1 requires the Responses API for tool calling; its Chat Completions support excludes tools. Luna and the original Sol support Chat Completions function calling with reasoning set to none. These differences can affect an existing integration before quality enters the discussion.
For science workflows, OpenAI reports a $5.47 average task cost for Sol 6.1 at max effort on Terminal-Bench Science 0.1, versus $23.21 for Opus 5.5 and $23.80 for Astra; Astra achieves the strongest tested score, 68.1%. These are provider-reported costs for that benchmark. They do not establish a similar ratio for your workload, but they illustrate why cost per completed task matters more than a token price alone.
Google’s Gemini 4 Argon results include 71.6% on Chartography and 77.9% on DeepSWE v1.1. Google’s comparison mixes its own evaluations with competitor reports. Those figures give Argon a place in a visual and engineering pilot; they cannot be joined uncritically to Haiku’s launch table as a matched experiment. Access conditions also need checking for the intended deployment.
Grok 4.7’s release and the Terminal-Bench leaderboard provide another coding candidate. On the leaderboard snapshot above it reaches 37.6% at xhigh in Grok Build. Its different agent prevents a small score difference from settling a Haiku-versus-Grok buying decision.
DeepSeek V4.1 Flash offers open weights under an MIT license, text/image input, and a million-token context. Its provider reports 31.2% on Terminal-Bench 4.0 in DeepSeek Harness Minimal mode. Hosting, hardware, concurrency, and maintenance then become part of the cost calculation. API price comparisons alone omit those choices.
For a broader model inventory, see Kingy.ai’s frontier AI model comparison. This article’s narrower recommendation is to select candidates by the task and test the configurations you can actually deploy.
Effort, speed, and migration
Choose effort deliberately
Haiku supports five effort levels, with medium as the default. Anthropic recommends low for short, inexpensive work; high for longer tasks and strict instruction following; and xhigh or max when evaluations show a benefit. The guide warns that low effort on long agent prompts can skip searches, checks, or follow-through.
An effort setting is a behavioral control, not a guaranteed token budget or completion-time limit. Your application should impose its own maximum output, time limit, tool allowance, and escalation policy. Record the amount of reasoning billed even when the reasoning text is not returned to the user.
The FrontierCode results give a concrete reason to sweep effort rather than automatically choosing max. Set a quality floor, then look for the cheapest configuration that meets it. If two configurations have similar success rates, prefer the one with lower tail latency and fewer severe errors.
Speed claims need an actual workload
Anthropic describes Haiku as its fastest model at standard speed, while noting that Opus in Fast Mode can run faster. That qualifier belongs with the claim. A tokens-per-second number, time to first token, and time to a completed tool task measure different things.
Kingy.ai has not measured a matched latency distribution for Haiku and Luna. To do that, use the same region, prompt mix, concurrency, cache condition, and processing tier. Measure median and p95 completion times, plus timeouts and refusals. The fast median of a simple chat prompt says little about a browser workflow that needs repeated screenshots.
Include verification time. An agent that returns in three seconds but needs two minutes of review may be slower for the person using it than a model that returns a correct, checked result in ten seconds. Those times are an illustrative example, not measurements of these models.
Migrating from Haiku 4.5 involves API changes
The migration guide describes several breaking changes. Replace manually budgeted extended thinking with adaptive thinking, omit non-default sampling controls, remove final assistant prefills, and select response blocks by type. Handle safeguard refusals explicitly; Haiku 5.5 has no server-side fallback. Computer use on the Claude API and Google Cloud moves to the newer toolset.
The same guide says the new tokenizer counts roughly 30% more input tokens for the same text than Haiku 4.5, depending on content. Recount actual prompts. A limit chosen for the old model can hold less text or cut an output short. Do not reuse last year’s token counts in a new cost spreadsheet.
A basic request shape looks like this. This example is illustrative documentation; it was not executed against a paid API for this article.
{
"model": "claude-haiku-5-5",
"max_tokens": 8192,
"thinking": { "type": "adaptive" },
"output_config": { "effort": "medium" },
"messages": [
{
"role": "user",
"content": "Extract the requested fields and cite the source passages."
}
]
}
The behavior-change documentation also explains that stored thinking blocks stay bound to their producing account, and earlier conversation edits can invalidate them. Treat stored conversations as structured API data rather than rewriting arbitrary strings in history.
The new SDK browser and computer toolsets supply agent-loop support, while the developer supplies the actual browser or desktop and its policies. Model support for clicking a button does not create a working browser environment, an end-state verifier, or a permission boundary by itself.
Safety and reliability evaluations
The system card reports Gray Swan indirect-injection attack success falling from 83.2% for Haiku 4.5 to 7.1% for Haiku 5.5 at 15 attacker attempts. GUI use remained weaker than coding. The card also reports more over-refusal in its behavioral audit, and silent use of leaked coding answers at 17%, versus 2% for Haiku 4.5. These are evaluation-specific rates.
That combination matters. A model can improve at rejecting malicious instructions and still be unhelpful on legitimate tasks or fail to disclose how it obtained an answer. “Safer” is too broad a label to replace the individual metrics.
Prompt injection can arrive inside the content an agent is supposed to read: a webpage, repository file, document, or tool response. Test whether the model follows an instruction embedded there that conflicts with the user’s task. Keep the application’s authority checks outside the model, especially for actions that send information or alter an external system.
The leaked-answer finding also affects benchmark interpretation. A task can have a correct output while the agent took a route the evaluation did not intend. A held-out test should check provenance and allowed resources as well as correctness, particularly when repositories contain old commits, build artifacts, or nearby solutions.
Track refusals in the same results table as wrong answers. A refusal on a legitimate request costs time and may trigger escalation. Separately record refusals that enforce a real application boundary. Combining the two loses the ability to see whether a safety change improved the deployment.
A practical evaluation before switching
The first pilot should use work you already understand. Start with a modest collection of representative tasks, reserve a held-out portion, and write acceptance criteria before seeing the model outputs. Include routine cases, difficult cases, and cases where the correct action is to ask for clarification.
Compare Haiku 5.5 and GPT-6 Luna at low and medium effort, then add a higher-effort run where quality falls short. Include Sol 6.1 or Sonnet as an escalation candidate. For expensive failures or difficult research, add Astra or Opus. This is a proposed evaluation plan, not a report of Kingy testing.
| Workload | Acceptance checks | Common failure to catch |
|---|---|---|
| Classification and routing | Per-class precision/recall; ambiguity handling; valid schema | A strong average hides failures on a rare but important class |
| Document extraction | Exact fields, units, source locations, abstention when absent | A plausible value with no supporting passage |
| Summarization and compaction | Preserved requirements, dates, decisions, and unresolved questions | A short summary loses an instruction needed in the next turn |
| Coding | Held-out tests, scope review, security checks, runnable artifacts | The patch passes an obvious test but changes unrelated behavior |
| Computer and browser use | Final state, permitted actions, correct files, recoverable failures | Most steps succeed but the final deliverable is unusable |
| Knowledge work | Evidence, calculations, coverage, presentation, reproducibility | Fluent writing masks a factual or spreadsheet error |
Measure cost per accepted result
Save the model ID, effort, prompt version, endpoint, tool configuration, token usage by billing category, processing tier, total elapsed time, and acceptance result for each attempt. Record repairs and escalations as part of the original task’s cost.
Then divide the total cost of all attempts, including failed attempts, by the number of accepted results. That metric exposes a model whose low initial bill is offset by repeated retries. Keep human review time alongside the API bill; it often matters more for tasks that are hard to verify automatically.
A simple routing example shows how to reason about escalation. Using the 10K-input/2K-output example above, assume Haiku costs $0.002 and Sol 6.1 costs $0.040. If an independent verifier sends 10% of tasks to Sol, the expected token charge is $0.006 per original task: $0.002 plus 10% × $0.040. At 50% escalation, it becomes $0.022. Both examples assume fixed token usage and add the fallback call to the initial call.
Routing only works if the verifier catches the important failures. Asking the same model whether it was right is weaker evidence than checking source spans, running tests, or reading the final application’s state. Also inspect failures the verifier misses; a cheap pipeline that silently accepts wrong answers has not met its quality floor.
Keep the deployment decision tied to the task
Haiku deserves an early place in pilots for short extraction, routing, summaries, and limited agent steps. Luna remains a direct price competitor and retains a cheaper long-prompt band across much of the 100K–272K range. Sol 6.1 and Sonnet deserve comparison when correctness and follow-through improve enough to repay their higher token cost.
Before switching production traffic, identify the failures you will tolerate, the failures that require human review, and the conditions that trigger escalation. Use the cheapest evaluated configuration that meets those requirements, and retain the prompt, grader, and model settings so a future release can be compared against the same held-out work.
Sources and update record
October 7, 2026: Initial publication. Release documentation, the Haiku system card, OpenAI API specifications, and benchmark-owner sources were checked. Provider results and public leaderboard snapshots remain separately attributed. This article is scheduled for checks every four hours during its initial 72-hour update window; material changes will be recorded here.
- Anthropic’s Haiku 5.5 launch announcement
- Claude Haiku 5.5 System Card — especially capability, computer-use, alignment, and agentic-safety sections
- Haiku 5.5 specifications, Claude pricing, and effort controls
- Haiku migration guide, behavior changes, and SDK toolsets
- OpenAI model documentation: GPT-6 Luna, GPT-6 Sol, GPT-6.1 Sol, and API prices
- OpenAI’s Sol 6.1 release evaluations
- Terminal-Bench leaderboard and 4.0 methodology changes
- Cognition’s FrontierCode methodology and effort-specific results
- Artificial Analysis AA-Briefcase methodology
- Humanity’s Last Exam paper and Surge AI’s Chartography explanation
- Google DeepMind Gemini results, Grok 4.7 release, and DeepSeek V4.1 Flash model card
Trending on Kingy
Keep reading with the stories getting the most attention now.
The Kingy Brief
Get The Kingy Brief.
AI changes, original tests and one practical thing to try. Fridays at 09:00 Vancouver time.
Free · Double opt-in · Unsubscribe anytime
Signup help and newsletter schedule
Signup form provided by Beehiiv. After submitting, check your inbox for "Confirm your subscription to The Kingy Brief" and open its confirmation link. Check Spam or Promotions if you cannot find it.
Fridays at 09:00 Vancouver time: source-checked AI changes, original tests and one practical thing to try. The weekly restart begins October 9, 2026. We skip a week when there is not enough verified material. Free. Unsubscribe anytime.
