AI News

GPT-6.1 Sol and GPT-6 Luna vs Claude Haiku 5.5: Specs, Benchmarks and Real Costs

Published October 7, 2026. Sources checked October 7, 2026. This is a source-based comparison of public documentation and evaluator results. Kingy has not run matched hands-on tests of these three models. Prices are US dollars before taxes; worked examples are calculations, not invoices.

Claude Haiku 5.5 gives OpenAI's inexpensive Luna model a serious new competitor. For prompts of up to 100,000 tokens, Haiku 5.5 and GPT-6 Luna have identical headline API prices: $0.10 per million input tokens and $0.50 per million output tokens. Haiku's launch evidence shows substantial gains over Luna in several practical tasks. Above that prompt length, however, Haiku's rates rise sharply. Anthropic's launch announcement, OpenAI's Luna specifications.

GPT-6.1 Sol occupies a different price tier: $2 input and $10 output per million tokens. Its clearest advantage in the evidence below is difficult terminal coding. Paying twenty times the small models' short-prompt rates can make sense when the cheaper model leaves work unfinished, but it needs to be justified by the task. OpenAI's Sol specifications.

Our reading of the evidence is that Haiku deserves a place in evaluations for economical extraction, document work and computer use; Luna remains compelling for high-volume work and longer prompts; Sol 6.1 has the clearest advantage here for difficult terminal coding. There is enough evidence to make those distinctions. There is not enough matched evidence to declare one model the winner at everything.

Which Sol and Luna are we comparing?

“Sol Luna” refers to two OpenAI models. The main comparison uses GPT-6.1 Sol (gpt-6.1-sol), GPT-6 Luna (gpt-6-luna) and Claude Haiku 5.5 (claude-haiku-5-5). OpenAI's current catalogue identifies Sol 6.1 as its balanced model and Luna 6 as its model for economical, high-volume work. Sol 6.1 arrived September 29; Luna 6 arrived September 22. Haiku 5.5 launched October 7. OpenAI model catalogue, OpenAI changelog, Haiku overview.

The older GPT-6 Sol still matters because some published comparisons refer to it. It is a separate model with a different knowledge cutoff and a $0.20 cached-input rate, compared with $0.10 for Sol 6.1. We identify it explicitly wherever it appears. A result labelled “GPT-6 Sol” cannot be silently reassigned to Sol 6.1. GPT-6 Sol documentation.

This is principally a model and API comparison. A ChatGPT or Claude subscription adds its own interface, tools, allowances and routing rules. API token rates do not establish what a subscriber pays for a completed task inside either app.

Specifications: context, output, inputs and reasoning

Specification Claude Haiku 5.5 GPT-6 Luna GPT-6.1 Sol
Model ID claude-haiku-5-5 gpt-6-luna gpt-6.1-sol
Published context window 1,000,000 tokens 1,050,000 tokens 1,050,000 tokens
Standard maximum output 128,000 tokens 128,000 tokens 128,000 tokens
Inputs → output Text and images → text Text and images → text Text and images → text
Knowledge cutoff June 2026 May 18, 2026 April 30, 2026
Default reasoning effort Medium Medium Medium
Effort levels Low, medium, high, xhigh, max None, low, medium, high, xhigh, max Low, medium, high, xhigh, max
Publicly disclosed parameter count Not established in sources checked Not established in sources checked Not established in sources checked
Downloadable weights No public download identified No public download identified No public download identified

Sources: Haiku model documentation, Claude effort documentation, Luna documentation, Sol 6.1 documentation.

The OpenAI models have a nominally 5% larger token window. That arithmetic is straightforward; the amount of source material that fits is less straightforward. Different tokenizers assign different counts to the same text, tables, code and images. A million tokens from one provider should not be treated as a million tokens from another provider.

Anthropic says the newer Haiku tokenizer produces approximately 30% more tokens for the same text than Haiku 4.5. That is a comparison with the previous Haiku tokenizer, not a measured 30% difference from OpenAI. An old 80,000-token Haiku prompt could, illustratively, become about 104,000 tokens and cross Haiku 5.5's price threshold. Count the actual prompt with the new model before budgeting. Haiku tokenizer changes.

Output limits also need context. The input, generated reasoning and answer must fit the applicable window. A 128,000-token output limit is a ceiling, not a promise of a useful response that long. Haiku additionally documents a 300,000-token output beta for Message Batches; this is a separate batch capability, not the standard interactive limit. Haiku capabilities.

All three can reason about supplied images and return text. A model seeing a screenshot does not, by itself, give it a functioning browser or permission to operate a computer. Those abilities come from tools and an application that executes the model's actions. Similarly, calling an image-generation tool is different from the language model natively producing an image.

The cutoff dates are useful for deciding when retrieval is necessary, but they do not rank factual accuracy. A later cutoff does not guarantee knowledge of a particular event, source or price. For current information, the retrieval process and source checks matter more than the date printed on the model card.

Pricing: the 100K and 272K boundaries change the comparison

These are direct-provider Standard text-token rates per million tokens. “Input” means uncached input. Cache writes and reads have their own rates.

Model and input length Input Cache read Cache write Output
Haiku 5.5, up to 100K $0.10 $0.01 $0.125, 5-minute cache $0.50
Haiku 5.5, over 100K $0.50 $0.05 $0.625, 5-minute cache $2.50
Luna 6, up to 272K $0.10 $0.01 $0.125 $0.50
Luna 6, over 272K $0.20 $0.02 $0.25 $0.75
Sol 6.1, up to 272K $2.00 $0.10 $2.50 $10.00
Sol 6.1, over 272K $4.00 $0.20 $5.00 $15.00

Sources: Anthropic pricing, OpenAI pricing, Sol 6.1 model rates. Haiku's one-hour cache writes cost $0.20 below its threshold and $1 above it.

Haiku charges five times its short-prompt rates once the prompt exceeds 100,000 tokens. OpenAI's threshold is above 272,000 input tokens, where input and cache rates double and output rates rise by 50% for the full request. These are request tiers, so budgeting as though only the excess tokens receive the higher price understates cost. Haiku long-context pricing, OpenAI model pricing conditions.

That creates three useful cost regions. Up to 100K, Haiku and Luna share the listed rates. Between 100K and 272K, Haiku's listed rates are five times Luna's. Above 272K, the ratio depends on the mix: Haiku input is 2.5 times Luna input, while Haiku output is about 3.33 times Luna output. Those are tariff comparisons for equal token counts, not measurements of how many tokens each model needs for the same job.

Sol 6.1 costs twenty times Luna on uncached input and output below OpenAI's threshold. Its cached-input rate is ten times Luna's. A cache-heavy workflow therefore has a different price ratio from a fresh-prompt workflow. Comparing only the input price hides that difference.

Worked costs for actual request shapes

The following calculations use uncached input and the stated number of total billable output tokens. They exclude tool fees, retries, regional premiums and taxes. If a model thinks for another 20,000 billable tokens, those tokens must be added; the answer length alone is insufficient.

Input / billable output per request Haiku 5.5 Luna 6 Sol 6.1
5K / 1K $0.001 $0.001 $0.020
100K / 2K $0.011 $0.011 $0.220
100,001 / 2K $0.0550005 $0.0110001 $0.220002
150K / 5K $0.0875 $0.0175 $0.3500
300K / 5K $0.1625 $0.06375 $1.2750
900K / 10K $0.4750 $0.18750 $3.7500

Calculated from the cited provider rate tables. For the 5K/1K request, 100,000 calls cost $100 on Haiku or Luna and $2,000 on Sol. For the 150K/5K request, the same volume costs $8,750, $1,750 and $35,000 respectively. These are hypothetical budgets with fixed token usage; they are not observed customer bills.

The one-token threshold example deliberately isolates the tariff change. Moving from 100,000 to 100,001 input tokens increases that Haiku request from $0.011 to $0.0550005. The input barely changes, but the request enters a different tier. A long system prompt, tool schema or accumulated conversation can push an application over that boundary even when the user's newest message is short.

For caching, take an eligible 80K-token shared prefix, 1K of fresh input and 1K of output. A first request priced as a cache write costs $0.0106 on Haiku or Luna and $0.212 on Sol. A subsequent full-prefix cache hit costs $0.0014, $0.0014 and $0.020. This assumes a genuine hit within the applicable retention period. It does not guarantee that every call to a growing conversation will hit the cache.

OpenAI states that cache-write pricing replaces the ordinary input rate for those tokens; it is not an extra input fee. Keep cache reads, cache writes and fresh input in separate buckets. Cache reuse also depends on an unchanged prefix, so moving volatile data to the end can matter more than a small model-price difference. OpenAI prompt caching.

Both providers document a 50% input/output discount for asynchronous Batch processing; OpenAI also lists Flex at half Standard rates. Batch can be useful for overnight classification or extraction, but it is a different service from an interactive answer. Provider routes, tools and residency choices can add other charges. Compare the route you will deploy, not an abstract lowest advertised price. Claude Batch pricing, OpenAI pricing modes.

Benchmarks: who ran each test matters

A benchmark result describes a particular model, effort level, agent, dataset version, tool configuration and grading method. It is evidence about that setup. Each table below retains those distinctions.

Evidence source What is established Main limit
Anthropic launch and system card Haiku launch results; Anthropic-run desktop-use comparison; evaluator-supplied results reproduced in the card Some results originate with the vendor; headline settings can be expensive
Artificial Analysis Independently evaluated knowledge-work results cited in Haiku's card; directly retrieved Sol/Luna model data A full public Haiku Intelligence Index was not verified at this check
Terminal-Bench public leaderboard Current Sol/Luna results in Codex Haiku was absent from the public table checked; its score comes from Anthropic
Cognition FrontierCode leaderboard Current Sol/Luna coding results and task methodology Haiku's announced result was not yet in the public table checked
Kingy Source verification, arithmetic and editorial analysis No matched model execution or measured production latency

Primary sources: Haiku system card, Artificial Analysis Sol, Artificial Analysis Luna, Terminal-Bench, FrontierCode. Public leaderboard availability is a launch-day observation, not a claim that the evaluator has never tested Haiku.

Knowledge work: Haiku is competitive, with an effort caveat

Evaluation Haiku 5.5, max Luna 6, max Sol 6.1, max
GDPval-AA v2.1, Elo 1,620 1,437 1,575
AA-Briefcase v1.1, Elo 1,578 1,336 1,564

Haiku figures come from independently run Artificial Analysis evaluations reported in Anthropic's launch material; Sol and Luna values were retrieved from AA's Sol page and Luna page. These are dated evaluator snapshots, not a Kingy-run simultaneous comparison.

GDPval-AA evaluates professional deliverables. AA-Briefcase examines longer projects with linked tasks and source files. The results support taking Haiku seriously for work that used to require a more expensive model. They provide a much more useful reason to trial Haiku than a general claim that every new model is smarter.

Haiku leads Luna by 183 Elo on GDPval-AA and 242 Elo on Briefcase in these snapshots. Its leads over Sol are much smaller: 45 and 14 Elo. Elo is a relative rating, not a percentage of tasks completed. A 14-point difference does not establish a decisive practical advantage, particularly without a matched uncertainty analysis across these published figures. AA benchmark methodology.

At Haiku's default medium effort, the system card reports GDPval-AA 1,277 and Briefcase 1,372. The corresponding output-token use is roughly one tenth and less than one quarter of max respectively. Those results make effort part of the product decision: max's strong work-product scores do not describe default-cost operation. System card, sections 8.10.2–8.10.3.

For a document team, useful checks would include extracting the correct revenue line, preserving units, reconciling contradictory source files, producing a valid workbook and attaching citations that support each conclusion. A polished answer with the wrong denominator is still wrong. Work-product benchmarks are relevant because they move beyond short question answering, but your own documents can expose different failure modes.

Terminal coding: Sol has the strongest evidence

Terminal-Bench 4.0 result Resolution rate Setting and source
GPT-6.1 Sol 58.2% ± 3.1 pp Max; Codex; public leaderboard
GPT-6 Sol, older version 49.4% ± 3.2 pp Max; Codex; public leaderboard
Claude Haiku 5.5 39.2% Max; Claude Code; Anthropic run
GPT-6 Luna 16.4% ± 2.7 pp Max; Codex; public leaderboard

Sources: live Terminal-Bench leaderboard and Haiku system card, section 8.4. Public leaderboard intervals are labelled 95% confidence intervals; pp means percentage points. Haiku's run had safeguards enabled, no fallback and restricted internet access.

The practical conclusion is that Sol 6.1 is the better-supported candidate for difficult terminal work among these three. Its numerical lead over Haiku is 19 percentage points in the cited runs. Haiku's lead over Luna is 22.8 points. Different agents and execution policies prevent these from being pure measurements of the underlying models under identical conditions.

Terminal-Bench 4.0 includes demanding work in command-line environments. Success can require navigating a repository, modifying code, diagnosing a dependency problem and verifying an outcome. A cheap model that writes plausible code but repeatedly fails to finish can consume developer time, retries and tool resources that dwarf its token bill.

There is another reason to retain the source labels. Artificial Analysis's separate runs show Sol at 56.1% and Luna at 12.6% on Terminal-Bench 4.0, rather than the public leaderboard's 58.2% and 16.4%. These are separate evaluation runs; averaging them or substituting one value into another source's table would obscure the difference. AA Sol data, AA Luna data.

Mergeable code: highest effort is not always best

FrontierCode 1.1 Main Score Effort Evidence
GPT-6.1 Sol 50.2% Medium, best displayed result Cognition public leaderboard
Claude Haiku 5.5 46.4% Max Cognition-run result reported in Anthropic's card
GPT-6 Luna 42.4% Max, best displayed result Cognition public leaderboard

Sources: Cognition leaderboard and Haiku system card, section 8.3. This is a best-displayed-result comparison, not an equal-effort sweep. Haiku's announced row was not yet visible in the public leaderboard checked.

FrontierCode evaluates whether a patch meets functional requirements, tests, scope and repository standards. Its score differs from its pass-rate column, so the two must not be swapped. Version 1.1 also zeros runs flagged for consulting solution-bearing internet sources. Cognition's methodology.

For an engineering team, scope discipline is a useful dimension. A model can pass tests while changing files the task never required, weakening a check or leaving a patch expensive to review. Evaluating the final diff helps catch problems that a single code-generation score misses.

Sol's displayed best result is at medium effort. This is a concrete reason to evaluate several reasoning settings rather than automatically selecting max. A larger reasoning budget can increase cost without improving the final patch. The gaps in this table are also small enough that a matched repeat evaluation would be more persuasive than declaring a universal coding winner from them alone.

Desktop use: partial progress and complete success differ

OSWorld 2.1 offline subset, max effort Partial-credit score Strict pass rate
GPT-6.1 Sol 76.6% 39.8%
Claude Haiku 5.5 72.4% 37.1%
GPT-6 Luna 48.9% 17.1%

Source: Anthropic's system card, Figure 8.9.3.A. Anthropic ran all three on the same 82 offline tasks, with five attempts per task and provider-specific compaction. This is a vendor-run matched task comparison.

The distinction between the columns is essential. Haiku's 72.4% score measures partial progress; its fully successful attempts were 37.1%. A desktop assistant that completes most checkpoints but saves the wrong file can feel capable while still leaving the user to finish the job.

Haiku is numerically much closer to Sol than to Luna on this test. Sol still leads, and the modest Haiku–Sol gaps should not be described as proof of equivalence or a statistically decisive difference. Offline tasks also do not establish reliability on live websites with changing content, sign-in steps or network errors.

For production browser work, measure complete task success and the number of interventions. Reading a page accurately, clicking the correct control, recovering from an error and leaving the final state correct are separate abilities. A screenshot-recognition benchmark captures only part of that workflow.

Reasoning, charts and document accuracy

Haiku's launch table reports Humanity's Last Exam at 45.9% without tools and 57.4% with tools, and Chartography at 46.4% without tools versus Luna's 29.1%. Keep the tool conditions attached to those numbers. The with-tools HLE figure is not a direct substitute for a no-tools result. Anthropic benchmark table.

Additional Artificial Analysis results, max effort Sol 6.1 Luna 6
Intelligence Index v4.3.2 51.8 38.1
Humanity's Last Exam, AA no-tools run 52.9% 38.5%
GDP.pdf, all-pass 31.0% 22.8%
AA-LCR v1.1 83.0% 83.3%

Sources: AA Sol, AA Luna, AA methods. We have not filled Haiku's missing independent full-index row with an estimate. Its launch HLE result and AA's runs have different provenance and should be compared cautiously even where both exclude tools.

These results show why a composite score needs a breakdown. Sol leads on the overall index and the cited PDF result. Luna is close on this particular long-context test. A similar score on one retrieval-heavy task does not imply comparable performance on difficult coding, and a larger context window does not guarantee better document reasoning.

For chart analysis, separate reading values from reasoning with them. A model might correctly identify a plotted line yet miss whether the axis is logarithmic, confuse a confidence interval with a range, or use an inappropriate baseline. For PDFs, ask it to identify the page and source cell behind each number before checking the calculation. Those checks provide clearer evidence than whether the prose sounds authoritative.

Speed: token generation is only one part of latency

Artificial Analysis's retrieved first-party measurements showed roughly 55.5 output tokens per second for Sol 6.1 at max and 127.9 for Luna at max. Its weighted index costs were approximately $0.72 and $0.07 per task. These values describe the evaluator's workloads and current serving measurements, not a promised response time or tariff for your application. AA Sol–Luna comparison.

Anthropic calls Haiku 5.5 its fastest model at standard speeds, with a qualification for Opus Fast Mode. That is a comparison within Anthropic's range. We have not independently verified a matched Haiku-versus-Luna throughput or time-to-answer measurement, so we cannot translate it into “Haiku is faster than Luna.” Haiku launch speed qualification.

An application has at least four relevant timings: input processing, reasoning before an answer, answer generation and tool execution. A model can generate tokens quickly and still spend a long time thinking before producing useful text. A streaming indicator may improve the user's experience without shortening the time to a correct final outcome.

Test latency at the effort level you intend to use. A support assistant that needs a short routing label has different requirements from a repository agent working for twenty minutes. Record median and slow-tail completion times, not only tokens per second. Include retries and failed calls when measuring the experience the user receives.

Luna's new Decisions API has different economics

OpenAI also released its Decisions API in public beta on October 7, initially supporting only gpt-6-luna. It evaluates text and images and returns a condition probability, a choice from supplied options, or a score against a rubric. OpenAI describes it as about ten times faster than Responses; that is a provider claim about two OpenAI endpoints, not a measured speed advantage over Haiku. OpenAI changelog, Decisions documentation.

For this endpoint, Luna costs $0.10 per million input tokens, with no cache-read, cache-write or output-token charges. Long-context multipliers and regional premiums still apply. A short 5K-input classification request therefore has a base token cost of $0.0005, or $50 for 100,000 such calls. This is a calculated example. The earlier cost table covers ordinary input/output requests and should not be applied unchanged to Decisions. Decisions pricing.

That makes the endpoint relevant to high-volume routing and fixed-label classification. Use Responses and structured outputs when the job requires a custom extracted object or written explanation, and function calling when it requires a tool request. The comparison should follow the task's required output, not just the model name.

Tool support and migration costs

Both OpenAI models support function calling, structured outputs and a broad set of tools through Responses. Sol 6.1 requires Responses for tool calling; its Chat Completions support excludes tools. Luna allows function calling through Chat Completions only with reasoning effort set to none. This can affect a migration even when the application changes only its model ID. OpenAI GPT-6 integration guidance.

Haiku 5.5 supports structured output and tool use, but an existing Haiku 4.5 integration needs more attention. Anthropic's migration guide changes thinking configuration, removes assistant prefill, restricts sampling parameters and changes the computer-use toolset. Its safety classifiers can return a refusal with no server-side fallback, and Priority Tier is not supported. Haiku migration guide.

The practical review is short: inspect request parameters, parse content blocks by type, preserve required thinking blocks in tool loops, recount tokens, and handle refusal or output-limit termination explicitly. A model that accepts the request but returns only a thinking block can break code that assumes the first block contains an answer.

Haiku supports low through max effort and defaults to medium. Thinking can be disabled at high effort or below, but that combination is rejected at xhigh and max. Anthropic recommends raising effort only where evaluations demonstrate a useful gain. Haiku effort guidance.

An integration that already has reliable retrieval, validated tool arguments and predictable output schemas can be more valuable than a small benchmark lead. Account for migration work in the decision. Retesting a mature application can cost more than the token savings from a modest traffic volume.

Safety, factual errors and refusal behavior

Haiku's system card reports improvements in prompt-injection resistance and honesty under pressure, alongside increased over-refusal in one behavioral audit. Those are vendor evaluations with particular test conditions. They do not establish a cross-provider safety ranking. Haiku safety assessment.

A refusal can be the correct behavior, an unnecessary interruption or evidence that the application is asking the wrong tool to perform a task. Log those cases separately. Counting every refusal as a failure encourages unsafe behavior; treating every refusal as a safety success hides usability problems. The correct outcome depends on the request and the model's permitted role.

Prompt injection also remains an application concern. An agent can encounter instructions inside a webpage, issue comment or document that conflict with the user's task. Keep untrusted content distinguishable from instructions, limit tool permissions, validate consequential actions and check the resulting state. Model resistance helps, but it cannot replace the application's controls.

For factual work, evaluate correct answers, unsupported answers and appropriate uncertainty separately. Requiring a citation is useful only if somebody checks that the cited page supports the claim. A low-cost model that extracts verified passages can be valuable; using the same model to invent a source-backed conclusion without verification defeats that purpose.

Which model should you choose for your workload?

The recommendations below are editorial judgments from the evidence above, rather than claims of a Kingy-tested production winner.

Workload Sensible starting point What should decide the final choice
Fixed-label classification and routing Evaluate Haiku and Luna, including Luna Decisions Correct labels, supported answer type, cost and completion latency
Structured extraction Evaluate Haiku and Luna with schema-constrained output Correct fields, source evidence and valid schemas
Summaries and professional documents under 100K input Haiku deserves an early trial Source fidelity, numeric accuracy and cost at the selected effort
Repeated 100K–272K prompts Luna has the clear listed-price advantage Whether Haiku's quality improvement justifies the higher tier
Difficult repository and terminal tasks Sol 6.1 Accepted patches, complete execution, review time and retries
Screenshot-driven desktop tasks Trial Haiku and Sol 6.1 Strict completion and interventions, not partial progress alone
Very long document workflows Compare Luna's cost with Sol's quality Retrieval accuracy, citations, prompt size and complete outcomes
Existing well-performing production integration Keep it as the control A challenger must improve the outcome enough to pay for migration

For customer support, the cheapest sensible first step may be classification and retrieval, followed by a stronger model only when the ticket needs reasoning across contradictory facts. A model's general coding score is less important than getting escalation, account details and the permitted response right.

For content and research teams, Haiku's work-product results justify testing it on bounded tasks such as pulling facts from a filing or preparing a first structured summary. Retain a separate check of names, dates, prices and calculations. Sol can handle the difficult synthesis step if the cheaper model's output proves unreliable.

For software teams, use Sol 6.1 as a serious candidate for tasks with expensive failure modes. Haiku or Luna may be enough for explaining a small function, extracting symbols or making a narrowly defined edit. Evaluate the finished patch and its tests; count the time a developer spends correcting it.

Routing can improve economics, but it needs evidence. Sending every task to a cheap model and escalating half of them can cost more than a well-chosen initial route. A fictional example illustrates the tradeoff: if 1,000 fixed-shape requests cost $1 on a small model and 10% need a $0.02 Sol retry, the model bill becomes $3. If repeated failures require long contexts or human review, that arithmetic changes quickly. This is a routing calculation, not a measured success rate.

How to evaluate these models fairly

Build a small, representative set of real tasks before migrating. Include straightforward examples and the difficult cases that consume time today: inconsistent documents, ambiguous instructions, long inputs, tool errors and tasks that require admitting uncertainty.

Keep the source files and expected outcomes fixed. Record the precise model ID, reasoning effort, prompt, tool access, token usage and date. If you compare provider-specific agents, say so; if you want to isolate the model, keep the agent's instructions and tools as consistent as possible. Either experiment can be useful, but they answer different questions.

Grade the result before calculating the savings. For classification, check the label. For extraction, check the field and evidence. For coding, run the necessary checks and review scope. For computer use, inspect the final file or application state. A single successful example is a demonstration; repeated success across representative cases is better evidence for adoption.

Then calculate cost per accepted outcome. Include unsuccessful attempts, billable thinking, cache misses, tool charges and review time. For example, a $0.001 call with a hypothetical 50% acceptance rate costs $0.002 per accepted result before retries or labor. That can be a worse purchase than a more expensive call that produces usable work consistently. The figures here illustrate the method, not these models' observed acceptance rates.

Measure enough repetitions to notice unstable behavior. A near tie should remain a near tie until the evidence separates it. Resist adjusting the test prompts for one model after seeing its failures while leaving the other model's prompts untouched. If you improve the prompts or agent, give every candidate the same opportunity and document the revision.

Updates and unresolved launch-day questions

This article will be checked every four hours during a bounded 72-hour update window. Material changes to specifications, rates, benchmark availability or methodology will receive dated revisions. A successful check that finds no relevant change will not be presented as a new model result.

The most useful next evidence would be a complete public Artificial Analysis result for Haiku 5.5, public leaderboard confirmation of its announced coding results, matched latency and cost measurements at practical effort settings, and reports with enough retained artifacts to inspect real task outcomes. Provider availability in a particular account or region also needs deployment-specific verification.

Use these published results to select candidates, then choose the model, endpoint and effort that produce correct outcomes on your own tasks at an acceptable total cost.

Revision history

  • October 7, 2026: Initial comparison published with official specifications, request-tier pricing, calculated costs, dated benchmark sources, effort distinctions and launch-day evidence limits. No Kingy-run model benchmark results are claimed.