October 7, 2026 · Subscription-client benchmark · No subjective model judge used.
We ran Claude Haiku 5.5 and GPT-6 Luna through 40 original tasks, twice per model. All 160 scored responses completed, after eight separate setup pilots. Added metered spending was $0.00 because the test used existing subscriptions. The account checks showed no consumed paid credits, and we bought no add-ons, credits or hosting.
On the frozen all-check metric, Haiku passed 60 of 80 outputs (75.0%) and Luna passed 51 of 80 (63.75%). Median end-to-end completion time was 6.83 seconds for Haiku and 2.92 seconds for Luna. Those are real observations, but neither the score nor the timing should be read as a universal model ranking.
The scoring catch matters. Some prompts allowed ordinary language in a field while the answer key expected one exact phrase, type or array order. Correctly preserving uncertainty could therefore fail a deterministic check. We kept the frozen scores, saved the responses and exposed that limitation instead of rewriting the rules after seeing the answers.
What was actually tested
The eight categories were AI news summaries, product comparisons, JSON extraction, instruction following, practical calculations, small coding fixes, YouTube hooks and outlines, and article or newsletter editing. Each had five tasks and carried 12.5% of the overall weight. The source packets were clearly labeled synthetic fixtures, with fictional products and fixed facts.
That design supports repeatability: the evidence does not change halfway through a run. It also limits generalization. This was not a test of live research, browsing, large repositories, breaking-news coverage or sustained creative collaboration. The models answered bounded requests using only the supplied packets.
Both exact model versions were retained. The Claude client requested and reported claude-haiku-5-5. Codex requested gpt-6-luna, and its authenticated catalog listed that model. Codex’s response events did not expose a separate backend model identifier, so the evidence distinguishes a verified selection from a returned backend identity.
The task bank, reference answers, tests, category weights and example-selection rules were written and hashed before scored outputs. Four different unscored pilots per model checked basic extraction, JSON handling, arithmetic and missing information. We repaired client configuration during that setup phase, preserved the failed local launch, then froze the scored configuration. No poor scored answer was retried.
The results, with the right denominator
| Category | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| News summaries | 4/10 | 2/10 |
| Product comparisons | 0/10 | 0/10 |
| JSON extraction | 10/10 | 10/10 |
| Instruction following | 10/10 | 10/10 |
| Calculations | 8/10 | 4/10 |
| Coding fixes | 10/10 | 10/10 |
| YouTube hooks/outlines | 10/10 | 10/10 |
| Editing | 8/10 | 5/10 |
| Total | 60/80 | 51/80 |
An output passed only if every applicable frozen deterministic check passed. Partial credit is retained separately in the machine-readable records. Request completion and answer correctness are different measures; the run recorded zero scored request failures.
For the overall comparison, we averaged repetitions within each task and then resampled whole task IDs within categories. The 10,000-draw paired bootstrap estimated a Haiku-minus-Luna difference of 11.25 percentage points, with a 95% interval from 3.75 to 20.0 points. The interval excludes zero on this strict metric, but it does not correct the exact-match limitations or establish a broader quality ranking.
Two answers to one prompt are not two independent tasks. Five tasks per category is also a small sample. The category table is useful for locating behaviors to inspect, but it cannot establish a general reliability rate for every editing, coding or extraction workflow.
Why a failed check is not automatically a false claim
The conflicting-release task, N3, supplied two same-date statements: one announced a launch date, while the other reported a delay without a replacement date. The models’ saved responses kept the release date null and explained the conflict. The key, however, expected the evidence-status field to contain exactly “conflicting.” A fuller explanation failed that equality check.
The missing-exchange-rate task, C5, offers another useful caution. The outputs left the converted amount null and named missing conversion information. The reference required a particular wording in the missing-information array. Those failures should be inspected as exact-key mismatches, not counted blindly as invented exchange rates.
Other checks test clearer requirements, such as explicit word bounds, arithmetic, strict root keys and independent code behavior. Even here, the reader should inspect what was actually specified. The saved evidence includes every failed check, its reference and the unedited response. It does not silently replace the original score with a more favorable interpretation.
There were also substantive numerical differences. C1 explicitly priced 10,000 uncached input tokens, 2,000 cached input tokens and 2,000 output tokens including reasoning. The correct total was $0.00202. Haiku returned that value twice; Luna returned $0.00087 and $0.00079. Conversely, both passed all ten extraction and all ten coding outputs. Both scored zero on comparisons, where broad exact-object equality makes the saved answers more informative than that category percentage alone.
The saved report’s six comparison examples follow rules set before scoring: shared success, shared failure, the largest advantage for each model, missing or conflicting information, and writing repeat disagreement. Both repetitions remain available, including when a selection category has no eligible example and the documented fallback applies.
The subscription route changes the interpretation
We first checked the requested API routes. The accessible Vercel account showed no active key or credit balance, and Haiku 5.5’s Gateway listing did not offer free-credit eligibility on the listed routes. Existing OpenRouter funding and activated Claude API credits were not verified. No credit was purchased.
We selected a route below $10 using existing ChatGPT Pro and Claude Max subscriptions through their supported command-line clients. Claude’s desktop login did not initially carry over to the separately launched client; completing the normal account connection resolved that setup issue.
This required a recorded protocol change. The OpenAI subscription route did not expose the original common 8,192-token ceiling. The run therefore used client defaults without claiming equal internal compute. Inputs and outputs, including available reasoning-token counts, are in the saved evidence. All observed total inputs stayed below 20,000 tokens; native tokenizer pre-counts were unavailable.
Identical task content does not make the surrounding clients identical. They add different wrappers, and their process startup and service paths differ. Calls ran one at a time in a saved randomized paired order, with fresh contexts and no external model tools. The reported latency includes the whole client invocation, not just model inference. It supports a statement about this tested workflow, not an intrinsic speed claim.
| Observed measure | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Median client completion time | 6.83 seconds | 2.92 seconds |
| Scored request failures | 0/80 | 0/80 |
| Additional metered cost per passing output | $0.00 | $0.00 |
What the zero-dollar figure means
The hard ceiling was $9.99. Before further requests, the ledger reserved a conservative $2 maximum covering documented context and output maxima at upper tariff rates. Missing usage would retain its hold. Returned usage reduced the estimate, and independent billing evidence allowed included-subscription calls to reconcile to zero.
Every Claude response reported that paid overage was not in use and was disabled at the organization level. Codex’s account credit balance was unchanged across the run, while ordinary subscription usage remained available. Gross additional metered charges, paid credits consumed and incremental cash were $0.00 each. Existing subscription purchase prices were excluded, and subscription allowances were used.
Cost per successful output was consequently $0 in additional metered charges for both clients. This is not a general API-cost tie or a claim that subscriptions are economically free. Anthropic’s model documentation and OpenAI’s model page provide the API tariffs. Recorded API-equivalent amounts are estimates, not invoices.
How to use the findings
For extraction and calculations, inspect schema validity, required types and exact values before focusing on prose. For code fixes, use the independent tests: generated JavaScript ran inside QuickJS WebAssembly with no exposed filesystem, network or environment, plus memory and time limits. The expected results stayed outside that sandbox.
For news and editing, preserving attribution and uncertainty matters more than a confident tone. The objective checks verify only their declared fields and constraints; they do not certify all prose as factual or publishable. No independent third-family judge was available through authenticated access, so subjective writing quality remains unscored in a blind review sheet. Neither contestant judged its own responses.
Luna’s observed turnaround makes its tested client worth considering when interactive latency matters. Use the category results and saved examples to assess strict output requirements; do not select either model solely from the aggregate percentage. A stronger follow-up would specify accepted enums and types explicitly before collecting a new, unseen task bank. This completed run is reproducible evidence of the tested clients and checks, with its limits kept visible.
Trending on Kingy
Keep reading with the stories getting the most attention now.
The Kingy Brief
Get The Kingy Brief.
AI changes, original tests and one practical thing to try. Fridays at 09:00 Vancouver time.
Free · Double opt-in · Unsubscribe anytime
Signup help and newsletter schedule
Signup form provided by Beehiiv. After submitting, check your inbox for "Confirm your subscription to The Kingy Brief" and open its confirmation link. Check Spam or Promotions if you cannot find it.
Fridays at 09:00 Vancouver time: source-checked AI changes, original tests and one practical thing to try. The weekly restart begins October 9, 2026. We skip a week when there is not enough verified material. Free. Unsubscribe anytime.
