AI News

AI comparison testing methodology

How the comparisons are measured

The initial study compares three direct API configurations on 12 Python coding tasks and 12 applied math tasks. Each model receives the same task prompt three times in separate requests. This is a comparison of those configurations on a small declared bank, with descriptive results and limited coverage.

Original v1 task selection and scoring

The coding bank covers data correctness, algorithms and edge cases. It measures Python function implementations, with a stated import list and execution limits. It does not measure IDE assistance or whole-repository agents. The math bank covers finance, statistics, probability and constrained optimization. Its numerical answers are checked against frozen reference values; explanation quality is not scored.

A coding attempt passes only if every frozen fixture passes. Math outputs must supply the requested JSON fields and satisfy an absolute tolerance of 0.000001 or relative tolerance of 0.00000001. Equivalent numeric strings are accepted. All 12 tasks have equal weight; all repetitions count. No best-output selection, prompt tuning during the cohort or automatic paid retry is allowed.

Comparing v1 and v2 benchmark results

V2 results are published from a separate, complete matched cohort. The original v1 study remains available in the archive.

Our first coding and math results use the original v1 task bank. V2 changes the prompts and scoring requirements and appears as a separate benchmark version. A score change between v1 and v2 alone cannot establish that a model improved or declined because the test changed too.

What changes in v2

Both versions contain 12 coding tasks and 12 applied math tasks. Each model receives every task three times in separate requests. Each attempt passes only when it satisfies every check for that task. All tasks have equal weight, and every repetition counts.

Original v1 and reviewed v2 requirements
AreaOriginal v1Reviewed v2
Coding instructionsSome execution restrictions and input assumptions were incompletely stated.Prompts list the available module functions, require input preservation and clarify CSV handling and dependency ordering.
Coding checks48 cases across 12 tasks.62 cases across the same 12 tasks, with added edge cases.
Output formatThe grader accepted whole Markdown fences and equivalent numeric strings in math, despite stricter prompt wording. Requested coding dictionary order was not enforced.Code and math answers must follow the stated format. Math requires one JSON object with unique requested fields and numeric values. Ordered coding dictionaries are checked where requested.
Math precisionPrompts did not disclose the scoring tolerance or clearly prohibit final rounding.Prompts distinguish exact counts and cents from values checked with tolerance, and explain the required precision. A small reliability-reference error is also corrected.

Each complete v2 cohort contains 36 scored attempts per model in each category: 12 tasks multiplied by three repetitions. The extra coding cases increase coverage within tasks; they do not give those tasks more weight. In v2, numerical scoring uses exact equality for declared discrete fields and, for continuous fields, an absolute tolerance of 0.000001 or relative tolerance of 0.00000001. We use deterministic checks and no LLM judge.

How we compare runs

We rank models within a complete run that uses the same task bank and grading rules. We keep v1 and v2 tables separate and identify each run by its bank version, grader fingerprint, exact model identity, settings and test date. Archived v1 coding · Archived v1 math.

Across versions, we can describe which tasks passed, what failed and how the requirements changed. We will not present the difference between a v1 score and a v2 score as a measured performance gain. Even a task with the same name may have clearer instructions or additional checks in v2.

Runs within the same version can show changes over time when the definitions remain unchanged and model identities and settings are disclosed. A claim about model improvement requires a matched comparison under the same test conditions. Three repetitions provide limited evidence for close rankings, so we report ties and task-level tradeoffs without claiming statistical superiority.

Cost figures belong to their dated runs. We include failed attempts, label usage-based estimates and keep billed reconciliation separate. Changed API prices or a changed task bank can affect cost per successful attempt; current prices never replace historical costs.

How we preserve the original evidence

The October 4 audit fixed a runner defect: the original coding instructions allowed built-in types, but the runner omitted bytes. Replaying the same retained answers corrected three Claude dependency-order verdicts, changing its coding display from 27/36 to 30/36. This was a scoring correction using existing answers. The other coding scores and all math scores stayed unchanged.

Original responses, grades, costs, test dates and JSON/CSV downloads remain intact. The correction is recorded separately and marked on affected attempts. The original CSV and rounding ambiguities remain disclosed. We have not applied v2 prompts, stricter format rules or extra cases to old outputs.

These comparisons measure bounded Python function implementation and applied quantitative answers. They do not establish performance on whole repositories, IDE assistance or mathematical explanation quality. The coding results and math results show the retained outputs so readers can inspect the evidence behind each score.

Matched conditions

Providers use their own medium reasoning settings and the same 8,192-token output cap. Similar setting names do not establish equal internal compute. Processing is identified on every dated row. Optional Flex uses a separate contract; it is not silently compared with Standard latency or pricing. No automatic paid retry or Standard fallback is allowed. There is no browsing, agent tooling, explicit prompt caching, batch discount or regional premium. Tasks are shuffled and provider order rotates across repetitions. Exact requests, returned identities and usage are retained.

Evidence and uncertainty

Original responses, checksums, usage and deterministic grades remain in an append-only run history. A complete category cohort contains 108 trials. Unresolved submissions, undocumented identity changes, missing evidence or grader infrastructure failures prevent a complete ranking. A provider’s documented availability is separate from verified account access.

Scores describe this task bank. We report ties and task-level tradeoffs, and make no claim of statistical superiority or universal ability from three repetitions. New task or grading rules create a new bank version and need approval. Changed model identities start a new cohort.

Cost accounting

All initiated attempts, including failures, contribute to cost. Estimates based on metered usage and reviewed rates are labeled; billed reconciliation is kept separately. Taxes, prepaid credit purchases, unused credit and subscription commitments are not silently converted into a task price. Cost per successful attempt is unavailable when none succeeds. Current-price estimates never overwrite historical run costs.

Continuous maintenance

Hosted checks record last attempt and last successful source fetch separately. Changed source text enters review; it is not automatically a verified launch or price correction. Paid monthly reruns are disabled; selected launch/backfill cohorts run only within the US$20 monthly ceiling and finite existing programme. If a check or release fails, the last verified publication remains available with its original test date.

Subjective categories

Essays, logos and video editing require a separately approved audience-voting protocol. Tool identities and identifying metadata stay private during voting. Presentation is randomized; required watermarks and other blinding limitations are disclosed. Preference rankings require actual participation meeting predeclared closing rules. Old votes cannot score new outputs.

Independence and disclosure

No provider-funded testing, affiliate influence or supplied credits have been recorded for this cohort. Actual future relationships must be declared before publication. Kingy’s wider business includes commercial sponsorships. Any connection to a tested provider must be declared for that run. Payment cannot improve a score or remove a failed output.

Bank fingerprint: fce29b4d66cfddeab39632d451935271a37df595374b40eacd300201e3726291
Protocol fingerprint: 90b114dd8a70ecbf3aecc35f58cea1ba85c786554aee9835a38a1c518293da05

Maintenance needs recovery: 3 checks failed on the latest attempt. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-07T22:38:59.132050+00:00.

Download the public task bank · Read the methodology