AI News

Tested AI comparisons

The same tasks. The complete outputs. Dated evidence.

Choose an AI tool using the work it produced

These comparisons test specified API models on the same practical tasks. Every scored attempt will remain visible, including mistakes and failures. Coding and applied math are the first categories.

Measured results

Python coding tasks

12 tasks · 6 dated configurations · 3 repetitions

Last tested: 2026-10-07T23:20:49.408221+00:00

Invoice reconciliation, CSV parsing, deployment order, rate limits, Unicode handling and other function tasks.

Measured results

Applied math tasks

12 tasks · 6 dated configurations · 3 repetitions

Last tested: 2026-10-07T23:26:26.061111+00:00

Cash flow, probability, production constraints, loan payments, margins and statistics.

What has been tested

648 provider requests are recorded in the benchmark cohorts.

How updates work

Official sources and page health are checked on the host every 24 hours. Selected new configurations are tested once after review, within a US$20 monthly ceiling for expansion costs. Historical baseline costs remain under their existing allowances. Paid monthly full-roster reruns are disabled. Eligible launches target publication within 72 hours; older catch-up runs are labelled backfills. A source check does not change a test result. A run must fit its approved spending allowance.

Launch coverage

  • Qwen3.8-Omni-Flash · announced 2026-09-18 · Pending provider access and contract; no test yet
  • Grok 4.7 · announced 2026-09-21 · Pending provider access and contract; no test yet
  • GPT-6 Sol · announced 2026-09-22 · Full Standard run exceeds the monthly expansion allowance
  • GPT-6 Luna · announced 2026-09-22 · Tested on the v2 bank
  • Claude Opus 5.5 · announced 2026-09-22 · Tested on the v2 bank
  • Reflection Beam · announced 2026-10-05 · Restricted access or unverified pricing; no test yet
  • Mistral Large 4 · announced 2026-10-06 · Pending provider access and contract; no test yet
  • GLM-5.3-Flash · announced 2026-10-03 · Pending provider access and contract; no test yet

gpt-6-luna · Backfill · original deadline 2026-09-25T00:00:00+00:00 to 2026-09-26T00:00:00+00:00 (official API availability date only) · completed

claude-opus-5-5 · Backfill · original deadline 2026-09-25T00:00:00+00:00 to 2026-09-26T00:00:00+00:00 (official API availability date only) · completed

claude-haiku-5-5 · Launch target · original deadline 2026-10-10T00:00:00+00:00 to 2026-10-11T00:00:00+00:00 (official API availability date only) · completed

Human preferences

Writing style, logo aesthetics and visual editing quality will use separate blind audience votes. Technical checks cannot establish taste. No language model scores the quality of another model’s output. Inspect the audience gallery status.

Maintenance needs recovery: 3 checks failed on the latest attempt. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-07T22:38:59.132050+00:00.

Download the public task bank · Read the methodology