AI News

Tested AI comparisons

The same tasks. The complete outputs. Dated evidence.

Choose an AI tool using the work it produced

These comparisons test specified API models on the same practical tasks. Every scored attempt will remain visible, including mistakes and failures. Coding and applied math are the first categories.

Measured results

Python coding tasks

12 tasks · 3 API models · 3 repetitions

Last tested: 2026-10-04T17:01:11.203235+00:00

Invoice reconciliation, CSV parsing, deployment order, rate limits, Unicode handling and other function tasks.

Measured results

Applied math tasks

12 tasks · 3 API models · 3 repetitions

Last tested: 2026-10-04T17:01:24.616847+00:00

Cash flow, probability, production constraints, loan payments, margins and statistics.

What has been tested

216 provider requests are recorded in the benchmark cohorts.

How updates work

Official sources and page health are checked on the host every 24 hours. Complete tests are due monthly and after a reviewed major launch affecting the selected tools. A source check does not change a test result. A run must fit its approved spending allowance.

Paid reruns are currently held. Daily source and page checks continue; existing test dates remain unchanged.

Human preferences

Writing style, logo aesthetics and visual editing quality will use separate blind audience votes. Technical checks cannot establish taste. No language model scores the quality of another model’s output. Inspect the audience gallery status.

Maintenance needs recovery: The daily source check is overdue. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-04T09:01:57.926856+00:00.

Download the public task bank · Read the methodology