AI News

AI Video Test Lab: 30 attempts across Veo, Kling and Seedance

Kingy Research · Controlled AI video test

AI video models look very different when every failed attempt stays in the ledger. Kingy retained 30 attempts across Veo 3.1 Lite, Seedance 2.0 Fast and Kling 3.0 Standard, keeping historical cost, latency and unusable outputs in the comparison.

AI Video Test Lab editorial graphic showing three controlled model-test lanes and thirty retained attempt markers.
Thirty charged attempts. Three exact model routes. No failed generation removed from the ledger.

Commercial disclosure

Kingy ran this test through OpenRouter using existing account credit. The retained ledger records $9.0666 in provider charges, and no credit purchase or auto top-up occurred during the replication run. The retained records do not establish whether any pre-existing credit was purchased or promotional, so Kingy does not claim the test was entirely self-funded.

No vendor reviewed or approved this report, and no other material support is documented in the retained test records.

Controlled results

Frozen model version Attempts Recorded cost Mean latency Latest agent-screened usable Cost per agent-usable Changed verdicts
google/veo-3.1-lite-20260331 10 $1.20 55.3s 10/10 (100%) $0.12 2
bytedance/seedance-2.0-fast-20260414 10 $3.6666 332.9s 3/10 (30%) $1.2222 1
kwaivgi/kling-v3.0-std-20260429 10 $4.20 189.3s 4/10 (40%) $1.05 1

Evidence limit: these are historical recorded costs, not current price quotes. The usability columns use the latest agent-assisted screen and must not be presented as independent human ground truth.

What failed

The most frequent retained labels were wrong camera move (7), product geometry change (5), motion flicker (5), missing camera move (5), and product text corruption (4). Background drift, registration-marker drift and general visual artifacts also appeared.

Counts are multi-label. One clip can carry several failures. There were no recorded execution failures and no technical retries; the observed defects were creative output failures.

How the test worked

  1. Freeze the brief. Each pack used one controlled prompt, four-second duration, 16:9 frame, 720p output and no generated audio.
  2. Retain every attempt. All 18 baseline and 12 replication generations remain in the attempt, cost and yield ledger.
  3. Separate transport from taste. A technical retry reconciles transport or polling; a creative reroll is another charged generation.
  4. Preserve disagreement. Changed verdicts remain visible instead of silently selecting one review.

Review disagreement

Two agent-assisted blind screens disagreed on four identical product clips. A later task-owner reviewer confirmed six blind decisions after receiving AI-generated recommendations: two usable and four unusable. After those decisions were frozen, the owner authorized opening the mapping. The human-confirmed, AI-assisted decisions matched the replication screen on all six disputed clips: Veo 2/2 usable, Kling 0/2 and Seedance 0/2. This reconciles only the disputed subset; it does not erase the earlier disagreement or create independent human ground truth.

What this test did not establish

  • Face identity, wardrobe or cross-shot character continuity
  • Hands, contact physics, speech or lip sync
  • Runway or Wan generation quality
  • Current model availability or current pricing
  • A universal best AI video model

Runway and Wan are disabled, untested adapters. They are not hidden losers and must not appear in the ranking.

Public evidence boundary

The publication candidate includes aggregate counts, tested version labels, controlled prompts, non-sensitive settings, aggregate failures, clearly qualified agent-assisted rates and the authorized model-level six-clip reconciliation. It excludes raw clips, first frames, blind IDs, per-clip mappings, request/job/generation IDs, provider logs, private scorecards, reviewer identity and local paths.

Sources and method

No current provider price or capability claim is made in this draft.