Kingy Research · Controlled AI video test
AI video models look very different when every failed attempt stays in the ledger. Kingy retained 30 attempts across Veo 3.1 Lite, Seedance 2.0 Fast and Kling 3.0 Standard, keeping historical cost, latency and unusable outputs in the comparison.

Commercial disclosure
Kingy ran this test through OpenRouter using existing account credit. The retained ledger records $9.0666 in provider charges, and no credit purchase or auto top-up occurred during the replication run. The retained records do not establish whether any pre-existing credit was purchased or promotional, so Kingy does not claim the test was entirely self-funded.
No vendor reviewed or approved this report, and no other material support is documented in the retained test records.
Controlled results
| Frozen model version | Attempts | Recorded cost | Mean latency | Latest agent-screened usable | Cost per agent-usable | Changed verdicts |
|---|---|---|---|---|---|---|
google/veo-3.1-lite-20260331 |
10 | $1.20 | 55.3s | 10/10 (100%) | $0.12 | 2 |
bytedance/seedance-2.0-fast-20260414 |
10 | $3.6666 | 332.9s | 3/10 (30%) | $1.2222 | 1 |
kwaivgi/kling-v3.0-std-20260429 |
10 | $4.20 | 189.3s | 4/10 (40%) | $1.05 | 1 |
Evidence limit: these are historical recorded costs, not current price quotes. The usability columns use the latest agent-assisted screen and must not be presented as independent human ground truth.
What failed
The most frequent retained labels were wrong camera move (7), product geometry change (5), motion flicker (5), missing camera move (5), and product text corruption (4). Background drift, registration-marker drift and general visual artifacts also appeared.
Counts are multi-label. One clip can carry several failures. There were no recorded execution failures and no technical retries; the observed defects were creative output failures.
How the test worked
- Freeze the brief. Each pack used one controlled prompt, four-second duration, 16:9 frame, 720p output and no generated audio.
- Retain every attempt. All 18 baseline and 12 replication generations remain in the attempt, cost and yield ledger.
- Separate transport from taste. A technical retry reconciles transport or polling; a creative reroll is another charged generation.
- Preserve disagreement. Changed verdicts remain visible instead of silently selecting one review.
Review disagreement
Two agent-assisted blind screens disagreed on four identical product clips. A later task-owner reviewer confirmed six blind decisions after receiving AI-generated recommendations: two usable and four unusable. After those decisions were frozen, the owner authorized opening the mapping. The human-confirmed, AI-assisted decisions matched the replication screen on all six disputed clips: Veo 2/2 usable, Kling 0/2 and Seedance 0/2. This reconciles only the disputed subset; it does not erase the earlier disagreement or create independent human ground truth.
What this test did not establish
- Face identity, wardrobe or cross-shot character continuity
- Hands, contact physics, speech or lip sync
- Runway or Wan generation quality
- Current model availability or current pricing
- A universal best AI video model
Runway and Wan are disabled, untested adapters. They are not hidden losers and must not appear in the ranking.
Public evidence boundary
The publication candidate includes aggregate counts, tested version labels, controlled prompts, non-sensitive settings, aggregate failures, clearly qualified agent-assisted rates and the authorized model-level six-clip reconciliation. It excludes raw clips, first frames, blind IDs, per-clip mappings, request/job/generation IDs, provider logs, private scorecards, reviewer identity and local paths.
Sources and method
- Sealed local attempt ledger with 30 completed generations
- Two retained agent-assisted blind scorecards
- Frozen human-confirmed, AI-assisted six-clip review form
- Kingy editorial and sponsorship standards
No current provider price or capability claim is made in this draft.