# Kingy Production Benchmarks · bottle ads v1

Preregistered September 23, 2026. Protocol ID `kpb-bottle-v1`. Versioned fixture, code and source hashes accompany the release. No confirmatory paid trials have run under this protocol.

## Decision and scope

Choose a workflow to turn one supplied, approved bottle packshot and approved copy into a 15-second, sound-off, 9:16 product ad. The buyer is a small ecommerce creative agency's production lead, choosing the repeatable method for monthly client creative. This is production acceptance, not an experiment on click-through, sales, ROAS or platform approval.

The free feasibility cohort uses the existing fictional Kingy Field Bottle. It is an AI-generated test object, not a photographed commercial product. Three representative briefs cover introduction, product detail and retargeting. They test different copy and creative goals for ONE SKU. Generalization across packaging, industries, operators or ad effectiveness is not supported. A paid client cohort substitutes three rights-cleared client SKUs before preregistration.

## Design

Three workflows × three briefs × three independent production sessions = 27 final-deliverable slots, nine per workflow. Workflow 1 uses the original packshot and a conventional template edit. Workflow 2 uses Runway `gen4.5` through Runway Dev API, then the same edit. Workflow 3 uses Google `veo3.1_fast` through Runway Dev API with audio off, then the same edit. These are complete specified workflows, not tests of all features in either provider. The shared gateway controls billing surface; results cannot establish direct Google API or vendor reliability comparisons. These established routes are a bounded initial shortlist, not a claim that they are the latest or best models. Newer Gemini Omni Flash is a named next-cohort candidate because Google's current docs recommend it as a default.

Each final ad contains two five-second visual sections and a five-second original-packshot end card. Models supply visuals only. Exact approved copy is composited conventionally in every workflow. Sound is deliberately excluded because the job is a sound-off placement. The baseline receives the same packshot, copy, output specification and active-time limit; a human may use the provided template or an ordinary editor but records the tool and every change. Baseline repeats measure operator/session variance; deterministic template copies are not independent creative evidence.

For paid routes, initiate shot 1 then shot 2. Each shot has one initial attempt and at most one retry, with identical prompt/settings and a different preregistered seed. First candidate that passes the operator's provisional gates proceeds; a retry requires a recorded rejection or terminal provider failure. No best-of-many selection. No prompt tuning within the scored cohort. The operator may trim the first contiguous five seconds only; no cherry-picked interval. One conventional edit revision is allowed within 30 active operator minutes per deliverable, including prompting, review and editing. Stop when the time or attempt cap is reached and record failure. Final blinded review cannot trigger extra generation. Baseline sessions have the same time cap.

Use the seeded shuffled session order in `schedule.json`; do not finish one provider before starting the other. Runs use seeds 101, 202 and 303 plus fixed shot/retry offsets. Record operator, date, account region, product surface, API version, returned model version if available, native billable duration, input hashes, seed, exact prompt and settings. When a provider hides a model snapshot, say so; a family ID is not an immutable model version. Price verification must be less than 24 hours old at submission. Do not substitute a route silently after access failure.

## Acceptance fixed before testing

All objective gates must pass: MP4/H.264; 720 × 1280; 24 fps; duration 15 ± 0.1 seconds; decodes without errors; silent. Inspect full playback and frames at every second. Exact hook appears by second 1 and remains through second 4; exact detail occupies seconds 5–9; exact CTA occupies seconds 10–14. Essential text stays inside x=72…648, y=140…1090 and is readable at 360px display width. The safe area is this test's fixed convention, not a promise about every ad platform's overlays.

Human product/content gates: complete cap, bottle and base visible in both hero sections; exact protected label when visible; no product geometry drift, extra product, invented claim, watermark or brand error; all approved copy exact. Any such blocking defect rejects the ad regardless of aesthetic score. A rejected shot cannot be concealed behind the end card or cropped away.

Two independent human reviewers, not the operator, review anonymized exports in independently randomized order. They see brief, source image and output, but not model, provider, cost, retries or production time. The blind key is private. They score product fidelity, copy legibility, motion stability and brief fit on 1–5: 1 unusable; 2 major repair; 3 visible correction needed; 4 acceptable with no further work; 5 clean. Every score must be ≥4 and every gate pass for acceptance by a reviewer. Both must accept; disagreement goes to a third blind reviewer whose decision and explanation are retained. Report disagreement count and adjudications. Style may reveal workflow; blinding is imperfect and stated. No LLM score substitutes for a human acceptance decision.

Automated dimensions/codec/duration checks are objective; OCR may flag copy differences but does not establish readability or product fidelity. Code records those separately. Operator preview is provisional, not acceptance. Missing review, raw output, human time or cost is missing evidence, never zero.

## What to measure

For each initiated request: terminal outcome, provider/request ID, submission/completion time, billable units, receipt evidence, estimated charge, observed generation consumption, tax/fee evidence and output checksum. Provider errors and safety blocks count as initiated requests if submitted. Authentication/preflight errors do not. Keep every output and discarded candidate. Distinguish generation retry, conventional edit revision and session rerun.

For each slot: accepted/rejected/pending, rejection reasons, selected attempt IDs, active prompting/review/edit/export minutes, queue/wall time, technical results and human scoring. Record active human time with start/stop logs or a contemporaneous timer, separately from automated render time and agent elapsed time. Count failed work and retry work in the numerator. Do not bill machine waiting as active human editing.

Per-workflow production cost per accepted deliverable = (all observed generation consumption for that cohort + directly attributable software/asset allocation + active operator minutes × stated hourly rate/60) / accepted final deliverables. Benchmark-only blinded-review labor and research/setup overhead are shown separately. The monetary value of human time is always a labeled rate scenario, unless an actual invoice supports it. Never call that a paid cash cost. Publish generation-only consumption per accepted deliverable separately. With zero accepted outputs, cost per accepted deliverable is unavailable, not $0. With missing costs or time, the fully loaded value remains unavailable. Report new cash purchases, unused credits and subscription commitments separately so they are not double-counted with consumed credits.

Report acceptance count/n, rejected/n, unresolved/n, retry count and failure classes; median and range of minutes; aggregate costs and per-brief cells. Also show a Wilson interval for final acceptance where there are completed decisions, with a warning that three briefs on one SKU are clustered and the interval is descriptive, not a population guarantee. Nine sessions/workflow cannot establish a stable universal winner. Do not use significance claims. Offer a conditional recommendation only if data are complete: workflows meeting the acceptance gates, observed cost at buyer's rate, and remaining uncertainty. No composite magic score.

## Retention, corrections and independence

Retain full-resolution originals, final exports, exact requests/responses, receipts, time logs, review forms, blind key, SQLite ledger and SHA-256 manifest. Download temporary provider outputs immediately. Publish a redacted reproducibility bundle with all non-private evidence. Client originals remain private unless publication rights are explicitly granted; publish hashes and explain access limits. Keep the raw cohort and source snapshot for at least 12 months after the report's final public date; no automatic deletion is configured. Inquiry handling is separate.

Freeze the protocol and input hash in every run. Corrections create a new report revision and append a reason; never overwrite raw records or silently rewrite costs. Regenerate JSON, CSV, tables and download bundle from the same ledger; list affected recommendations. Major model change, prompt/edit template change, new input class, billing-surface change or acceptance rule change starts a new cohort. Price-only changes get a dated estimate overlay and do not rewrite historical consumption. Failed download or undocumented model switch blocks publication of a complete ranking.

Kingy accepts no payment to improve a score. Testing fees purchase the work and readout; providers cannot veto failures. Disclose supplied credits, affiliate relationships and any sponsorship for each cohort. Kingy's broader media business sells sponsorships, so readers should inspect the separate disclosure. This launch has no new provider-funded testing and no paying pilot customers. Historical records do not establish independent human acceptance.
