Trending on Kingy
Keep reading with the stories getting the most attention now.
Both apps passed all 12 fixed checks. That is the clearest result from our SignalDesk test: GPT-6 Astra and GPT-5.6 Sol each built a local dashboard that met the same functional bar, using the same prompt and fictional data.
Sol’s recorded time to first claimed completion was 14 minutes 41 seconds, against Astra’s 16 minutes 35 seconds. The 1 minute 54 second difference describes these two runs. It does not establish which model is generally faster.
A fresh isolated review preferred Astra’s finished app, scoring it 24/25 against Sol’s 22/25. Its clearest advantage was the narrow-screen comparison: both products could be read together. That preference is separate from the functional tie and Sol’s shorter recorded completion time.
The job: make the records useful
SignalDesk is a fictional launch-intelligence product. The input file contains 12 fictional records. The app needed to make those records searchable and filterable, open their details by keyboard, compare two items, preserve the user’s state across a reload, and remain usable on a 390-pixel-wide screen.
The constraints were deliberate: plain HTML, CSS and JavaScript; local data; no package installation, remote APIs, authentication or deployment. This was a local prototype test. We did not test a backend, production traffic or handling of customer information.
Both runs used recorded Extra High / xhigh, with a single agent. They received the same 1,887-byte prompt, product brief, dataset and comparison protocol. The model wrote the app; Codex supplied file and shell access; local browser tools exercised it. None of this isolates the model from its surrounding tools, and it does not test Ultra.
What actually passed
The independent browser records contain real fills, clicks, keyboard actions, reloads and viewport changes. They establish more than screenshots alone.
| Recorded result | Astra | Sol |
|---|---|---|
| Fixed checks passed | 12/12 | 12/12 |
| Desktop check viewport | 1440 × 900 | 1440 × 900 |
| Mobile check viewport | 390 × 844 | 390 × 844 |
| Console, page or failed-request errors in the independent sequence | 0 | 0 |
| Operator app edits | 0 | 0 |
| Post-completion correction prompts | 0 | 0 |
The checks covered the initial record count, a fixed search, selected filters and reset, a two-record comparison and third-item limit, keyboard detail interaction, persistence and a zero-results state. The two sizes and browser console were checked too. This was a fixed sample of behavior, not an exhaustive accessibility or cross-browser audit. It did not independently exercise every filter combination or every possible keyboard path.
How the finished interfaces compared
| Review dimension | Astra | Sol |
|---|---|---|
| Clarity | 5 | 4 |
| Craft | 4 | 4 |
| Correctness | 5 | 5 |
| Completeness | 5 | 5 |
| Usefulness | 5 | 4 |
| Total | 24/25 | 22/25 |
The new reviewer used four fresh Chrome contexts—one per candidate and viewport—and exercised the required interactions at 1440 × 900 and 390 × 844. No console errors or non-loopback page requests were observed. These are fresh review observations, separate from the historical fixed-check certification; the mobile test used desktop Chrome viewport emulation, not physical phone hardware.
In Astra’s mobile comparison, both selected products and all six requested fields were visible together. Sol kept the information available, but its 660-pixel-wide table sat inside a 362-pixel pane and required horizontal panning. Astra also retained category and status in its mobile records and placed the full filter set in the first mobile view. Sol’s compact list required more detail-opening to find that context.
That difference explains the clarity and usefulness scores. Both apps received 5s for correctness and completeness within the sampled requirements, and 4s for craft: their desktop introductions, metrics and filters left only one complete launch row visible initially. No extra-feature bonus was awarded.
| Astra mobile comparison | Sol mobile comparison |
|---|---|
![]() |
![]() |
These are new, matched review captures of the preserved apps. They are not frames from the original authoring process.
The replacement review used neutral copies, randomized labels and a fresh AI reviewer whose input contained only that packet. Its scorecard and report were locked before the mapping was read for this article. One AI review is a limited preference observation, not a human usability study or a measure of repeatability.
An earlier anonymous-label review scored Astra 24/25 and Sol 22/25. We preserved that record but superseded it because its context did not establish clean isolation from producer identities. Those historical scores are not the article’s blind result.
The build included repairs
“First claimed completion” includes the models’ own construction, checks and revisions. Astra’s record identifies a repair to dialog focus wrapping. Sol’s identifies repairs to a favicon request and comparison-checkbox updates. Neither required the evaluator’s permitted post-completion correction. These defect records are not counts of every edit made during authoring.
The tools also failed along the way. Astra recorded ten failed command executions. Sol recorded nine failed command executions and one failed MCP call. Sol’s Playwright wrapper made failed npm-registry lookups in three command executions, with five logged failure blocks; no download or installation completed. Calling the entire run free of network attempts would be inaccurate, even though the independent final browser checks recorded only loopback traffic.
Try the task
Copy the fictional product brief and 12-record JSON into a new directory, then submit the exact tested prompt. The prompt is plain text and includes all required behavior. It is suitable for a local environment with file access, a local HTTP server and browser inspection already available.
The beginner guide walks through the files. The advanced reproduction guide explains the fixed checks and correction budget. The delivered Astra app and Sol app must be served over loopback to load their local JSON; double-clicking HTML is not the tested setup.
The practical lesson is to define the interactions before asking for polish. Twelve passing checks tell you what worked in this case. Looking at both interfaces afterwards helps you decide which presentation you prefer without turning visual taste into a functional failure.
Limits attached to this result
This is one time-separated run per model on one brief. The completion intervals come from the retained run metadata; the app JSONL streams do not provide independent per-event timestamps. Neither interval is a repeated benchmark or a controlled inference-speed measurement.
Astra has no literal screen recording of its model/effort selection. Its evidence consists of recorded CLI settings, transcripts, artifacts and verification. Sol has a visible runner banner rather than a native selector recording; the raw footage includes unrelated foreground content and remains private. The new screenshots are review captures of preserved artifacts, not images of the original authoring process.
No publication, deployment, external-app write, successful installation or paid-service call is recorded. Account-wide allowance changes do not tell us either run’s cost. The fixed results say nothing about maintainability over months, security under hostile inputs, production reliability, native image/video output or a broad coding advantage.
For a different task with a different failure mode, read the research-to-spreadsheet comparison. The Field Guide hub keeps completed evidence separate from planned demonstrations.


