AI News

What Can Astra Do? A Field Guide With Local Evidence

Astra and Sol each produced a local web app and a research workbook with recorded final fixed checks of 12/12 for its task. This guide gives you the prompts, fictional inputs, finished files, and checks behind those historical results, alongside limitations found in later review.

Start with the outcome you want. The beginner prompt guide explains how to prepare the files and includes copyable prompts. The advanced reproduction guide covers the run conditions, fixed checks, corrections, and limits.

What the evidence labels mean

Label What it means here
Observed A retained local artifact or execution record supports the statement. The scope of its checks still matters.
Saved-documentation A statement comes from official pages saved for the research runs. It describes those source snapshots, not a fresh check of today’s product.
Untested This delivery has no completed local test for the stated workflow or adaptation.
Planned A defined demonstration exists, but no model result is available.

The models worked through Codex and its surrounding tools. Local file and shell tools wrote the app files; a browser exercised their interactions. The research workflow used official-page retrieval and a spreadsheet runtime to build, inspect, and render XLSX files. Those tools contributed to the outcomes.

Beginner route: choose a finished example

Build a local app from a fictional brief

Observed. Both models received the same SignalDesk brief, 12 fictional launch records, and exact task prompt. The app had to support search, filters, details, two-record comparison, saved state, and a narrow mobile layout.

Both final apps have recorded results of 12/12 on the independent fixed checks. Neither received a correction prompt after claimed completion. Both made changes while authoring, so this does not mean either app emerged in one uninterrupted generation.

Read the working-app comparison for the final-artifact review and its limits. You can also inspect the Astra app files and Sol app files. Serve each through a local HTTP server using the reproduction instructions; opening the HTML as a file may prevent its data from loading.

The practical lesson is to put interactions in the brief and check them in a browser. A screenshot cannot establish that filters work or state survives a reload.

Turn documented sources into an auditable workbook

Observed. Each model produced a five-sheet XLSX with a model comparison, capability matrix, source table, and formula-driven calculator using fictional workload inputs.

Recorded checkpoint Astra Sol
Initial fixed checks 12/12 10/12, with an audit inconsistency
Post-completion correction prompts 0 1
Final fixed checks 12/12 12/12

Sol’s original record passed check #12 despite a source defect that also conflicted with that check. The recorded 10/12 is preserved with this caveat. A later in-memory review found that Sol blanks an affected total when a base rate is removed but does not dynamically identify the missing rate; Astra displays “Missing API rate.” Treat 12/12 as the recorded historical result. The research comparison explains the correction and this additional limitation.

Open the Astra workbook or Sol workbook, then follow a calculator rate back to its source. The figures are saved-documentation scenarios using fictional volumes. They are not bills or a current pricing recommendation. Preserved verification tested recalculation through a spreadsheet runtime; it does not establish native Excel interaction.

The practical lesson is to require a visible gap when an input or rate is unsupported. A complete-looking number can conceal a missing assumption.

Advanced route: inspect the conditions

The two comparisons used recorded gpt-6-astra and gpt-5.6-sol runs at Extra High (xhigh), with a single agent. Each pair used byte-identical supplied prompts and inputs. The reproduction guide explains how to separate authoring changes, post-completion corrections, operator verification, and presentation review.

These are individual case studies. They do not establish a general model ranking, repeated-run reliability, or real per-task cost. The app and workbook checks measure different artifacts; their scores should not be pooled. Preference findings belong with the method and limitations in each comparison article.

Motion remains planned

Planned. The eight-second SignalDesk motion brief and exact planned prompt are ready to inspect. Existing HyperFrames, Chromium, Node, and FFmpeg executables passed version probes, but a complete composition render remains unverified.

The visible-recording preflight stopped before either model received the task. The probe captured an unrelated foreground window rather than the required run evidence. No authoring turn or candidate exists, so there is no motion result to score. See the motion demonstration plan.

Image generation or editing, a dedicated computer-use test, external-app writes, and Ultra orchestration remain untested in this field guide. The completed app and workbook cases provide no hands-on proof for those workflows, and no native media-generation claim follows from the planned HyperFrames demonstration.

Choose the beginner guide to try a prompt, or the advanced guide to inspect the evidence before running anything.