AI News

Astra vs Sol: A Spreadsheet That Has to Show Its Sources

A spreadsheet can look finished while one unsupported assumption sits inside it. Our research-to-spreadsheet test exposed that problem: Astra’s first completed workbook passed the recorded 12 checks; Sol’s needed its one permitted correction before reaching the same recorded final check count. A fresh review later found a narrower missing-rate diagnostic limitation in Sol’s final workbook.

In that fresh isolated review, Astra’s final workbook scored 22/25 and Sol’s 20/25. The preference came from how clearly the workbook explains a missing rate. That reader-facing result is separate from the first-completion check count and from elapsed time.

This article covers one matched task. It evaluates the produced workbooks and their local evidence. It is not a current pricing guide, a real spending comparison or a general ranking of research ability.

One brief, five sheets

Both models received the same 1,784-byte prompt and research brief, ran at xhigh with a single agent, and used Codex CLI 0.153.1 with the same bundled spreadsheet runtime, version 26.903.11726. The requested workbook compared the two products using permitted official OpenAI documentation and four fictional monthly workload inputs: 3 million uncached input tokens, 2 million cached input, 500,000 cache-write and 750,000 output tokens.

The five required sheets were Read me, Model comparison, Capability matrix, Cost calculator and Sources. Numeric inputs had to remain editable. Claims needed source URLs and access dates. A calculation with missing or incompatible definitions had to remain blank with an explanation.

The model selected and organized source material and authored formulas. Codex and the bundled spreadsheet tool created, inspected and rendered the local XLSX. A populated workbook is evidence of that tool-assisted workflow; it does not show that every documented model capability was exercised.

First completion and final completion

Recorded measure Astra Sol
Initial fixed checks 12/12 10/12, with audit caveat below
Post-completion correction passes 0 1
Recorded final fixed checks 12/12 12/12
Launch to first final response 15m 00.6s 11m 39.1s
Separate correction turn 5m 30.6s

Sol’s initial check record fails a source check and a presentation-evidence check. It nevertheless passes scope check #12, whose wording also prohibits unsupported values. That is an inconsistency in the evaluator’s initial total. We preserve 10/12 as recorded, rather than silently inventing a revised score. The source error and the need for correction are supported independently of that counting issue.

Both authoring runs performed self-checks and refinements before their first final response. Sol also repaired clipped titles before initial completion. “First completion” therefore means before an evaluator-requested correction, not a single unedited generation. The displayed times exclude preflight and operator verification, and the correction time is not an end-to-end turnaround total.

The assumption that needed correction

Sol’s initial workbook inferred a 2× long-context cache-rate multiplier that its cited source set did not explicitly define. The permitted generic correction prompt made it recheck the brief. It replaced the unsupported multiplier with an unavailable label, separated input and cache-rate multipliers, and clarified the explanation.

The wider saved official-source collection did contain explicit long-context cache rates: Astra had followed a linked pricing page that Sol had not retained. The finding is therefore “unsupported by Sol’s cited sources,” not “OpenAI published no rate.” The runs retrieved different source sets at different times. That limits matching; it does not prove that the documentation changed between them.

Sol’s initial calculator render also left the last wrapped line ambiguous at its auto-crop boundary. An operator later produced an explicit-range proof render without editing the workbook. That is verification evidence, not an extra authoring improvement by Sol. Its adjusted monthly estimates were already blank before the correction; the correction did not newly suppress a previously populated estimate.

What the calculator really demonstrated

In the saved all-short-context scenarios, both workbooks produced $75.75 for Astra and $30.30 for Sol using the fictional volumes and recorded rates. Changing uncached input from 3 million to 4 million changed those scenario totals to $85.75 and $34.30. These are historical example calculations, not bills or newly verified prices.

The preserved operator scripts imported the delivered XLSX files, changed values in memory, checked recalculation, removed a required rate, and confirmed that the affected result became blank while the other model’s result remained available. They did not save edited workbooks. That is actual runtime behavior evidence, separate from visual inspection.

Neither monthly token totals nor a polished table reveal how tokens are distributed across individual short- and long-context requests. Both final workbooks explain why a mixed-workload monthly estimate cannot be determined from the supplied inputs. Astra additionally provides an all-long-context scenario; its broader source set and extra formulas are descriptive differences, not bonus benchmark points.

The finished-workbook review

Review dimension Astra Sol
Clarity 4 4
Craft 4 4
Correctness 5 4
Completeness 5 4
Usefulness 4 4
Total 22/25 20/25

The reviewer preferred Astra’s explanation of calculation boundaries. In a new in-memory check, removing a required base cache-write rate blanked the affected total in both workbooks. Astra also displayed Missing API rate. Sol retained notes about long-context allocation and base assumptions, but none identified the missing base rate; the linked rate area displayed zero in the runtime while the total stayed blank.

That diagnostic gap limits Sol’s correctness and completeness scores for this review. It does not overwrite the old evaluator’s final 12/12 record or establish how native Excel would recalculate the edit. The final workbook was left unchanged; its single correction allowance had already been used.

The reviewer found both workbooks strong but gave clarity, craft and usefulness 4s: Sol was compact, while Astra’s long tables and repeated URLs required more scrolling. Astra received no bonus simply for more sources, more formulas or extra scenarios.

The fresh AI reviewer received randomized neutral copies and a pooled local source collection, without the producer mapping or prior review history. Workbook subject matter still names both products. ZIP container timestamps were neutralized; every uncompressed workbook part remained byte-identical. Scores were locked before the mapping was revealed.

The review imported both delivered XLSX files with the bundled artifact runtime, inspected all five sheets, viewed all ten supplied renders at original scale with readable pixel crops, and tested calculator edits in memory. It did not use native Excel or an ordinary-zoom workbook viewer. Saved workbook hashes remained unchanged. This provides runtime and raster evidence, not a native spreadsheet usability study.

The earlier anonymous-label review gave Astra 22/25 and Sol 24/25. It had context exposure and inspected renders rather than opening the workbooks as instructed. We retain it as a historical record and use the replacement review here. A render-only review must not be described as interactive workbook testing.

Reuse the useful part

Start with the exact historical prompt and fictional brief. That prompt requests current official sources, so running it unchanged requires permission for external retrieval. This package did not rerun that research. The beginner guide provides a separately labelled offline adaptation using saved snapshots; it is a suggested new exercise, not the measured prompt.

You can inspect the preserved Astra workbook, Sol workbook and local sheet gallery. The most useful requirement to borrow is simple: if a rate, definition or input is missing, show the gap beside the blank result. Then test that behavior by removing a value deliberately.

What remains uncertain

This is one time-separated cell with different retrieval paths: eight preserved official pages or sections for Astra and four for Sol. They were not authored from a common frozen corpus. Sol’s original first-completion XLSX was not retained separately; its initial hash and inspection/calculator records remain. Initial behavior claims rely on those records.

The fresh review is one AI judgment per pair, with the display/runtime limits stated above. Formula counts, source counts and token telemetry do not establish efficiency or cost. No Ultra experiment was run. A nonzero Astra assertion command and a recovered Sol retrieval error also mean the process was not error-free.

Both runs retain pre-submission recordings, but unrelated foreground windows make the raw footage private. Pinned CLI launch records establish the worker model and effort; captured GUI selectors belonged to the operator’s task. No paid API call, installation, external-app write, publication or deployment is recorded. All source claims here refer to saved September 2026 snapshots, not a fresh check of current availability or prices.

The working-app comparison tests a different deliverable. The Field Guide hub links the two without adding their scores into a universal winner.