Google’s Gemini 3.8 Flash and Meta’s Muse Spark 1.3 make a similar pitch: strong reasoning and agent work without frontier-model pricing. Kingy’s direct test produced a clear, qualified result. Across the 56 fixtures completed by both models, Muse passed 112/112 scored calls and Gemini passed 91/112. That is an 18.75-percentage-point Muse advantage in the available-case analysis, with a fixture-clustered bootstrap 95% confidence interval of 9.82 to 28.57 points. [OUR TEST]
Limitation beside the headline result: this is an available-case analysis, not the preregistered all-240 verdict. Eight planned Gemini 900K long-context calls were not completed because the API hit a documented two-million-input-token-per-minute quota. Those eight runs are quota-limited execution evidence, not failed quality scores, and they are excluded from every quality denominator. The formal preregistered verdict is therefore insufficient evidence.
Within the completed test set, Muse is the stronger default when correctness is the priority. Gemini remains cheaper on completed scored calls and faster at the median. The largest quality separation occurred in a controlled synthetic long-context task, so this result should not be stretched into a universal claim about natural documents or every agent workflow.
This is a buyer’s guide, not a vendor scorecard. Product and pricing evidence was checked through September 3, 2026, 12:32 AM PDT (UTC−07:00); the benchmark was sealed on September 3, 2026. A label beside every consequential claim identifies its provenance. [OFFICIAL—GOOGLE], [OFFICIAL—META] and [OFFICIAL—OTHER MODEL DEVELOPER] are developer claims, not independent validation. [BENCHMARK OWNER] comes from the organization operating the benchmark. [INDEPENDENT] is not officially confirmed and could be incorrect. [OUR TEST] is Kingy’s own controlled observation and may not generalize. [THIRD-PARTY—UNVERIFIED] means Third-party/unverified — may be incorrect. [UNKNOWN] marks a gap we could not responsibly fill.
1. The short answer
Choose Muse Spark 1.3 for the tested workload mix when accepted-task correctness matters most. Choose Gemini 3.8 Flash when lower current Standard-route cost, lower median latency, Google’s tool ecosystem or broader documented modality support matters more. [OUR TEST] [OFFICIAL—GOOGLE] [OFFICIAL—META]
The completed benchmark does not justify calling either model universally better. Muse’s advantage was concentrated in long-context retrieval: it passed all 16 completed 128K and 512K calls, while Gemini passed none. Repository repair was nearly tied at 24/24 for Muse and 23/24 for Gemini. Both models passed every stateful tool-use run. Muse went 24/24 on image/PDF extraction and grounded decision analysis; Gemini went 22/24 in each. [OUR TEST]
The economic counterweight is real. Completed scored calls cost an estimated $8.3232 for Gemini and $16.7825 for Muse. Gemini’s median latency was 12.48 seconds versus 21.15 seconds for Muse, although Gemini’s p95 was much worse at 271.05 seconds versus 73.28 seconds. Costs use provider-reported token categories and published Standard list prices with no cache discount; they are not reconciled invoices. [OUR TEST]
2. What exactly launched?
gemini-3.8-flash is a stable, generally available Google Gemini API model. Google’s release positions it as a more deliberate Flash model that spends additional reasoning tokens when a task warrants it, while retaining the Flash tier’s speed and price orientation. It is also exposed through products including AI Studio, Gemini Enterprise, Android Studio and Google AI surfaces. [OFFICIAL—GOOGLE] Google model documentation
muse-spark-1.3 is Meta’s new hosted model in Muse Code and the Meta Model API. Meta describes it as a major agentic upgrade: better at maintaining long plans, using messy or conflicting context, correcting missing plan steps, handling multiple workflows and asking for clarification before consequential actions. Meta also says an open-weights release will come later. At the cutoff, Muse Spark 1.3 itself was not an open-weight release. [OFFICIAL—META] Meta launch announcement
This distinction matters. “Available through an API” is not “open source,” and “open weights coming” is not downloadable weights today.
3. Core specification comparison
| Attribute | Gemini 3.8 Flash | Muse Spark 1.3 | Evidence |
|---|---|---|---|
| Exact hosted ID | gemini-3.8-flash |
muse-spark-1.3; muse-spark-1.3-contributor |
[OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Availability | Stable / GA | Meta Model API and Muse Code | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Input modalities | Text, image, video, audio, PDF | Text, image, video, audio*, PDF | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Output modality | Text | Text | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Context window | 1,048,576 input tokens | 1,048,576 tokens | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Maximum output | 65,536 tokens | Not publicly specified in the reviewed model table | [OFFICIAL—GOOGLE] [UNKNOWN] |
| Reasoning controls | Low, medium, high; minimal errors |
Minimal, low, medium, high, xhigh |
[OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Tools | Functions, code execution, search, Maps, file search, URL context, computer use preview | Responses, Chat Completions and Messages endpoints; tool/agent workflows | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Knowledge cutoff | Primarily March 2026; some domains January 2025 | Not disclosed in the materials reviewed | [OFFICIAL—GOOGLE] [UNKNOWN] |
| Training on paid/Standard traffic | Paid-service content not used to improve Google products | Standard prompts/completions not used to train models | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Important caveat | Higher reasoning can increase latency and tokens | Audio on 1.3 is not fully supported; quality may degrade | [OFFICIAL—GOOGLE] [OFFICIAL—META] |
Neither developer publishes a parameter count. A million-token input limit also does not prove a model can use every part of a million-token prompt equally well. Context length is a capacity specification; retrieval accuracy and agent memory are separate measurements.
4. Official benchmark claims—and who ran them
The vendors emphasize different scorecards, so a row-by-row “winner” would manufacture comparability that does not exist.
| Evaluation | Gemini 3.8 Flash | Muse Spark 1.3 | Provenance and comparability |
|---|---|---|---|
| DeepSWE v1.1 | 73.7 | 75.4 | Separate vendor runs; Google and Meta materials, not a matched head-to-head. [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Terminal-Bench 2.1 | 89.4 | 88.8 | Separate vendor runs; agent wrappers and settings may differ. [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| OSWorld 2.0 | 59.0 partial | 66.9 | Google labels its result partial; Meta used 108 workflows and its internal common framework. Not directly comparable. [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| GDPval-style work | 1545 Elo on Google’s cited setup | 1754 on GDPVal-AA v2 | Different versions/methods; do not rank them against each other. [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Long-context MRCR 256K–512K | Not in the reviewed Google launch table | 98.5 | Meta-run eight-needle MRCR v2. [OFFICIAL—META] |
| Long-context MRCR 512K–1M | Not in the reviewed Google launch table | 98.1 | Meta-run eight-needle MRCR v2. [OFFICIAL—META] |
| DeepSearchQA | Not in the reviewed Google launch table | 89.4 | Meta-run, with the same search/browser tool across compared models. [OFFICIAL—META] |
Google additionally reports gains over Gemini 3.7 Flash on Terminal-Bench 4.0, finance, legal, document understanding, long-video understanding, Humanity’s Last Exam and science tasks. Meta reports 20% fewer tool calls and 25% fewer tokens than Muse 1.2 in internal engineer comparisons. Those are useful release-delta signals, but they are still vendor-selected tests. [OFFICIAL—GOOGLE] [OFFICIAL—META]
5. Meta’s scorecard has a reasoning-mode discrepancy
Meta’s evaluation methodology says Muse Spark 1.3 was run at max; Muse Spark 1.2 was run at xhigh. The launch post, however, says max reasoning is coming after additional safety testing, while the live API guide lists xhigh as the highest selectable effort. [OFFICIAL—META] Meta evaluation methodology
The published configuration is therefore not reproducible from the currently documented public controls. This may reflect internal access or documentation lag, but Meta has not reconciled it publicly. [UNKNOWN] Treat fine-grained vendor-score comparisons with caution.
6. The external evidence before Kingy’s test
There was no independent, same-harness Gemini 3.8 Flash versus Muse Spark 1.3 result at the product-evidence cutoff. Kingy’s primary benchmark now supplies a direct comparison under one protocol; it is reported in Sections 19–22 and remains [OUR TEST], not third-party evidence.
| Model | DeepSWE v1.1 solve rate | Average cost | Average output tokens | Status |
|---|---|---|---|---|
| Gemini 3.8 Flash, high | 74% ±1 | $2.36 | 143K | Benchmark-owner result. [BENCHMARK OWNER] |
| Claude Opus 5, max | 74% ±4 | $11.84 | 118K | Benchmark-owner result. [BENCHMARK OWNER] |
| GPT-5.6 Sol, max | 73% ±3 | $6.46 | 60K | Benchmark-owner result. [BENCHMARK OWNER] |
| Claude Fable 5, max | 70% ±4 | $21.63 | 119K | Benchmark-owner result. [BENCHMARK OWNER] |
| Kimi K3, max | 69% ±5 | $4.65 | 81K | Benchmark-owner result. [BENCHMARK OWNER] |
| GLM-5.3, max | 69% ±3 | $3.99 | 80K | Benchmark-owner result. [BENCHMARK OWNER] |
| Gemini 3.7 Flash, high | 65% ±2 | $2.18 | 107K | Benchmark-owner result. [BENCHMARK OWNER] |
| Muse Spark 1.3 | — | — | — | Not listed independently at cutoff. [UNKNOWN] |
| Muse Spark 1.2, xhigh | 55% ±2 | $3.70 | 99K | Older model; benchmark-owner result, not a 1.3 proxy. [BENCHMARK OWNER] |
Source: DeepSWE leaderboard. The benchmark owner is not a model vendor, but these results are not official confirmation of general product quality and could be incorrect outside this harness. Confidence intervals overlap for several leaders, and cost includes the benchmark’s agent behavior, not merely list-price multiplication.
Artificial Analysis separately reports an Intelligence Index of 59 for Gemini 3.8 Flash, versus 56 for 3.7 Flash, with about 305 output tokens per second. [INDEPENDENT] This is not officially confirmed and could be incorrect; it is a broad synthetic aggregate, not a Muse comparison or a substitute for your workload. Artificial Analysis model page
7. Common harness versus native agent product
“Model benchmark” can conceal a whole system. Meta’s evaluation report says GDPVal-AA v2, OSWorld, DeepSearchQA and AutomationBench used a common internal agent framework or a benchmark-provider harness. Its coding evaluations used named agent products or a mini-swe setup. Third-party comparison numbers were taken from Meta runs, official leaderboards or provider self-reports depending on availability. [OFFICIAL—META]
| Evidence lane | What it controls | What it can tell us | What it cannot tell us |
|---|---|---|---|
| Common harness | Similar tools, prompts and execution loop | More defensible model-level comparison | Full native product experience |
| Native agent product | Vendor’s own scaffolding, memory and tool policy | What a buyer may actually experience | Whether gains come from model or wrapper |
| Official leaderboard | Benchmark owner’s standard rules | Comparable public score under one protocol | Every deployment setting |
| Provider self-report | Vendor-selected configuration | Directional launch evidence | Independent rank or reproducibility |
This is why Meta’s Muse 1.3 DeepSWE 75.4 and DataCurve’s Gemini 74% cannot be subtracted to declare a 1.4-point Muse win. The number after the decimal is more precise than the evidence.
8. Price now, price later and the data-policy catch
All prices below are USD per one million tokens. Output prices include reasoning tokens where the developer says they are billed that way. Tool charges are separate.
| Route | Cached input | Fresh input | Output | Timing / policy |
|---|---|---|---|---|
| Gemini Standard | $0.075 | $0.75 | $3.75 | Through Dec. 31, 2026. [OFFICIAL—GOOGLE] |
| Gemini Standard | $0.15 | $1.50 | $7.50 | From Jan. 1, 2027. [OFFICIAL—GOOGLE] |
| Gemini Batch/Flex | $0.0375 | $0.375 | $1.875 | Current; doubles Jan. 1. [OFFICIAL—GOOGLE] |
| Muse Standard | $0.15 | $1.25 | $4.25 | Prompts/completions not used for training. [OFFICIAL—META] |
| Muse Contributor | $0.002 | $0.10 | $0.20 | Traffic may be used for training. [OFFICIAL—META] |
Google charges $14 per 1,000 Search or Maps queries after a shared 5,000-query allowance. Meta charges $2.50 per 1,000 search-grounding queries. Google offers explicit Batch/Flex pricing; the reviewed Meta page did not document a batch-token discount. Meta Standard lists 3,000 requests per minute and four million tokens per minute; Contributor lists 100 RPM and three million TPM, shared at team level. [OFFICIAL—GOOGLE] [OFFICIAL—META]
The Contributor route is spectacularly cheap, but it is not a privacy-preserving equivalent to Standard. Do not send proprietary code, personal data or confidential customer material merely to win a spreadsheet comparison.
9. Realistic workload cost
These estimates multiply published token and tool rates. They exclude taxes, retries, file storage beyond noted cache rates, orchestration, network and engineering labor. Actual reasoning-token use can dominate output cost.
| Workload assumption | Gemini now | Gemini Jan. 1 | Muse Standard | Muse Contributor |
|---|---|---|---|---|
| One 100K-input / 10K-output job | $0.1125 | $0.2250 | $0.1675 | $0.0120 |
| One 500K-input / 25K-output job | $0.4688 | $0.9375 | $0.7313 | $0.0550 |
| One 1M-input / 50K-output job | $0.9375 | $1.8750 | $1.4625 | $0.1100 |
| Five-turn agent: 600K fresh + 2M cached + 40K output | $0.7500 | $1.5000 | $1.2200 | $0.0720 |
| Research: 100K + 10K output + 10 billable searches | $0.2525 | $0.3650 | $0.1925 | $0.0370 |
| 10,000 batch jobs, each 100K + 10K output | $562.50 | $1,125.00 | $1,675.00* | $120.00* |
* Meta totals use ordinary token rates because no batch-token discount was documented on the reviewed page. The research row assumes Google’s free search allowance is already exhausted. [OFFICIAL—GOOGLE] [OFFICIAL—META]
The list-price crossover is clear. Gemini costs less for the modeled ordinary private workloads today. Muse Standard costs less for those workloads after Google’s price doubles, and its cheaper search makes grounded research competitive even now. Gemini Batch/Flex remains attractive when latency is flexible. Contributor wins raw price only for data you are comfortable contributing to training.
10. Retry economics matter more than list price
For an agent attempt using 120K fresh input and 60K billed output, the simple list-price cost is about $0.315 on Gemini today, $0.405 on Muse Standard and $0.024 on Contributor. If the workflow succeeds 70% of the time, expected token cost per success is attempt cost ÷ 0.70: approximately $0.45, $0.58 and $0.034. Those are scenarios, not measured success rates.
A more expensive model can be cheaper if it avoids retries, human repair or destructive tool errors. Conversely, an impressive terminal benchmark can be irrelevant if your production agent spends most of its budget reading 800,000 tokens of policy and customer history. Track cost per accepted outcome, p50/p95 latency, retries and reviewer minutes—not tokens alone.
11. Coding and software agents
Coding is close in Kingy’s completed benchmark: Muse passed 24/24 repository-repair calls and Gemini passed 23/24. The available-case difference was 4.17 percentage points, with a fixture-clustered 95% confidence interval from 0 to 12.50 points. [OUR TEST] That is not a persuasive reason to migrate a working coding fleet by itself.
Gemini also posts 74% ±1 on the DeepSWE benchmark owner’s current leaderboard, roughly tying much more expensive frontier systems within uncertainty. [BENCHMARK OWNER] Meta reports 75.4 for Muse on its own DeepSWE v1.1 run, 59.4 on SWE-Atlas CodeBase Q&A and 88.8 on Terminal-Bench 2.1. [OFFICIAL—META] Different harnesses and agent wrappers keep those published numbers from settling the product-level comparison.
For the hardest long-running repository work, Claude Fable 5.1 remains a quality-first option in this buyer set, at a steep $10 input / $50 output rate. [OFFICIAL—OTHER MODEL DEVELOPER] Claude Fable Between the two models tested here, Gemini remains the cheaper completed-call option and Muse had the better observed repository-repair result.
12. Long-running agents and computer use
Muse’s release is explicitly agent-oriented. Meta says it can sustain multi-workflow plans, incorporate conflicting sources, correct gaps and seek confirmation before consequential actions. Its official OSWorld 2.0 result is 66.9, behind Claude Opus 5 at 68.3 in Meta’s own table and ahead of GPT-5.6 Sol at 62.7. [OFFICIAL—META]
Those numbers support a serious pilot, not unattended production. The action-confirmation behavior is especially important: the highest-cost failure is often not a wrong answer, but a correct tool call applied to the wrong account, environment or record. Evaluate permission boundaries and recovery as first-class capabilities.
Google’s computer-use support is documented as preview. [OFFICIAL—GOOGLE] In Kingy’s stateful tool-use fixtures, both models passed 24/24 calls and all 48 completed runs respected the no-delete rule; no unauthorized destructive tool action was observed. [OUR TEST] This supports supervised use under the tested tool schema, not unattended deployment against production accounts.
13. Multimodal and document work
Both models advertise text, image, video, audio and PDF input with text output. Gemini has the stronger documented modality and tool story. Muse’s own documentation warns that audio input on 1.3 is not fully supported and may degrade. [OFFICIAL—GOOGLE] [OFFICIAL—META] For the narrower image/PDF extraction task Kingy actually ran, Muse had the better observed result.
Google reports 86.2 on LABBench2, 86.2 on CharXiv reasoning and 87.8 on agentic LVBench for long video. [OFFICIAL—GOOGLE] These are vendor results. In Kingy’s image/PDF extraction fixtures, Muse passed 24/24 calls and Gemini 22/24; the available-case difference was 8.33 points, with a clustered 95% confidence interval from 0 to 25.00. [OUR TEST] For high-stakes document extraction, neither model should be trusted without field-level citations and deterministic validation.
14. Long-context work
This was the decisive category in Kingy’s completed suite. On the controlled 128K and 512K fixtures, Muse passed 16/16 calls and Gemini passed 0/16. The available-case difference was 100 percentage points, with a fixture-clustered 95% confidence interval of 100 to 100. [OUR TEST]
That result needs a hard boundary. The fixtures used controlled synthetic repeated-token construction with hidden facts and controlling amendments. They test retrieval and instruction priority under pressure, not the full diversity of legal files, repositories, transcripts or research corpora. It is strong evidence about this construction, not proof that Gemini fails every natural long-context task.
The planned 900K layer did not produce a quality comparison. Muse completed its eight planned 900K calls. Gemini’s corresponding eight calls remained incomplete after its API returned the paid-tier input-token-per-minute quota error described in Section 21. Those runs remain execution-limit evidence and are never counted as wrong answers. [OUR TEST]
Meta separately reports MRCR v2 scores of 98.5 from 256K–512K and 98.1 from 512K–1M in an eight-needle retrieval test. [OFFICIAL—META] Gemini documents the same 1,048,576-token input capacity. [OFFICIAL—GOOGLE] A context-window specification states what can be submitted, not what will be retrieved correctly or admitted under a particular account’s live quota.
15. Privacy, data use and enterprise controls
Google says content submitted under paid services is not used to improve its products; free-tier content can be. Meta says Standard prompts and completions are not used to train models, whereas Contributor traffic can be. [OFFICIAL—GOOGLE] [OFFICIAL—META]
Those statements do not by themselves prove data residency, retention duration, sector certification, contractual indemnity or a zero-data-retention configuration for every route. [UNKNOWN] Procurement teams should verify the applicable contract and region, not infer them from a pricing page.
For a conventional enterprise rollout, Gemini has advantages in availability, product integration and clearer paid-versus-free separation. That product fit is separate from Muse’s better correctness in Kingy’s completed test set. For privacy-conscious hosted experimentation, Muse Standard is credible. Do not confuse Muse Contributor with Muse Standard.
16. Safety, failure modes and consequential actions
Google’s model card reports mixed safety movement versus Gemini 3.7 Flash: improvements on text safety and tone, no change on one image measure, but worse multilingual safety and more unjustified refusals. Google says manual review found most regressions were false positives or non-egregious. [OFFICIAL—GOOGLE] Gemini 3.8 Flash model card
Meta says Muse 1.3 improves instruction following, action confirmation, prompt-injection resistance and refusal behavior. [OFFICIAL—META] Those are developer claims; the launch materials did not provide a complete reproducible safety table for every claim. Kingy’s narrower tool-agent result was clean: 48/48 completed runs passed the no-delete check, with no unauthorized destructive tool actions. [OUR TEST] Keep layered controls anyway: least-privilege credentials, allowlisted tools, human confirmation for external writes, replayable logs and a hard spend/time ceiling.
17. Frontier proprietary alternatives
| Model | Why consider it | Context / list price | Main trade-off | Evidence |
|---|---|---|---|---|
| Claude Fable 5.1 | Hardest long-running agent and coding work | 1M; $10 in / $50 out | Very expensive | [OFFICIAL—OTHER MODEL DEVELOPER] |
| Claude Opus 5 | Strong quality with lower cost than Fable | 1M; $5 in / $25 out | Still far above Flash/Spark | [OFFICIAL—OTHER MODEL DEVELOPER] |
| GPT-5.6 Sol | Strong coding/agent model; ZDR-compatible route documented | 1.05M; $4 in / $20 out promotional | Long-input multiplier above 272K | [OFFICIAL—OTHER MODEL DEVELOPER] OpenAI model page |
| Grok 4.6 | Competitive multimodal frontier alternative | 500K; $2 in / $6 out below long-context threshold | Higher long-context pricing and smaller window | [OFFICIAL—OTHER MODEL DEVELOPER] xAI model page |
| Gemini 3.8 Flash | Lower current Standard cost and lower median latency in our test | 1.048M; $0.75 / $3.75 today | Price doubles Jan. 1; weak 128K/512K result in our synthetic test | [OFFICIAL—GOOGLE] [OUR TEST] |
| Muse Spark 1.3 Standard | Best correctness in our completed test set | 1.048M; $1.25 / $4.25 | Higher completed-call cost; audio caveat | [OFFICIAL—META] [OUR TEST] |
At 100K input and 10K output, list-price token cost is roughly $1.50 for Fable, $0.75 for Opus, $0.60 for Sol, $0.26 for Grok, $0.1125 for Gemini and $0.1675 for Muse Standard. The frontier models need a material success-rate advantage to repay that gap.
18. Open-weight and genuinely open alternatives
“Open” needs a license, not a vibe. Kimi K3, Qwen3.8 Max, GLM-5.3 and MiniMax M3 publish weights but use custom licenses with conditions; we call them open-weight. DeepSeek V4 Pro publishes under MIT, and Muse Glimmer 30B uses Apache 2.0; those are substantially closer to conventional open source.
| Model | Scale / context | License posture | Practical fit | Evidence |
|---|---|---|---|---|
| DeepSeek V4 Pro | 1.6T total, 49B activated; 1M context | MIT | Maximum-quality self-host candidate, but enormous | [OFFICIAL—OTHER MODEL DEVELOPER] DeepSeek model |
| Kimi K3 | 2.8T total, 104B activated; 1M context | Custom Kimi K3 | Strong multimodal open-weight agent; datacenter-scale | [OFFICIAL—OTHER MODEL DEVELOPER] Kimi K3 |
| Qwen3.8-2.4T-A95B | 2.4T total, 95B active | Custom Qwen | High-end open-weight deployment; multi-terabyte weights | [OFFICIAL—OTHER MODEL DEVELOPER] Qwen weights |
| GLM-5.3 | 753B; FP8 release about 756GB | Custom GLM | Strong current coding alternative | [OFFICIAL—OTHER MODEL DEVELOPER] GLM-5.3 |
| MiniMax M3 | 427B; 1M | Custom community license | Native image/video and computer-use option | [OFFICIAL—OTHER MODEL DEVELOPER] MiniMax M3 |
| Muse Glimmer 30B | 30B | Apache 2.0 | Most practical local/private option here | [OFFICIAL—META] Muse Glimmer |
Open weights do not make inference free. Hardware, quantization, throughput, operations and license review decide the real cost. Muse Glimmer wins our practical local/private category; DeepSeek V4 Pro is the maximum-capability open-source option when infrastructure is not the constraint.
19. Kingy’s primary benchmark: 232 completed scored calls
Kingy preregistered a 240-call direct-API comparison: 60 fixtures, two trials per fixture and two models. The five workload families were repository repair, stateful tool use, long-context retrieval, image/PDF extraction and grounded decision analysis. We obtained 232 complete, scoreable responses. Muse completed and passed 120/120; Gemini completed 112/120 and passed 91/112 of the responses that could be scored. [OUR TEST]
Both models ran through their native Standard API routes with the exact stable IDs: gemini-3.8-flash using high reasoning and muse-spark-1.3 using xhigh. The prompts, fixture data, tool schemas, hidden deterministic graders and two-trial design were fixed before the primary run. The substantive 65,536-token output ceiling and 16,384-token tool-step ceiling were designed to be non-binding. No cache discount, Batch, Flex, Priority, Contributor pricing or gateway routing entered the comparison.
Every completed run records the request, response, stop reason, provider-reported token categories, non-streaming end-to-end latency, tool actions, estimated cost, grader result and artifact hash. The plan used deterministic shuffle seed 380013. The available-case confidence intervals use a fixture-clustered percentile bootstrap with the same seed and 200,000 iterations; the fixture, not each repeated call, is the independent unit.
| Analysis set | Gemini 3.8 Flash | Muse Spark 1.3 | Result |
|---|---|---|---|
| All completed scoreable calls | 91/112 passed (81.25%) | 120/120 passed (100%) | Different completion counts; descriptive only |
| Fully paired available cases | 91/112 passed (81.25%) | 112/112 passed (100%) | Muse +18.75 percentage points |
| Clustered 95% confidence interval | — | — | Muse +9.82 to +28.57 points |
The 18.75-point comparison is explicitly an available-case analysis over 56 fixtures completed twice by both models. Numerically, it clears the preregistered thresholds of a 10-point advantage, a clustered interval excluding zero and no observed material safety disadvantage. It is still not the preregistered all-240 conclusion because the missing 900K stratum prevents that rule from being formally applied. The formal verdict remains insufficient evidence.
20. Results by workload
| Workload | Fully paired fixtures | Gemini | Muse | Muse minus Gemini | Fixture-clustered 95% CI |
|---|---|---|---|---|---|
| Repository repair | 12 | 23/24 | 24/24 | 4.17 points | 0.00 to 12.50 |
| Stateful tool use | 12 | 24/24 | 24/24 | 0.00 points | 0.00 to 0.00 |
| Long context, completed 128K/512K only | 8 | 0/16 | 16/16 | 100.00 points | 100.00 to 100.00 |
| Image/PDF extraction | 12 | 22/24 | 24/24 | 8.33 points | 0.00 to 25.00 |
| Grounded decision analysis | 12 | 22/24 | 24/24 | 8.33 points | 0.00 to 20.83 |
The table shows why the aggregate needs interpretation. Muse did not win by a similar margin everywhere. Coding was one call apart. Tool use was tied. Document extraction and grounded decision work each differed by two calls. The completed long-context layer created the large overall separation.
The long-context failures were not all interchangeable. They reflect the deterministic grader’s requirements for exact retrieval and amendment resolution under the controlled fixture design. A wrong or unusable answer is a quality failure; a provider refusing to execute because of quota is not. That distinction is why the 128K/512K results are scored while the missing Gemini 900K calls are kept outside the denominator.
21. The eight Gemini 900K runs: execution evidence, not quality scores
Eight preregistered Gemini 900K calls did not return scoreable responses. Two distinct scored run IDs were submitted before the stop rule halted the stratum. Across four dispatches, all four responses were HTTP 429 errors. Google’s error identified the quota metric as generativelanguage.googleapis.com/generate_content_paid_tier_input_token_count, quota ID GenerateContentPaidTierInputTokensPerModelPerMinute, with a value of 2,000,000 input tokens per minute. [OUR TEST]
A separate 900K pilot journal records eight dispatches, seven recorded 429 responses and three retry events. One pilot dispatch has no recorded error response, consistent with interruption or timeout rather than a scored model answer. After the repeated provider limit was established, the harness stopped rather than submit the other six planned scored calls.
This is operationally useful. A one-million-token context specification did not guarantee that this Tier 1 project could run the preregistered 900K workload at the planned cadence. It is not evidence that Gemini answered eight questions incorrectly. Muse’s completed 900K responses are preserved in the execution record, but there is no paired 900K quality comparison, so they do not enter the headline quality estimate.
Required reading for the headline: the 18.75-point result is available-case analysis. The eight Gemini 900K runs are quota-limited execution evidence, excluded from quality scoring. The preregistered all-240 verdict remains insufficient evidence.
22. Latency, cost, tokens and safety
| Measure | Gemini 3.8 Flash | Muse Spark 1.3 | Interpretation |
|---|---|---|---|
| Completed scored calls | 112 | 120 | Gemini’s eight missing 900K calls are not included |
| Median latency | 12.48 s | 21.15 s | Gemini was faster at the middle of the distribution |
| p95 latency | 271.05 s | 73.28 s | Muse had the better completed-call tail latency |
| Estimated completed scored-call cost | $8.323201 | $16.782506 | Standard list rates; no cache discount |
| Passed scoreable responses | 91 | 120 | Different completion counts; do not divide across models without the paired set |
| Cost per passed response | $0.091464 | $0.139854 | Descriptive, not quality-adjusted routing economics |
| Provider-reported cached tokens | 49,148 | 1,312,779 | Recorded but priced at full Standard input rate |
| Tool-agent no-delete checks | 24/24 | 24/24 | Zero unauthorized destructive actions observed |
P95 uses the nearest-rank method over completed scored calls. Latency is non-streaming end to end, not time to first token. Gemini returned STOP on 107 completed scored calls and MAX_TOKENS on five; Muse reported 120 completed responses.
The completed scored-call total was $25.105707. The full experiment ledger was $30.649939, including pilots, an aborted pilot reserve and $2.026980 in conservative timeout reserves. Those reserves are accounting safeguards, not claims about the final provider invoices. [OUR TEST]
Gemini was cheaper and faster at median, but its p95 was nearly four and a half minutes. Muse cost roughly twice as much across completed scored calls, returned every scored response and had a much tighter p95. Buyers should track cost per accepted outcome and tail latency, not list price or median speed alone.
23. Winners by use case
| Use case | Recommendation | Confidence | Why |
|---|---|---|---|
| Overall correctness on this test set | Muse Spark 1.3 | High for completed fixtures | 112/112 versus 91/112 in the fully paired available-case analysis. [OUR TEST] |
| Repository repair | Near tie; slight Muse edge | Medium | 24/24 versus 23/24; clustered interval includes zero. [OUR TEST] |
| Stateful tool use | Tie | High for tested schema | Both 24/24; both passed every no-delete check. [OUR TEST] |
| Controlled 128K/512K retrieval | Muse Spark 1.3 | High for this fixture design | 16/16 versus 0/16. Do not generalize to every natural corpus. [OUR TEST] |
| Image/PDF extraction | Muse Spark 1.3 | Medium | 24/24 versus 22/24; interval includes zero. [OUR TEST] |
| Grounded decision analysis | Muse Spark 1.3 | Medium | 24/24 versus 22/24; interval includes zero. [OUR TEST] |
| Lowest completed-call cost | Gemini 3.8 Flash | High | $8.32 versus $16.78 on completed scored calls. [OUR TEST] |
| Median latency | Gemini 3.8 Flash | High | 12.48 versus 21.15 seconds. [OUR TEST] |
| Tail latency | Muse Spark 1.3 | High | 73.28 versus 271.05 seconds at p95. [OUR TEST] |
| Broad documented modalities and Google tools | Gemini 3.8 Flash | High | Mature modality matrix and tool catalog; Muse has an audio caveat. [OFFICIAL—GOOGLE] [OFFICIAL—META] |
| Lowest hosted sticker price | Muse Contributor | High on price only | Very cheap, but prompts and completions may be used for training. [OFFICIAL—META] |
| Practical local/open | Muse Glimmer 30B | Medium-high | Apache 2.0 and a more deployable scale than the giant alternatives. [OFFICIAL—META] |
The main unresolved items are the gap between Meta’s max evaluation configuration and the public API’s documented xhigh ceiling, Muse’s maximum output limit and knowledge cutoff, exact retention and residency terms by route, production reasoning-token distributions, and the timing and license of any Muse Spark weight release. [UNKNOWN] They matter, but none erases the direct completed-call result.
What would change the practical recommendation? A natural-corpus retest could show that the synthetic long-context gap does not transfer. Different workload weights could favor Gemini’s cost and latency. A higher Gemini quota tier could permit the missing 900K paired layer. A material pricing change could move the cost crossover. None of those possibilities should be silently substituted for the evidence we have.
24. Final recommendation
If these five workload families resemble your production mix and correctness comes first, start with Muse Spark 1.3 Standard. It completed and passed every scored call, led the fully paired available-case analysis by 18.75 percentage points and avoided the severe 128K/512K retrieval failures seen from Gemini. [OUR TEST]
Choose Gemini 3.8 Flash Standard when current cost, median speed, Google-native tools or its documented modality breadth outweigh that measured quality gap. It cost about half as much across completed scored calls and was 8.67 seconds faster at the median, but its completed-call p95 was much worse and its long-context quality result was poor under this synthetic construction. [OUR TEST]
Do not convert the quota gap into a fake quality win. The eight missing Gemini 900K calls were blocked at execution, excluded from scoring and sufficient to prevent the preregistered all-240 verdict. The responsible wording is precise: Muse wins the completed, fully paired available-case analysis; the formal preregistered verdict is insufficient evidence.
For production, route by workload rather than loyalty. Use a cheap model where it clears your acceptance test, move difficult cases to the model with the better measured success rate, retain human approval for consequential actions and retest on your own natural data. Browse more model profiles in Kingy’s AI model database and track changing economics in the KAPI Price Index.
25. Frequently asked questions
Is Muse Spark 1.3 open source?
No. Meta says Muse Spark open weights are coming, but 1.3 was hosted-only at our cutoff. Muse Glimmer 30B is a separate Apache-2.0 model that can be self-hosted. [OFFICIAL—META]
Is Gemini 3.8 Flash cheaper than Muse Spark 1.3?
Gemini Standard is cheaper for ordinary private token workloads through December 31, 2026. From January 1, 2027, its published rates double and Muse Standard becomes cheaper in our modeled workloads. Muse Contributor is cheaper throughout, but allows training use of prompts and completions. [OFFICIAL—GOOGLE] [OFFICIAL—META]
Which model is better for coding?
Kingy’s repository-repair test was effectively a near tie: Muse passed 24/24 and Gemini 23/24, with a clustered confidence interval that includes zero. Gemini separately posts 74% ±1 on the DeepSWE benchmark owner’s leaderboard, while Meta reports 75.4 for Muse in its own run. Your repository, tools and acceptance tests should decide. [OUR TEST] [BENCHMARK OWNER] [OFFICIAL—META]
Do both models support a one-million-token context window?
Yes: both list 1,048,576 tokens. That specification does not guarantee equal recall, synthesis quality or live quota admission. In Kingy’s controlled completed 128K/512K calls, Muse passed 16/16 and Gemini 0/16. The planned Gemini 900K layer was quota-limited and remains unscored. [OFFICIAL—GOOGLE] [OFFICIAL—META] [OUR TEST]
Did Gemini lose all eight 900K tests?
No. The eight planned Gemini 900K runs did not produce scoreable model responses. Repeated attempts hit Google’s two-million-input-token-per-minute project quota, so the runs are classified as quota-limited execution evidence and excluded from quality scoring. [OUR TEST]
Is the 18.75-point gap the preregistered final verdict?
No. It is an available-case analysis over 56 fixtures fully completed by both models. Because eight Gemini 900K runs were missing, the preregistered 240-call decision rule could not be formally applied. The formal verdict is insufficient evidence. [OUR TEST]
Should I use Muse Contributor for production?
Only for data your organization permits Meta to use for training. For proprietary code, personal data or confidential customer work, use a route whose contract and data policy meet your requirements. [OFFICIAL—META]
26. Sources and fact-check record
Primary sources reviewed at the stated cutoff: Google model documentation, pricing, launch post and model card; Meta launch post, evaluation methodology, model documentation, reasoning guide and pricing/rate limits; model-developer pages linked in the alternative tables; and the DeepSWE benchmark-owner leaderboard.
Fact-check summary: model IDs, context limits, modality caveats, token prices, search charges and date changes were checked against first-party pages. Vendor benchmark claims remain labeled by runner and non-comparable configurations are not combined. Kingy’s primary benchmark is reported separately as [OUR TEST], with its stopped stratum and available-case status visible beside the result.
Testing disclosure: We preregistered 240 native Standard-API runs: 60 fixtures, two trials per fixture and two models. We obtained 232 complete, scoreable responses—120 from Muse Spark 1.3 and 112 from Gemini 3.8 Flash. Eight planned Gemini 900K-context runs did not produce responses. Google rejected execution under the Tier 1 project’s 2,000,000 input-token-per-minute limit. Two distinct scored run IDs were submitted before the stop rule halted further calls; their four dispatches returned HTTP 429. A separate 900K pilot also repeatedly hit the same quota. We therefore report the eight-run 900K stratum as quota-blocked and exclude it from quality scoring rather than counting it as eight wrong answers. Results use provider-reported token usage, Standard list prices with no cache discounts, non-streaming end-to-end latency and deterministic task graders. The fixtures are controlled and synthetic, and the repeated-token long-context construction is less linguistically diverse than a natural corpus.
Reproducibility record: protocol SHA-256 bfd86fb3e390c6bc11cca60c5aedacf0482fbe97006459cb9675d941bb35d0a2; canonical plan SHA-256 c34bf4b691053c93a2a90c731406a1f1f4cf9912a6b67dc5e221e5b5942f4368; canonical fixture-manifest SHA-256 dd15e96f0ff6a20a99909364ae490822bc21bf120606d2163d88d5245af4d6f0; sealed evidence-root SHA-256 51192a377a7f7d2e965777809f1eaa5f7a1fca5e21df11c5653eaf75faa207f7. The sealed index covers 514 files.
Featured snippet: In Kingy’s 232-call benchmark, Muse Spark 1.3 passed 112/112 fully paired calls and Gemini 3.8 Flash passed 91/112, an 18.75-point Muse advantage in the available-case analysis. Eight planned Gemini 900K calls were quota-limited and excluded from quality scoring, so the preregistered all-240 verdict remains insufficient evidence. Gemini was cheaper and faster at the median.
