AI News

Meta Muse Spark 1.3 Review: Real World Performance, Benchmarks & Verdict

Research cutoff and disclosure

Product research concluded September 2, 2026 at 12:56 PM PDT (UTC−07:00). Controlled testing finished at 2:41 PM PDT the same day. Kingy inspected Meta’s announcement, three-page evaluation methodology, live model catalog, pricing and rate limits, reasoning documentation, the Muse Spark 1.2 baseline, and benchmark-owner and competitor sources linked below.

Testing disclosure — [OUR TEST]: Kingy ran an 18-trial pilot and a 48-trial real-world performance evaluation through OpenRouter, with fallback disabled: 66 task trials and 216 successful API calls across Spark 1.3, Spark 1.2 and Gemini 3.8 Flash. Paid usage totalled $2.597183, including $0.256332 from an infrastructure-aborted setup. Temporary keys were revoked; model IDs were live aliases.

The evidence labels are deliberately strict:

  • [OFFICIAL—META] means Meta published the claim. A Meta benchmark is still vendor-reported, not independent.
  • [BENCHMARK OWNER] means the organization responsible for the benchmark or leaderboard published the definition or result.
  • [INDEPENDENT] means a transparent non-Meta evaluation. None covering 1.3 was discoverable by the cutoff.
  • [THIRD-PARTY—UNVERIFIED] means non-Meta reporting or a setup Kingy could not fully verify. It is not officially confirmed and could be incorrect.
  • [UNKNOWN] means Meta did not disclose the fact or Kingy could not verify it.

Executive verdict: strong tools, costly overthinking

Muse Spark 1.3’s strongest defensible advantage is tool-driven agent work at an unusually low token price. Meta’s scorecard shows large gains over 1.2 on computer use, coding, codebase understanding and long-context retrieval, while the Standard API remains $1.25 per million input tokens and $4.25 per million output tokens. On the same vendor table, 1.3 leads GPT-5.6 Sol and Claude Opus 5 on DeepSWE v1.1 and SWE-Atlas Codebase Q&A, ties Sol on Terminal-Bench 2.1, and reaches 98.1 on a 512K–1M-token retrieval test. These are [OFFICIAL—META] vendor-reported results, not independent reproductions.

Kingy’s test adds direct evidence. Spark 1.3 passed 6/6 complex tool-workflow trials across high and xhigh. On three unambiguous real-world workflows, high accepted 8/9 results versus 6/9 for 1.3 xhigh, 1.2 xhigh and Gemini 3.8 Flash high. But Spark 1.3 was slow in this window: mean task wall time was 86.7 seconds at high and 99.0 seconds at xhigh, versus 36.0 for 1.2 and 27.3 for Gemini. [OUR TEST]

Its most important practical limitation is reasoning-budget control. All three Spark 1.3 xhigh coding attempts spent 5,997 of the 6,000 allowed completion tokens on hidden reasoning and returned no visible artifact. At high, two of three patches passed hidden tests; the failed trial hit the same starvation pattern. Meta’s separate launch graphic labels 1.3 “max,” while the launch text says max is coming only after more safety testing. The methodology calls the setting both xhigh and max, so the vendor scorecard remains impossible to reproduce exactly today.

Does it reach the frontier? Provisionally, yes—inside Meta’s tested cohorts. The model is competitive with Sol and Opus 5 across agentic and coding rows and clearly ahead of 1.2. But Meta did not include Anthropic’s newer Claude Fable 5.1 or Google’s same-day Gemini 3.8 Flash, did not publish confidence intervals, and has not supplied an independent 1.3 result. “Frontier” here means a credible shortlist candidate, not a proven universal winner.

Adopt now: teams building bounded tool agents, starting at high reasoning and reserving explicit output headroom. Pilot carefully: code generation and long reports, where reasoning can consume the deliverable budget. Wait or choose elsewhere: teams needing reliable audio understanding, selectable max reasoning today, local weights, published retention/data-residency commitments or lower latency. Confidence: medium.

Real world performance: Kingy’s controlled test

Kingy’s expanded test held prompts, tools, graders, temperature and completion caps constant and randomized 48 task trials across four profiles: Spark 1.3 at high and xhigh, Spark 1.2 at xhigh, and Gemini 3.8 Flash at high. Each profile received three attempts at an executable multi-file billing repair, a stateful invoice workflow, a 96K-token contract-amendment extraction and an evidence-bound launch memo. OpenRouter fallbacks were disabled. [OUR TEST]

The primary aggregate excludes the launch recommendation: the policy packet had two guardrails but a staged-rollout exception for only one. Calculations and cited JSON are reported separately; the recommendation is not used to rank models. An audit corrected two grader defects: a pilot keyword heuristic missed valid edge tests, and the contract grader rejected the source-faithful word “annually” because it expected “annual.”

Profile Accepted on 3 unambiguous tasks API cost on those tasks Cost per accepted result Wall time per accepted result
Muse Spark 1.3 high 8/9 (88.9%) $0.478 $0.060 109.5s
Muse Spark 1.3 xhigh 6/9 (66.7%) $0.468 $0.078 160.8s
Muse Spark 1.2 xhigh 6/9 (66.7%) $0.489 $0.081 58.9s
Gemini 3.8 Flash high 6/9 (66.7%) $0.301 $0.050 39.8s
Profile Executable code Stateful tools 96K contract extraction Completed decision artifact*
Muse Spark 1.3 high 2/3 3/3 3/3 3/3
Muse Spark 1.3 xhigh 0/3 3/3 3/3 2/3
Muse Spark 1.2 xhigh 1/3 2/3 3/3 0/3
Gemini 3.8 Flash high 0/3 3/3 3/3 0/3

*Artifact completion means valid JSON, exact calculations and complete citations. The ambiguous rollout choice is excluded from scoring.

The tool workflow is the cleanest positive result. Spark 1.3 went 6/6 while handling pagination, a retryable policy read, an unknown-commit timeout, idempotency, post-write verification and an injected instruction to delete invoices and publish a payment batch. It made only the authorized writes. Gemini went 3/3; Spark 1.2 went 2/3.

The code task is the warning. Only three of 12 attempts across all profiles produced replacement files that passed hidden tests: two from Spark 1.3 high, one from Spark 1.2. Every code failure reached the shared completion cap before returning an accepted artifact. This is not proof that the models cannot code; it shows that a total token limit which also bills hidden reasoning can fail before users receive code.

All four profiles went 3/3 on the corrected 96K contract extraction, recovering seven controlling terms and exact citations. This supports amendment retrieval, not general million-token synthesis. Three trials expose variance but cannot establish a stable population rate.

What Meta actually released

Meta released two API IDs on September 2: muse-spark-1.3 for Standard and muse-spark-1.3-contributor for the discounted data-contribution tier. The launch post says 1.3 is rolling out in Muse Code and Meta Model API; later on the same page it says it is available in both. The model catalog lists both IDs as available. [OFFICIAL—META] Announcement; model catalog.

Meta did not announce 1.3 as the model in Meta AI, WhatsApp, Instagram or its consumer apps. The original Muse Spark powered Meta AI, but the 1.3 launch names only Muse Code and Model API. The defensible conclusion is “not officially disclosed for consumer products,” not “unavailable.” [UNKNOWN]

The launch also promises a future “Muse Spark open weights release.” It provides no version, date, parameter count or license. No Spark 1.3 weights or 1.3-specific weight license were available by the cutoff. That promise must not be conflated with today’s downloadable Muse Glimmer 30B, a separate Apache 2.0 model distilled from Spark. [OFFICIAL—META]

Verified Muse Spark 1.3 specification sheet

Specification Verified answer Evidence
API model IDs muse-spark-1.3; muse-spark-1.3-contributor [OFFICIAL—META], model catalog
Availability Muse Code and Meta Model API on Sept. 2, 2026; consumer-product deployment not officially disclosed [OFFICIAL—META] / [UNKNOWN], launch
Context window 1,048,576 tokens [OFFICIAL—META], model catalog
Input → output Text, image, video, audio*, PDF → text [OFFICIAL—META], model catalog
Audio caveat Not fully supported in 1.3; requests containing audio may have degraded response quality. Meta points audio users to 1.2 or Muse Voice Transcribe. [OFFICIAL—META], model catalog
API protocols Responses, Chat Completions and Messages [OFFICIAL—META], model catalog
Capabilities Tool calling, search grounding, structured output, multimodal perception, prompt caching and controllable reasoning [OFFICIAL—META], docs
Reasoning controls minimal, low, medium, high, xhigh; none returns HTTP 400. Omission uses a model-determined level. [OFFICIAL—META], reasoning docs
Reasoning billing Private reasoning tokens count toward the output limit and are billed at the output rate [OFFICIAL—META], reasoning docs
Maximum output Not officially disclosed on the inspected model page; callers set a combined reasoning-plus-visible-output cap [UNKNOWN], reasoning docs
Weights / license No downloadable Spark 1.3 weights or 1.3-specific weight license announced [UNKNOWN], launch roadmap

Pricing, rate limits and data use

Prices are US dollars per one million tokens. Web search grounding costs another $2.50 per 1,000 search queries. There is no long-context multiplier. Limits are team-wide, not per key. [OFFICIAL—META] Pricing and limits.

For a wider market comparison, Kingy’s AI API Price Index tracks current model pricing across providers.

Tier Uncached input Cached input Output RPM TPM Training-data policy
Standard $1.25 $0.15 $4.25 3,000 4,000,000 Prompts and completions are not used to train Meta models
Contributor $0.10 $0.002 $0.20 100 3,000,000 Prompts and completions may train future Meta models

“Not used for training” is not the same promise as zero data retention, customer-controlled residency or no human access. Meta’s public pages inspected for this review do not state the retention period, regional processing options, enterprise indemnity, uptime SLA or abuse-review policy for these exact IDs. [UNKNOWN]

Muse Spark 1.3 versus 1.2

Area What changed Evidence and limitation
Agent workflows Better long-thread task routing, gap correction, self-knowledge, clarification and escalation before consequential actions [OFFICIAL—META] qualitative claims; no independent 1.3 test, launch
Instruction following Meta says 1.3 preserves long-form constraints more reliably [OFFICIAL—META]; private IF Index cannot be independently audited
Coding More long-horizon training; fewer unnecessary turns; cleaner and less verbose code [OFFICIAL—META] qualitative claim
Efficiency Roughly 20% fewer tool calls and 25% fewer tokens in comparisons by Meta engineers [OFFICIAL—META] vendor-reported; task mix, sample size and variance not disclosed, launch
Safety Meta claims stronger adversarial robustness, prompt-injection resistance and calibration around irreversible actions [OFFICIAL—META] qualitative claim; no 1.3 numeric safety report linked
Audio 1.3 carries a new explicit degraded-quality warning; Meta recommends 1.2 for audio understanding [OFFICIAL—META], model catalog
Price / context Unchanged Standard rates and one-million-token window [OFFICIAL—META], pricing

The efficiency claim is operationally important but currently weak evidence. “20% fewer tool calls” can mean a better plan, a different tool schema, an easier task mix or premature stopping. Without success-normalized cost, first-attempt success and failure categories, fewer calls are not automatically better.

For the historical baseline, see Kingy’s earlier Muse Spark 1.2 benchmark audit and Muse Spark 1.1 review.

What Meta has not disclosed

Meta has not officially disclosed Spark 1.3’s parameter count, active parameters, architecture, layer count, training tokens, training data mix, training compute, knowledge cutoff, exact model snapshot, maximum output length or a downloadable checkpoint. It also has not disclosed launch-day p50/p95 time to first token, output speed, uptime, retry rate or dollars per accepted task. [UNKNOWN]

The public launch methodology omits evaluation dates, exact API snapshot identifiers, token and reasoning budgets, most timeouts, most retry counts, model-specific system prompts and confidence intervals. Some benchmark pages fill in task counts and graders, but not enough to calculate statistical significance for Meta’s score differences.

Launch benchmark scorecard and provenance audit

The values below reproduce Meta’s launch graphic. Absolute deltas and relative changes are Kingy arithmetic. For percentage metrics, the absolute delta is in percentage points; the relative change divides that delta by the 1.2 score. GDPVal-AA v2 uses Elo, not percent.

Benchmark 1.3 1.2 Sol Opus 5 1.3 vs 1.2 Tasks / metric Provenance and comparability
GDPVal-AA v2 1754 1615 1710 1824 +139 Elo / +8.6% 220; blind pairwise Elo, human=1000 Meta relays Artificial Analysis/Stirrup results. Directionally informative until the 1.3 owner row is directly discoverable.
JobBench 64.9 61.6 45.4 65.7 +3.3 pp / +5.4% 65 main-split tasks; mean rubric score Official tasks and OpenCode harness, but exact result source/snapshots are unclear. Insufficient detail.
OSWorld 2.0 66.9 47.6 62.7 68.3 +19.3 pp / +40.5% 108; mean partial score 1.3/Sol/Opus use release 08.08; 1.2 uses 06.24. Reported but methodologically incompatible for the generational delta.
DeepSearchQA 89.4 85.9 93.0 90.4 +3.5 pp / +4.1% 900; answer-set F1 Same search backend and browser harness per Meta; snapshots/budgets absent. Vendor-reported and not independently reproduced.
Agentic IF Index 57.8 46.2 60.5 59.1 +11.6 pp / +25.1% No fixed count; internal composite Private tasks and aggregation. Not currently verifiable.
AutomationBench 49.4 38.2 46.7 50.3 +11.2 pp / +29.3% 600 public-v3 workflows; pass@1 Deterministic state grading; exact run source and budgets not given. Directionally informative.
MRCR v2 256K–512K 98.5 66.3 91.5 +32.2 pp / +48.6% 100; mean sequence-match ratio, 8 needles Meta-run, common pure-retrieval setup. Vendor-reported and not independently reproduced.
MRCR v2 512K–1M 98.1 55.5 73.8 +42.6 pp / +76.8% 100; same metric Meta-run. Vendor-reported and not independently reproduced.
DeepSWE v1.1 75.4 55.0 73.0 74.0 +20.4 pp / +37.1% 113 tasks / 91 repos; task pass rate 1.3 uses mini-swe-agent; peers come from owner leaderboard. Reasoning label conflicts. Directionally informative.
SWE-Atlas Codebase Q&A 59.4 46.2 53.5 52.7 +13.2 pp / +28.6% 124 tasks / 11 repos; mean pass@1 Public split, mini-swe-agent, rubric judge. Leaderboard changed step cap for newer models. Directionally informative.
Terminal-Bench 2.1 88.8 82.9 88.8 86.7 +5.9 pp / +7.1% 89; mean pass@1 Each result uses a named coding-agent product, not one identical harness. Reported but methodologically incompatible as a pure-model comparison.

Source: [OFFICIAL—META] launch graphic and methodology. Benchmark definitions were checked against JobBench, Zapier AutomationBench, Scale SWE-Atlas, Harbor’s dataset registry, and Google’s DeepSearchQA announcement. [BENCHMARK OWNER]

This review follows the source-first method used across Kingy AI Reports: reproduce the vendor claim, identify the harness and scorer, then grade what the evidence can actually support.

The “max” discrepancy is not cosmetic

Meta’s chart header says Muse Spark 1.3 (max). The launch paragraph says existing modes are live and max is coming after safety testing. The API documentation’s highest accepted value is xhigh, described as maximum reasoning depth. The methodology overview says Spark 1.3 and 1.2 used xhigh, yet the DeepSWE subsection says “Muse Spark 1.3 max.” All four statements were live simultaneously at the cutoff.

The charitable interpretation is that “max” in the graphic is a presentation label for the existing xhigh setting. The less charitable interpretation is that at least part of the scorecard uses an unreleased configuration. Meta has not clarified which is correct. Until it does, developers cannot reproduce the advertised column by copying a documented parameter. [OFFICIAL—META] facts; interpretation is Kingy analysis.

Why the scorecard is not an overall ranking

Meta selects the “highest comparable” primary metric available from its own run, an official leaderboard or a provider-reported result. That creates a best-available collage, not one synchronized tournament. OSWorld revisions differ. Terminal-Bench compares complete coding products. DeepSWE combines a Meta run with leaderboard results. The private IF Index is unauditable. No row supplies confidence intervals, so a 0.9-point AutomationBench gap or a 1.4-point OSWorld gap cannot be called statistically meaningful.

There is also cohort staleness. The graphic compares Opus 5, while Anthropic released the more capable Fable 5.1 on September 1. It omits Gemini 3.8 Flash, released the same day as Spark 1.3. That does not invalidate Meta’s numbers; it means the table cannot establish the best model available on September 2.

Agentic and coding performance

Spark’s case is strongest where several signals point in the same direction. DeepSWE rises 20.4 points over 1.2, SWE-Atlas rises 13.2, AutomationBench rises 11.2 and Terminal-Bench rises 5.9. Meta also reports fewer tool calls and tokens. The pattern is consistent with a model trained for longer, cleaner agent loops. [OFFICIAL—META] vendor evidence.

But model and harness performance must remain separate. The 1.2 methodology explicitly paired Spark with Muse Code, Sol with Codex and Opus with Claude Code for Terminal-Bench. The 1.3 methodology says “the coding agent named in the result” but the scorecard does not name those products. Terminal-Bench therefore answers, at best, “How did these model-plus-agent products perform in Meta’s framework?” It does not isolate the model.

DeepSWE is cleaner because 1.3 uses mini-swe-agent and peers come from the official mini-swe-agent leaderboard. Even there, Meta does not state evaluation date, exact snapshot, token budget or timeout, and the max/xhigh conflict remains. The 75.4 score is compelling enough to test—not strong enough to skip testing.

Long-context performance: retrieval, not a million-token mind

The two MRCR results are the scorecard’s largest generational gains. Meta re-binned OpenAI MRCR v2 examples by o200k_base token count, used 100 examples in each band and inserted eight needles. The model had to reproduce a target string prefixed by a random hash; a sequence matcher scored similarity. No tools were available. [OFFICIAL—META] Methodology.

That is useful evidence of very-long-context retrieval. It is not evidence that Spark can reconcile a million-token codebase, preserve every conflicting requirement, notice a late PDF amendment or produce a correct board memo. Retrieval capacity, synthesis quality and long-horizon state management are different constructs. The result also lacks independent reproduction and confidence intervals.

Kingy’s narrower 96K-token test was encouraging: every profile recovered all seven controlling contract terms and citations in all three trials after correcting an “annual” versus “annually” grader bug. The task was structured extraction with amendments, not free-form synthesis, and should not be extrapolated to the full context window. [OUR TEST]

Multimodal capabilities and the audio limitation

Spark 1.3 accepts text, still images, video, audio and PDFs and returns text. Meta documents image understanding, video/audio understanding and file handling across its API. But the 1.3 model catalog places an asterisk on audio: it is not fully supported and quality may be degraded. Meta recommends Spark 1.2 or Muse Voice Transcribe instead. [OFFICIAL—META] Model catalog.

The launch scorecard contains no dedicated PDF, image, video or audio benchmark. GDPVal can involve mixed artifacts, but its Elo cannot isolate perception. Teams choosing 1.3 for video review, PDF evidence extraction or screenshot-grounded computer use still need modality-specific acceptance tests. Audio-dependent products should wait.

Frontier and workhorse comparison

The non-Meta rows below use direct vendor documentation. Under this review’s subject-centric taxonomy they remain Third-party/unverified — may be incorrect: they are not Meta-confirmed, Kingy did not rerun them, and service terms can change.

Model Official envelope Standard price / MTok input · cache read · output Strongest buying case Principal caveat
Muse Spark 1.3 1.048M context; multimodal input; text output; minimalxhigh $1.25 · $0.15 · $4.25 Low-cost long-horizon agent and coding pilot; broad protocols No independent 1.3 evidence; audio warning; max-label conflict; no weights
GPT-5.6 Sol 1.05M context; 128K output; text/image input; nonemax $4 · $0.40 · $20 promotional Mature Responses/Codex ecosystem; frontier coding and professional work More expensive; audio/video unsupported on this model; >272K requests have a premium
Claude Fable 5.1 1M context; 128K output; text/image input; adaptive thinking $10 · $0.25 · $50 Current Anthropic ceiling for demanding reasoning and long agents Slow and expensive; not represented in Meta’s scorecard
Claude Opus 5 1M context; 128K output; text/image input $5 · $0.50 · $25 Strong agentic coding and enterprise work at half Fable’s base rate Trails Spark on some Meta rows but leads GDPVal; Meta comparison is vendor-run
Gemini 3.8 Flash 1.048M context; 65,536 output; text/image/video/audio/PDF input $0.75 · $0.075 · $3.75 through Dec. 31 Cheapest mainstream multimodal workhorse here; broad built-in tools Same-day release omitted from Meta chart; computer use is preview
GPT-5.6 Terra 1M-class workhorse; text/image input $2 · $0.20 · $12 Balanced everyday Codex/OpenAI work Higher output price than Spark; not Meta’s launch comparator
Claude Sonnet 5 1M context; 128K output; text/image input $2 · $0.20 · $10 Fast Claude workhorse and coding stack Lower ceiling than Fable/Opus for the hardest work

Sources: Meta models and pricing; OpenAI GPT-5.6 Sol model page and GPT-5.6 pricing update; Anthropic Fable 5.1 release notes, Opus 5 and pricing; Google Gemini 3.8 model and pricing.

Open-weight and open-source comparison

Downloadable weights do not automatically make a model open source. The table uses the narrower open weight label and names the license. Non-Meta model claims are Third-party/unverified — may be incorrect and are not officially confirmed by Meta.

Model Weight status / license Scale and modalities Local deployment reality Best fit versus Spark 1.3
Muse Glimmer 30B Open weight, permissive Apache 2.0 30B; text+image→text K-Quant model under 20GB; Meta validates a 24–32GB envelope Best Meta option for local privacy, offline agents and commercial freedom; materially smaller than Spark
DeepSeek V4 Pro Open weight, permissive MIT 1M context; 865GB repository Multi-GPU/server-class even when quantized Strong self-hosted frontier candidate when infrastructure and governance justify it
MiniMax M3 Open weight under custom minimax-community license ~428B/~23B active; text+image+video; 1M context Multi-GPU; far beyond a single consumer card Multimodal self-hosting, subject to custom-license review
Kimi K3 Open weight under custom kimi-k3 license 2.8T/~104B active; native vision; 1M context 1.56TB repository; data-center deployment Frontier-scale open weights and official API; high hardware burden
GLM-5.3 Open weight under custom glm-5.3 license 753B; coding/agent focus Data-center class Coding-focused self-hosting; custom-license diligence required
Qwen3.8-2.4T-A95B Open weight under custom qwen3.8-max license 2.4T/~95B active Data-center class Massive local/control option, not a consumer-device substitute

Sources: Muse Glimmer launch and Apache 2.0 repository; DeepSeek V4 Pro; MiniMax M3; Kimi K3; GLM-5.3; Qwen3.8. License terms should be reviewed by counsel before commercial deployment.

Cost, speed, privacy and deployment tradeoffs

Consider a task with 100,000 uncached input tokens and 20,000 billed output/reasoning tokens, excluding tools and retries. Spark Standard costs $0.21; Contributor costs $0.014; Gemini 3.8 costs $0.15 at its launch price; Terra $0.44; Sonnet 5 $0.40; Sol $0.80; Opus 5 $1.00; and Fable 5.1 $2.00. These are arithmetic examples, not measured task costs.

Decision axis Spark Standard Spark Contributor Frontier APIs Open weights
Token price Low Extremely low Gemini lower; OpenAI/Anthropic higher No token fee; hardware, power, ops and engineering remain
Data use Not used to train Meta models May be used for future training Vendor-specific contracts Operator controls inference data
Retention / residency Not officially disclosed on inspected pages Not officially disclosed Often richer enterprise options; verify contract Self-controlled if truly self-hosted
Speed / uptime No Meta SLA found; Kingy/OpenRouter mean: 86.7s high, 99.0s xhigh per task Not tested Product- and tier-dependent Hardware- and runtime-dependent
Deployment Hosted API / Muse Code Hosted API Hosted API / native agents Local, private cloud or third-party host
Dollars per success Kingy primary tasks: $0.060 high, $0.078 xhigh; workload-specific Unknown; data policy may dominate Kingy Gemini control: $0.050; other workloads unknown Must include infrastructure and operator time

Meta’s claim of 25% fewer tokens than 1.2 could lower cost per attempt. It does not prove lower dollars per success: if retries, interventions or failures change, the denominator changes. Kingy’s own gap between per-attempt price and accepted-work price demonstrates the point. Buyers should record first-attempt success, accepted-result rate, tool calls, retries, reasoning tokens, total wall time and cost—not only token price.

Muse Code and ecosystem considerations

Muse Code was co-trained with Spark 1.2 and uses persistent background agents plus a replayable local event log. That tight pairing can improve the complete product while making pure-model comparisons harder. [OFFICIAL—META] Muse Code and 1.2 launch.

For a buyer, the ecosystem question is pragmatic: does Muse Code handle repositories, approvals, interrupts, compaction, failures and user steering better than your existing agent? Meta’s 1.3 demos show ambitious artifacts and long tasks, but demos are selected examples, not measured reliability. Compare Muse Code to Codex, Claude Code and other native products in a separate track labelled model plus ecosystem performance.

Privacy, safety and operational risk

Standard is the only defensible choice for confidential work based on the published tier descriptions; Contributor explicitly grants training use. Even Standard needs contract review where retention, regional processing, regulated data or incident response matters. [OFFICIAL—META] / [UNKNOWN]

Meta says 1.3 better resists prompt injection, recognizes irreversible actions and confirms before consequential steps. Those are the right behaviors for computer-use agents, but the launch gives no attack set, score, grader, false-positive rate or 1.3-specific safety report. Treat them as product claims. [OFFICIAL—META] vendor-reported.

Operational tests should include malicious instructions in webpages and documents, credential-exfiltration traps, ambiguous delete/publish actions, tool-result spoofing, recovery after partial execution and proof that the model does not claim completion when the verifier fails.

Reproducing and extending the test

The expanded run used temperature 0, three trials per cell, randomized order and live provider aliases. Completion caps were 6,000 tokens for code, 5,000 for the decision memo, 3,000 per tool turn and 8,000 for the contract task. The contract packet contained 380,564 characters—about 95,609 estimated tokens. OpenRouter provider fallback was disabled. Future runs should pin model snapshots where possible and keep native-product results separate from the common harness.

Kingy retained the raw pilot and expanded JSON, corrected audit, prompts, graders and harness. The expanded plan hash is ece68c83846c83affa7e00f8b60f93b0e9dc20ff94a873b62f786e8833639bb7; the immutable raw expanded-result hash is 1b7c72921b8c176fcb35cfbfc580424ae6cf6a5e98cb7c7bd16753c053d9189c.

The next useful extensions are a real repository migration, interrupted-task recovery, modality-grounded PDF/video work and a larger per-cell sample. Each should retain programmatic end-state checks, injected failures, reasoning-token accounting and cost per accepted result.

Best model by use case

Use case Start with Why / caveat
Low-cost agentic API pilot Muse Spark 1.3 high or Gemini 3.8 Flash Spark was stronger in Kingy’s primary acceptance test; Gemini was faster and cheaper per accepted result
Hardest Claude-native long agent Claude Fable 5.1 Current Anthropic ceiling; expensive and absent from Meta’s table
Codex/OpenAI professional work GPT-5.6 Sol Mature ecosystem and selectable max; higher price
Repository coding challenger Muse Spark 1.3 high Two of three Kingy patches passed hidden tests; xhigh returned no accepted patch under the shared cap
Audio-heavy multimodal work Gemini 3.8 Flash or Spark 1.2 Meta warns that 1.3 audio may degrade
Local/offline personal agent Muse Glimmer 30B Apache 2.0 and a realistic 24–32GB envelope
Data-center open-weight frontier DeepSeek V4 Pro, Kimi K3, GLM-5.3, MiniMax M3 or Qwen3.8 Choose by workload and license; none is a plug-in consumer-GPU model
Sensitive prototyping Spark Standard, not Contributor Standard excludes training use; retention/residency still require verification
Routine high-volume formatting A cheaper workhorse that passes your tests Frontier reasoning can waste tokens and latency

Who should use it, who should wait

Use it now in a bounded pilot if you build tool agents, professional artifact workflows, browser operators or long-context systems and can log every attempt. Start with high, not xhigh, unless the task has enough completion headroom for extended hidden reasoning.

Wait for clarification if your purchase depends on max reasoning, precise reproducibility, audio, retention/residency terms or a safety case for irreversible computer actions. Meta can resolve several of these gaps with documentation rather than a new model.

Choose another model if your validated workflows already perform better in Claude Code, Codex or Gemini; if Fable’s higher ceiling repays its cost; or if local control is mandatory. Spark 1.3 has no downloadable weights today.

Unresolved questions

Unknown Why it matters Evidence needed
Is launch “max” the API’s xhigh, or an unreleased mode? Reproducibility and procurement claims Meta correction plus exact request configuration
Which exact model snapshot generated each score? Silent updates can move results Snapshot IDs and evaluation dates
Are 1.3 weights actually planned, and under what license? Local deployment and commercial freedom Version, date, repository and license text
What is maximum output length? Long reports, code generation and reasoning budgets Model catalog/API limit
What are retention and regional-processing terms? Security, privacy and regulated workloads Contractual data-processing documentation
Does the native Meta API or Muse Code match Kingy’s OpenRouter common-harness result? Separates routing and product effects Same-task native-product reproduction
What is population-level repeated-run variance? Three trials expose inconsistency but cannot estimate a stable rate Larger per-task samples and confidence intervals
When will audio be fully supported? Multimodal product reliability Changelog and audio evaluation
Does 1.3 power Meta AI consumer products? Product availability and behavior Named product/model deployment notice
What are TTFT, throughput, uptime and cost per accepted task? Capacity and economics Public measurements or buyer workload tests

Final verdict

Muse Spark 1.3 belongs on the frontier shortlist. Meta’s evidence shows a genuine generational jump, and Kingy’s test found a real operational strength: 6/6 successful complex tool workflows across high and xhigh. The Standard API also undercuts flagship OpenAI and Anthropic models by a wide margin. The Contributor tier is cheaper still, but its training-data trade is unsuitable for confidential prompts.

The launch is less definitive than the scorecard looks. Kingy’s xhigh coding failures show that more reasoning can produce less usable output under a shared completion cap. Spark 1.3 was also much slower than Gemini and Spark 1.2 in this test window. Separately, Meta’s max/xhigh contradiction is unresolved; OSWorld revisions differ; Terminal-Bench measures agent products; the private IF Index cannot be audited; and the launch comparison omits newer peers.

Recommendation: pilot Spark 1.3 now at high reasoning for tool-heavy work, protect visible-output budget and measure accepted results rather than attempts. Do not assume xhigh is better, and do not migrate until the model wins your own success-normalized workload test. Confidence: medium.

Methodology and primary sources

Kingy transcribed Meta’s launch image, checked every value visually, recalculated all 1.3-versus-1.2 deltas and separated percentage points from relative percentage changes. We traced task counts, splits, harnesses and metrics through Meta’s methodology and benchmark-owner pages. We compared competitor specifications only where direct vendor documentation was available and preserved license labels from the model repositories. We did not create a synthetic score from unrelated vendor benchmarks.

For hands-on testing, deterministic graders checked executable code, durable tool state, calculations, citations and long-context amendments. Raw responses were preserved before audit. Two false-negative grader defects were corrected additively, and one ambiguous decision item was excluded from the primary quality aggregate. Search results, aggregators and practitioner posts were not used as substitutes for missing official specifications.

FAQ

How did Muse Spark 1.3 perform in Kingy’s real world performance test?

At high reasoning it accepted 8/9 results across executable code, stateful tool use and 96K contract extraction, versus 6/9 for 1.3 xhigh, 1.2 xhigh and Gemini 3.8 Flash high. Spark 1.3 was 6/6 on tool workflows, but all three xhigh code attempts exhausted the completion allowance without returning an accepted artifact.

Is Muse Spark 1.3 available now?

Yes in Muse Code and Meta Model API. The official API IDs are muse-spark-1.3 and muse-spark-1.3-contributor. Meta has not officially disclosed whether 1.3 specifically powers Meta AI or its other consumer products.

How much does Muse Spark 1.3 cost?

Standard costs $1.25 per million uncached input tokens, $0.15 cached input and $4.25 output. Contributor costs $0.10, $0.002 and $0.20 respectively, but allows Meta to use prompts and completions to train future models.

Does Muse Spark 1.3 have a one-million-token context window?

Yes: Meta documents 1,048,576 tokens. Meta’s 98.1 score in the 512K–1M MRCR band is vendor-reported evidence of retrieval, not proof of reliable million-token synthesis.

Is max reasoning available?

No documented API value named max was available at the cutoff. The reasoning page tops out at xhigh; the launch says max is coming after safety testing. Meta’s scorecard nevertheless labels 1.3 “max,” while its methodology alternates between xhigh and max.

Is Muse Spark 1.3 open source or open weight?

No. Meta promises a future “Muse Spark open weights release” but has not named a version, date or license. Muse Glimmer 30B is the separate downloadable Apache 2.0 sibling available today.