Launch analysis · Sources checked October 5, 2026 · Kingy AI has not independently tested Beam.
Reflection Beam is a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and tool-using agents. Reflection announced its first model on October 5 with selective preview access and an Apache 2.0 weights release planned later this month. See the official announcement.
The appeal is practical: a large pool of model capacity with a relatively small amount activated for each token. That could make capable coding agents cheaper to serve. Whether it does depends on the complete job, including reading context, running tools, retrying failed work and reviewing the result.
Business context: Our Reflection AI factory strategy analysis examines the Korean sovereign-cloud proposal, enterprise partnerships and operating economics behind the company’s model roadmap.
Active is about 4.6% of total.
Input and output share this limit.
Weights planned for October.
Reflection Beam specs and availability
| Item | Verified reporting detail |
|---|---|
| Developer | Reflection AI |
| API model identifier | Beam-501B-A23B |
| Architecture | Sparse mixture of experts (MoE) |
| Total / active parameters | 501B / 23B |
| Primary focus | Coding, reasoning and agentic tasks |
| Input and output | Text; native image, audio and file inputs are not supported by the documented API |
| Model context claim | 1M tokens after midtraining |
| Documented beta API context | 262,144 tokens, counting input and generated output together |
| Documented maximum API output | 131,072 tokens, subject to the combined context limit |
| Documented knowledge cutoff | June 30, 2026 |
| Reasoning controls | low, medium, high, xhigh, max; default medium |
| Weights | Planned for October 2026; no exact release day in the announcement |
| Intended weights license | Apache 2.0 |
| Public token pricing | No per-million-token rates in the announcement or developer pages reviewed for this guide |
The Models reference supplies the API limits and model ID; the compatibility reference covers inputs and supported features. The planned weights release should remain a future-tense claim until the checkpoint and license are available.
What the architecture tells us
| Area | What the primary sources describe |
|---|---|
| Transformer depth | 52 layers |
| Attention and routing | Interleaved local/global attention and fine-grained routed experts |
| Training numerics | FP32 residual accumulation is discussed; it does not establish the released inference precision |
| Tool integration | Function calling, tool choice and parallel tool calls are documented |
| Structured responses | JSON object and JSON Schema response formats are documented |
| Streaming | Supported by the beta API |
| Architecture still to confirm | The release configuration is needed for exact expert count, experts selected per token, hidden size and deployment-specific memory |
| Release still to confirm | Checkpoint packaging, supported quantizations, reproducible benchmark harnesses and exact release date |
In a sparse MoE, a learned router selects expert networks for individual tokens. Total parameters describe the full weight inventory; active parameters describe the subset used at a step. Hugging Face’s MoE explanation covers this compute-versus-memory tradeoff.
Our calculation, 23 ÷ 501, gives about 4.6%. That ratio describes parameter activation. Speed also depends on attention, memory traffic, routing, batching and communication between devices. Storage planning still starts with the full 501B weights.
The context window has two different numbers
Reflection says midtraining extended Beam’s effective context to 1 million tokens. The hosted beta currently documents 262,144 combined input/output tokens, with generation capped at 131,072 tokens within that same window. Those are different specifications.
For example, a 200,000-token prompt leaves at most 62,144 tokens within the documented combined limit. It does not leave another 128K on top. Reserve room for reasoning, the answer and conversation history, and check the metadata returned for your available model before building around a fixed allowance.
Beam benchmarks against GLM, Qwen, Kimi and DeepSeek
These are selected vendor-reported scores from Reflection’s launch table, checked October 5. They show where Beam sits in that published comparison; Kingy has not rerun the evaluations. All values below are reported percentage scores, with higher being better. NR means the source table did not report a value.

| Benchmark | Beam | GLM 5.2 | GLM 5.3 | Qwen 3.8 Max | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | 61.0 | 51.0 | 68.0 | 74.2 |
| SWE Bench Pro v1 | 65.5 | 62.1 | NR | 67.7 | NR | NR |
| SWE Bench Pro v2-Hard | 77.2 | NR | 84.3 | NR | 88.2 | NR |
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.2 | 86.6 | 88.3 | 90.6 |
| SWE Bench Verified | 80.9 | NR | NR | NR | NR | NR |
| HLE, no tools | 36.2 | 40.5 | 42.3 | 43.6 | 46.9 | 39.1 |
Against GLM 5.2, Beam’s result is mixed. It leads by 3.4 percentage points on SWE Bench Pro v1 and 0.4 points on DeepSWE, while trailing by 0.9 points on Terminal Bench and 4.3 points on HLE without tools. Small gaps deserve restraint when run counts and variance are unavailable.
Qwen 3.8 Max is ahead on the four rows where both models have a reported value. GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash also lead Beam across their available rows here. The gap on DeepSWE is substantial: Beam’s 44.4 compares with Kimi’s 68.0 and DeepSeek’s 74.2. A team whose difficult repository work resembles those tasks should evaluate that gap alongside any serving savings.
Beam versus Inkling and Nemotron 3 Ultra
| Benchmark | Beam | Inkling | Nemotron 3 Ultra |
|---|---|---|---|
| SWE Bench Pro v1 | 65.5 | 54.3 | 46.4 |
| Terminal Bench v2.1 | 80.1 | 63.8 | 56.4 |
| SWE Bench Verified | 80.9 | 77.6 | 70.7 |
| HLE, no tools | 36.2 | 29.7 | 26.7 |
Beam leads Inkling and Nemotron 3 Ultra on all four of these displayed rows. That supports Reflection’s claim of advancing the Western open-weight field in this comparison. It does not establish a lead over the broader open-model field.
Reasoning, tool use and long-context results
| Evaluation | Beam | GLM 5.2 | Qwen 3.8 Max | Kimi K3 |
|---|---|---|---|---|
| AIME 2026 | 97.8 | 99.2 | NR | NR |
| GPQA Diamond | 90.5 | 91.2 | 92.6 | 93.5 |
| SciCode | 49.7 | NR | 52.1 | 58.7 |
| AutomationBench public | 37.0 | 26.2 | 39.8 | 46.7 |
| MCP Atlas | 78.7 | 77.8 | 84.5 | 82.3 |
| tau3 banking | 38.0 | 37.1 | 55.2 | 37.1 |
| AA-LCR | 79.3 | 78.3 | 80.3 | 88.7 |
| LongBench v2 | 65.5 | 64.0 | 66.3 | NR |
| IFBench | 79.7 | 73.3 | 82.8 | NR |
Beam leads GLM 5.2 on AutomationBench, MCP Atlas, tau3 banking, AA-LCR, LongBench v2 and IFBench in these selected rows. Qwen and Kimi set higher targets on several of the same evaluations. Treat math, tool use, long-context reasoning and instruction following as separate requirements; a strong patch-writing score cannot answer every deployment question.
Why the source and harness matter
There is a concrete reporting discrepancy: Reflection’s table lists 51.0 for Qwen on DeepSWE v1.1, while the Qwen model card reports 56.6 for DeepSWE 1.1. Qwen says it reports its highest result across Claude Code and mini-SWE-agent. The pages do not explain why Reflection uses 51.0. We preserve Reflection’s value in its comparison rather than silently substituting the higher score.
SWE-bench Verified contains 500 human-validated instances, and its maintainers distinguish comparisons of complete coding systems from more controlled model comparisons. The agent harness, tools and retry budget all affect the outcome. Terminal-Bench 2.1 revised 28 of version 2.0’s 89 tasks, so its version label matters too.
Before treating two scores as a head-to-head result, match the task version, split, harness, reasoning effort, token budget and available tools. A “no tools” HLE score describes a different setup from HLE with search or code execution. Averaging unrelated raw scores would hide those distinctions.
Model size and multimodal capability
| Model or release | Total parameters | Active parameters | Context and modality distinction |
|---|---|---|---|
| Beam | 501B | 23B | Text only; 1M training claim, 256K documented beta API |
| Inkling | 975B | 41B | Text, image and audio inputs; model capability up to 1M |
| Qwen3.8-2.4T-A95B / hosted Qwen3.8-Max | 2.4T for the open checkpoint | 95B for the open checkpoint | Checkpoint: text only, 256K native with extension; hosted Max adds vision and a default 1M context |
| Kimi K3 | 2.8T | Not specified in the primary announcement checked here | Native vision and a 1M context claim |
Sources: the Inkling model card, Qwen checkpoint card and Moonshot’s Kimi K3 announcement. The hosted Qwen3.8-Max service adds features to the text-only checkpoint, including vision and a default 1M context. Keep the checkpoint and hosted product specifications separate.
Our arithmetic puts Beam at about 56% of Inkling’s active parameter count and 24% of the Qwen checkpoint’s. These ratios help explain Reflection’s efficiency thesis. They do not measure throughput or account for the extra modalities offered by competitors.
For GPT or Claude, the sources used here do not provide a controlled Beam comparison against the current systems. Run the same workload through both complete agents and compare acceptance, latency, tool behavior, modalities and total cost. Add hosting choice and weight adaptation when those affect your application.
What Beam’s efficiency claim means for pricing
Reflection estimates 3–4× less generation compute than GLM 5.2 on advanced reasoning comparisons. Its published method approximates forward-pass work as 2 × active parameters × mean generated tokens per attempt, counting reasoning and final-answer tokens. It excludes prompt prefill, context-dependent attention and serving overhead.
That measures an approximation of computational work. A provider’s token price, elapsed time and cost per completed job also depend on hardware, utilization, input processing, tool calls and failed attempts. A model with half the active parameters but twice the generated tokens would tie under the simplified formula; that is an illustration, not a measured Beam result.
No public per-million-token rates appeared in the launch announcement or developer pages reviewed for this guide. Confirm your account’s pricing and allowance on the Reflection platform. Until rates or infrastructure measurements are available, the compute estimate cannot establish a dollar saving.
Training details behind the efficiency argument
Reflection reports 23.8 trillion pretraining tokens and more than 100 million reinforcement-learning rollouts on 10,500 NVIDIA GB300 GPUs over four weeks. It describes asynchronous training and a length penalty that rewards successful solutions while discouraging unnecessary tokens. These are company-reported details from the launch post.
The useful behavior to test is completion. A shorter answer that leaves a patch unfinished can create another call and more human work. Longer reasoning can earn its cost if it produces an accepted fix with fewer retries. Training scale alone cannot tell a buyer which outcome Beam will deliver.
How to try Beam through the beta API
Start with the Reflection platform and official quickstart. New sign-ups join a waitlist; API keys become available after access is enabled. The platform is in beta.
The documented compatible base URL is https://api.reflection.ai/openai/v1. It supports Chat Completions and Models. Reflection’s compatibility guide explicitly excludes Responses, Embeddings, Images, Audio, Files, Batch and Assistants endpoints. Image, audio and file message inputs are also unsupported.
The Python example below uses the standard library. Set REFLECTION_API_KEY in your environment first. We checked the syntax and request fields against the documentation; we have not executed it against an authenticated Beam account.
import json
import os
import urllib.request
payload = {
"model": "Beam-501B-A23B",
"reasoning_effort": "medium",
"max_completion_tokens": 4096,
"messages": [{
"role": "user",
"content": "Write a CSV validator and explain its edge cases."
}]
}
request = urllib.request.Request(
"https://api.reflection.ai/openai/v1/chat/completions",
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": "Bearer " + os.environ["REFLECTION_API_KEY"],
"Content-Type": "application/json"
},
method="POST"
)
with urllib.request.urlopen(request, timeout=120) as response:
result = json.load(response)
choice = result["choices"][0]
answer = choice["message"].get("content")
if not answer:
raise RuntimeError(
"No final answer. Finish reason: " + str(choice.get("finish_reason"))
+ ". If length-limited, increase the completion budget or lower effort."
)
print(answer)
print("Finish reason:", choice.get("finish_reason"))
print("Usage:", result.get("usage"))
This is a connection test. Once it works, move to a real task with defined acceptance criteria and preserve the prompt, output and usage. A plausible answer to a toy prompt is weak evidence about repository work.
Reasoning effort and integration traps
Beam accepts low, medium, high, xhigh and max, with medium as the default. It always reasons; none and minimal are unsupported. Start at medium, then compare effort settings on the same tasks, as the Reasoning guide recommends.
- Reserve completion tokens for reasoning and the answer. Both use
max_completion_tokens. Exhausting it during reasoning can returnfinish_reason: "length"with no final content. - Preserve tool-call messages. In a multi-turn tool loop, return the assistant message with its
tool_callsandreasoning_contentunchanged. - Audit framework defaults. Parameters outside the supported reference can cause a 400 error. The accepted
stopparameter has no effect;nmust be 1. - Limit concurrency and honor retry guidance. Request, token and concurrency limits are shared across the organization. Follow
Retry-After, and respectx-should-retry: falsewhen a daily allowance is spent.
Sources: the compatibility guide, reasoning guide and rate-limit documentation. Check the coding-agent quickstart for supported agent setup.
Beam hardware requirements start with all 501B weights
The 23B-active count does not turn Beam into a 23B model for storage. These are our calculations for a uniform raw weight payload, before runtime overhead. They are not released checkpoint sizes or official deployment requirements.
| Weight precision | Calculation | Raw payload, decimal GB | Raw payload, GiB |
|---|---|---|---|
| 16-bit | 501B × 2 bytes | 1,002 | ≈933.2 |
| 8-bit | 501B × 1 byte | 501 | ≈466.6 |
| 4-bit | 501B × 0.5 bytes | 250.5 | ≈233.3 |
Add quantization metadata, buffers, activations, the context cache and framework overhead. Packaging can mix precisions, and quantization needs its own quality tests. A 128 GB workstation cannot hold the full approximately 250.5 GB raw payload at ordinary 4-bit storage.
Distributed serving or offloading can change where those weights sit, while adding bandwidth and latency tradeoffs. Before buying hardware, require a released checkpoint, a supported inference engine and quantization format, per-device memory requirements, and measured throughput at your expected context and concurrency.
What an open-weight release would give developers
The announced Apache 2.0 release could give developers a choice of hosting and a checkpoint they can inspect and adapt. Those benefits depend on what ships. Check the actual license, configuration, tokenizer, serving instructions and evaluation recipe when the artifacts arrive.
Open weights also do not establish full training reproducibility. The Open Source Initiative’s AI definition considers data information, code and parameters. Use the released materials to judge which parts of Beam’s pipeline are inspectable and reproducible.
How to decide whether Beam belongs in your stack
Our recommendation is a bounded pilot for teams with text-based coding or tool workflows. Beam’s reported results justify evaluation. Teams that require native image or audio input need another model or a separate conversion layer. Teams choosing primarily for the strongest published scores should test the capability gaps shown above before making an efficiency tradeoff.
Begin with 20–30 diverse tasks from your actual workload: real bug fixes, a small interface migration, a CSV transformation and documentation of an unfamiliar module. This is an engineering screen, not a statistically definitive ranking. Increase the sample when results are close or mistakes are expensive.
- Define acceptance first. A bug fix must pass the original failing test and relevant regression checks. A data transformation must match reference counts and edge cases. Research claims need traceable sources.
- Keep the environment comparable. Use the same repository snapshot, tools, permissions and time budget where feasible. Record any differences between systems.
- Record every attempt. Save effort, input and output usage, elapsed time, tool calls, retries, the final diff or artifact, and the reason for acceptance or rejection.
- Compare cost per accepted result. Include unsuccessful attempts, model calls and sandbox or tool charges. Track human review time separately or state the labor-cost assumption.
- Inspect failures before expanding the pilot. Separate incomplete work, incorrect outputs, fabricated evidence, excessive changes and infrastructure failures.
Cost per accepted result = total trial cost ÷ accepted results. As a hypothetical example, $0.80 per attempt at 40% acceptance averages $2.00 per accepted result; $1.20 at 80% acceptance averages $1.50. These are illustrative numbers, unrelated to Beam’s rates, and exclude review and tool costs.
Keep initial agent tasks isolated and review the action logs and final diff. Reflection’s training and safety descriptions provide context, but only your trial can establish whether its behavior fits your tool environment.
A released checkpoint, independent evaluations with documented budgets, and measured price, latency and acceptance under load would make the adoption decision firmer. Until then, use a coding pilot to test the specific bargain Beam offers: enough capability for your work at a lower delivered cost.
Frequently asked questions
Can I download Reflection Beam now?
The October 5 announcement describes a preview and a weights release planned later in October 2026. It gives no exact release day. Check Reflection’s release artifacts before treating the weights as available.
Is Beam’s context window 1M or 256K tokens?
Reflection reports 1M tokens after midtraining. Its beta API documentation lists 262,144 tokens shared by input and generated output. Use the hosted model’s current metadata for application planning.
Does Beam support images, audio or files?
The documented API accepts text. Native image, audio and file inputs are unsupported. An agent can use separate tools that convert other inputs to text; that capability belongs to the surrounding system.
Can I run Beam on a gaming PC?
A uniform 4-bit raw weight payload is approximately 250.5 GB before overhead. The 23B-active count reduces the computation used per token, while storage still depends on the full weights. Supported serving configurations have to be assessed after release.
Is Beam better than Qwen, Kimi or GLM?
Reflection’s reported results are mixed against GLM 5.2 and trail the stronger competitors on several displayed tasks. Beam leads Inkling and Nemotron 3 Ultra on the four comparison rows above. There is no universal winner across workloads.
What does the Beam API cost?
The launch announcement and developer pages reviewed for this guide do not publish per-million-token rates. Confirm platform pricing and your account’s allowances. Estimated generation compute does not establish an API price.
Sources and reporting notes
Checked October 5, 2026. This article is launch analysis. Benchmark values and training details are vendor reports; parameter ratios, memory payloads and illustrative economics are identified as Kingy calculations or hypothetical examples. No hands-on Beam performance trial was conducted. The API example was syntax-checked without making a paid or authenticated request.
- Reflection’s Beam announcement: specifications, benchmark tables, compute method, training and planned weights release.
- Reflection Models reference, quickstart, compatibility guide, reasoning guide and rate-limit reference: API access, limits, fields and behavior.
- Inkling model card, Qwen checkpoint card and Kimi K3 announcement: comparator size, modality and context specifications; Qwen’s separately reported DeepSWE result.
- SWE-bench Verified and Terminal-Bench 2.1: benchmark-maintainer methodology and version changes.
- Hugging Face’s MoE guide and the Open Source AI Definition: architecture and openness context.
Explore more on Kingy: AI models, AI guides and reviews and tests.
The Kingy Brief
Get future Kingy Brief editions.
Source-checked AI changes, original tests and one practical thing to try.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
Regular sending is paused; no restart date is set.
