AI Tools

Reflection Beam: Specs, Benchmarks, API Limits and Our Verdict

Launch analysis · Sources checked October 5, 2026 · Kingy AI has not independently tested Beam.

Reflection Beam is a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and tool-using agents. Reflection announced its first model on October 5 with selective preview access and an Apache 2.0 weights release planned later this month. See the official announcement.

The appeal is practical: a large pool of model capacity with a relatively small amount activated for each token. That could make capable coding agents cheaper to serve. Whether it does depends on the complete job, including reading context, running tools, retrying failed work and reviewing the result.

Business context: Our Reflection AI factory strategy analysis examines the Korean sovereign-cloud proposal, enterprise partnerships and operating economics behind the company’s model roadmap.

501B / 23BTotal / active parameters.
Active is about 4.6% of total.
256KDocumented beta API context.
Input and output share this limit.
PreviewEarly access through a waitlist.
Weights planned for October.

Reflection Beam specs and availability

ItemVerified reporting detail
DeveloperReflection AI
API model identifierBeam-501B-A23B
ArchitectureSparse mixture of experts (MoE)
Total / active parameters501B / 23B
Primary focusCoding, reasoning and agentic tasks
Input and outputText; native image, audio and file inputs are not supported by the documented API
Model context claim1M tokens after midtraining
Documented beta API context262,144 tokens, counting input and generated output together
Documented maximum API output131,072 tokens, subject to the combined context limit
Documented knowledge cutoffJune 30, 2026
Reasoning controlslow, medium, high, xhigh, max; default medium
WeightsPlanned for October 2026; no exact release day in the announcement
Intended weights licenseApache 2.0
Public token pricingNo per-million-token rates in the announcement or developer pages reviewed for this guide
Announcement and developer documentation checked October 5, 2026. Hosted limits may change during the beta.

The Models reference supplies the API limits and model ID; the compatibility reference covers inputs and supported features. The planned weights release should remain a future-tense claim until the checkpoint and license are available.

What the architecture tells us

AreaWhat the primary sources describe
Transformer depth52 layers
Attention and routingInterleaved local/global attention and fine-grained routed experts
Training numericsFP32 residual accumulation is discussed; it does not establish the released inference precision
Tool integrationFunction calling, tool choice and parallel tool calls are documented
Structured responsesJSON object and JSON Schema response formats are documented
StreamingSupported by the beta API
Architecture still to confirmThe release configuration is needed for exact expert count, experts selected per token, hidden size and deployment-specific memory
Release still to confirmCheckpoint packaging, supported quantizations, reproducible benchmark harnesses and exact release date
Source: Reflection’s technical launch discussion and API compatibility documentation.

In a sparse MoE, a learned router selects expert networks for individual tokens. Total parameters describe the full weight inventory; active parameters describe the subset used at a step. Hugging Face’s MoE explanation covers this compute-versus-memory tradeoff.

Our calculation, 23 ÷ 501, gives about 4.6%. That ratio describes parameter activation. Speed also depends on attention, memory traffic, routing, batching and communication between devices. Storage planning still starts with the full 501B weights.

The context window has two different numbers

Reflection says midtraining extended Beam’s effective context to 1 million tokens. The hosted beta currently documents 262,144 combined input/output tokens, with generation capped at 131,072 tokens within that same window. Those are different specifications.

For example, a 200,000-token prompt leaves at most 62,144 tokens within the documented combined limit. It does not leave another 128K on top. Reserve room for reasoning, the answer and conversation history, and check the metadata returned for your available model before building around a fixed allowance.

Beam benchmarks against GLM, Qwen, Kimi and DeepSeek

These are selected vendor-reported scores from Reflection’s launch table, checked October 5. They show where Beam sits in that published comparison; Kingy has not rerun the evaluations. All values below are reported percentage scores, with higher being better. NR means the source table did not report a value.

Terminal Bench v2.1 reported scores: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM 5.3 88.2, Qwen 3.8 Max 86.6, GLM 5.2 81.0, Beam 80.1, Inkling 63.8 and Nemotron 3 Ultra 56.4.
Kingy visualization of Reflection’s launch comparison. Vendor-reported scores; this is a single benchmark, not an overall model ranking.
BenchmarkBeamGLM 5.2GLM 5.3Qwen 3.8 MaxKimi K3DeepSeek V4.1 Flash
DeepSWE v1.144.444.061.051.068.074.2
SWE Bench Pro v165.562.1NR67.7NRNR
SWE Bench Pro v2-Hard77.2NR84.3NR88.2NR
Terminal Bench v2.180.181.088.286.688.390.6
SWE Bench Verified80.9NRNRNRNRNR
HLE, no tools36.240.542.343.646.939.1
Source: Reflection’s October 5 launch comparison. NR = not reported. Preserve each benchmark’s exact version and setup.

Against GLM 5.2, Beam’s result is mixed. It leads by 3.4 percentage points on SWE Bench Pro v1 and 0.4 points on DeepSWE, while trailing by 0.9 points on Terminal Bench and 4.3 points on HLE without tools. Small gaps deserve restraint when run counts and variance are unavailable.

Qwen 3.8 Max is ahead on the four rows where both models have a reported value. GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash also lead Beam across their available rows here. The gap on DeepSWE is substantial: Beam’s 44.4 compares with Kimi’s 68.0 and DeepSeek’s 74.2. A team whose difficult repository work resembles those tasks should evaluate that gap alongside any serving savings.

Beam versus Inkling and Nemotron 3 Ultra

BenchmarkBeamInklingNemotron 3 Ultra
SWE Bench Pro v165.554.346.4
Terminal Bench v2.180.163.856.4
SWE Bench Verified80.977.670.7
HLE, no tools36.229.726.7
Source: the same Reflection launch table. Reported percentage scores; higher is better.

Beam leads Inkling and Nemotron 3 Ultra on all four of these displayed rows. That supports Reflection’s claim of advancing the Western open-weight field in this comparison. It does not establish a lead over the broader open-model field.

Reasoning, tool use and long-context results

EvaluationBeamGLM 5.2Qwen 3.8 MaxKimi K3
AIME 202697.899.2NRNR
GPQA Diamond90.591.292.693.5
SciCode49.7NR52.158.7
AutomationBench public37.026.239.846.7
MCP Atlas78.777.884.582.3
tau3 banking38.037.155.237.1
AA-LCR79.378.380.388.7
LongBench v265.564.066.3NR
IFBench79.773.382.8NR
Selected percentage scores from Reflection’s reasoning, tool-calling and general-capability tables. NR = not reported.

Beam leads GLM 5.2 on AutomationBench, MCP Atlas, tau3 banking, AA-LCR, LongBench v2 and IFBench in these selected rows. Qwen and Kimi set higher targets on several of the same evaluations. Treat math, tool use, long-context reasoning and instruction following as separate requirements; a strong patch-writing score cannot answer every deployment question.

Why the source and harness matter

There is a concrete reporting discrepancy: Reflection’s table lists 51.0 for Qwen on DeepSWE v1.1, while the Qwen model card reports 56.6 for DeepSWE 1.1. Qwen says it reports its highest result across Claude Code and mini-SWE-agent. The pages do not explain why Reflection uses 51.0. We preserve Reflection’s value in its comparison rather than silently substituting the higher score.

SWE-bench Verified contains 500 human-validated instances, and its maintainers distinguish comparisons of complete coding systems from more controlled model comparisons. The agent harness, tools and retry budget all affect the outcome. Terminal-Bench 2.1 revised 28 of version 2.0’s 89 tasks, so its version label matters too.

Before treating two scores as a head-to-head result, match the task version, split, harness, reasoning effort, token budget and available tools. A “no tools” HLE score describes a different setup from HLE with search or code execution. Averaging unrelated raw scores would hide those distinctions.

Model size and multimodal capability

Model or releaseTotal parametersActive parametersContext and modality distinction
Beam501B23BText only; 1M training claim, 256K documented beta API
Inkling975B41BText, image and audio inputs; model capability up to 1M
Qwen3.8-2.4T-A95B / hosted Qwen3.8-Max2.4T for the open checkpoint95B for the open checkpointCheckpoint: text only, 256K native with extension; hosted Max adds vision and a default 1M context
Kimi K32.8TNot specified in the primary announcement checked hereNative vision and a 1M context claim
Model and product specifications from their respective primary sources. A model capability is not a guarantee of identical limits across hosts.

Sources: the Inkling model card, Qwen checkpoint card and Moonshot’s Kimi K3 announcement. The hosted Qwen3.8-Max service adds features to the text-only checkpoint, including vision and a default 1M context. Keep the checkpoint and hosted product specifications separate.

Our arithmetic puts Beam at about 56% of Inkling’s active parameter count and 24% of the Qwen checkpoint’s. These ratios help explain Reflection’s efficiency thesis. They do not measure throughput or account for the extra modalities offered by competitors.

For GPT or Claude, the sources used here do not provide a controlled Beam comparison against the current systems. Run the same workload through both complete agents and compare acceptance, latency, tool behavior, modalities and total cost. Add hosting choice and weight adaptation when those affect your application.

What Beam’s efficiency claim means for pricing

Reflection estimates 3–4× less generation compute than GLM 5.2 on advanced reasoning comparisons. Its published method approximates forward-pass work as 2 × active parameters × mean generated tokens per attempt, counting reasoning and final-answer tokens. It excludes prompt prefill, context-dependent attention and serving overhead.

That measures an approximation of computational work. A provider’s token price, elapsed time and cost per completed job also depend on hardware, utilization, input processing, tool calls and failed attempts. A model with half the active parameters but twice the generated tokens would tie under the simplified formula; that is an illustration, not a measured Beam result.

No public per-million-token rates appeared in the launch announcement or developer pages reviewed for this guide. Confirm your account’s pricing and allowance on the Reflection platform. Until rates or infrastructure measurements are available, the compute estimate cannot establish a dollar saving.

Training details behind the efficiency argument

Reflection reports 23.8 trillion pretraining tokens and more than 100 million reinforcement-learning rollouts on 10,500 NVIDIA GB300 GPUs over four weeks. It describes asynchronous training and a length penalty that rewards successful solutions while discouraging unnecessary tokens. These are company-reported details from the launch post.

The useful behavior to test is completion. A shorter answer that leaves a patch unfinished can create another call and more human work. Longer reasoning can earn its cost if it produces an accepted fix with fewer retries. Training scale alone cannot tell a buyer which outcome Beam will deliver.

How to try Beam through the beta API

Start with the Reflection platform and official quickstart. New sign-ups join a waitlist; API keys become available after access is enabled. The platform is in beta.

The documented compatible base URL is https://api.reflection.ai/openai/v1. It supports Chat Completions and Models. Reflection’s compatibility guide explicitly excludes Responses, Embeddings, Images, Audio, Files, Batch and Assistants endpoints. Image, audio and file message inputs are also unsupported.

The Python example below uses the standard library. Set REFLECTION_API_KEY in your environment first. We checked the syntax and request fields against the documentation; we have not executed it against an authenticated Beam account.

import json
import os
import urllib.request

payload = {
    "model": "Beam-501B-A23B",
    "reasoning_effort": "medium",
    "max_completion_tokens": 4096,
    "messages": [{
        "role": "user",
        "content": "Write a CSV validator and explain its edge cases."
    }]
}

request = urllib.request.Request(
    "https://api.reflection.ai/openai/v1/chat/completions",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "Authorization": "Bearer " + os.environ["REFLECTION_API_KEY"],
        "Content-Type": "application/json"
    },
    method="POST"
)

with urllib.request.urlopen(request, timeout=120) as response:
    result = json.load(response)

choice = result["choices"][0]
answer = choice["message"].get("content")
if not answer:
    raise RuntimeError(
        "No final answer. Finish reason: " + str(choice.get("finish_reason"))
        + ". If length-limited, increase the completion budget or lower effort."
    )

print(answer)
print("Finish reason:", choice.get("finish_reason"))
print("Usage:", result.get("usage"))

This is a connection test. Once it works, move to a real task with defined acceptance criteria and preserve the prompt, output and usage. A plausible answer to a toy prompt is weak evidence about repository work.

Reasoning effort and integration traps

Beam accepts low, medium, high, xhigh and max, with medium as the default. It always reasons; none and minimal are unsupported. Start at medium, then compare effort settings on the same tasks, as the Reasoning guide recommends.

  • Reserve completion tokens for reasoning and the answer. Both use max_completion_tokens. Exhausting it during reasoning can return finish_reason: "length" with no final content.
  • Preserve tool-call messages. In a multi-turn tool loop, return the assistant message with its tool_calls and reasoning_content unchanged.
  • Audit framework defaults. Parameters outside the supported reference can cause a 400 error. The accepted stop parameter has no effect; n must be 1.
  • Limit concurrency and honor retry guidance. Request, token and concurrency limits are shared across the organization. Follow Retry-After, and respect x-should-retry: false when a daily allowance is spent.

Sources: the compatibility guide, reasoning guide and rate-limit documentation. Check the coding-agent quickstart for supported agent setup.

Beam hardware requirements start with all 501B weights

The 23B-active count does not turn Beam into a 23B model for storage. These are our calculations for a uniform raw weight payload, before runtime overhead. They are not released checkpoint sizes or official deployment requirements.

Weight precisionCalculationRaw payload, decimal GBRaw payload, GiB
16-bit501B × 2 bytes1,002≈933.2
8-bit501B × 1 byte501≈466.6
4-bit501B × 0.5 bytes250.5≈233.3
Kingy calculations: 501 billion parameters × bytes per parameter. Decimal GB and binary GiB are shown separately.

Add quantization metadata, buffers, activations, the context cache and framework overhead. Packaging can mix precisions, and quantization needs its own quality tests. A 128 GB workstation cannot hold the full approximately 250.5 GB raw payload at ordinary 4-bit storage.

Distributed serving or offloading can change where those weights sit, while adding bandwidth and latency tradeoffs. Before buying hardware, require a released checkpoint, a supported inference engine and quantization format, per-device memory requirements, and measured throughput at your expected context and concurrency.

What an open-weight release would give developers

The announced Apache 2.0 release could give developers a choice of hosting and a checkpoint they can inspect and adapt. Those benefits depend on what ships. Check the actual license, configuration, tokenizer, serving instructions and evaluation recipe when the artifacts arrive.

Open weights also do not establish full training reproducibility. The Open Source Initiative’s AI definition considers data information, code and parameters. Use the released materials to judge which parts of Beam’s pipeline are inspectable and reproducible.

How to decide whether Beam belongs in your stack

Our recommendation is a bounded pilot for teams with text-based coding or tool workflows. Beam’s reported results justify evaluation. Teams that require native image or audio input need another model or a separate conversion layer. Teams choosing primarily for the strongest published scores should test the capability gaps shown above before making an efficiency tradeoff.

Begin with 20–30 diverse tasks from your actual workload: real bug fixes, a small interface migration, a CSV transformation and documentation of an unfamiliar module. This is an engineering screen, not a statistically definitive ranking. Increase the sample when results are close or mistakes are expensive.

  1. Define acceptance first. A bug fix must pass the original failing test and relevant regression checks. A data transformation must match reference counts and edge cases. Research claims need traceable sources.
  2. Keep the environment comparable. Use the same repository snapshot, tools, permissions and time budget where feasible. Record any differences between systems.
  3. Record every attempt. Save effort, input and output usage, elapsed time, tool calls, retries, the final diff or artifact, and the reason for acceptance or rejection.
  4. Compare cost per accepted result. Include unsuccessful attempts, model calls and sandbox or tool charges. Track human review time separately or state the labor-cost assumption.
  5. Inspect failures before expanding the pilot. Separate incomplete work, incorrect outputs, fabricated evidence, excessive changes and infrastructure failures.

Cost per accepted result = total trial cost ÷ accepted results. As a hypothetical example, $0.80 per attempt at 40% acceptance averages $2.00 per accepted result; $1.20 at 80% acceptance averages $1.50. These are illustrative numbers, unrelated to Beam’s rates, and exclude review and tool costs.

Keep initial agent tasks isolated and review the action logs and final diff. Reflection’s training and safety descriptions provide context, but only your trial can establish whether its behavior fits your tool environment.

A released checkpoint, independent evaluations with documented budgets, and measured price, latency and acceptance under load would make the adoption decision firmer. Until then, use a coding pilot to test the specific bargain Beam offers: enough capability for your work at a lower delivered cost.

Frequently asked questions

Can I download Reflection Beam now?

The October 5 announcement describes a preview and a weights release planned later in October 2026. It gives no exact release day. Check Reflection’s release artifacts before treating the weights as available.

Is Beam’s context window 1M or 256K tokens?

Reflection reports 1M tokens after midtraining. Its beta API documentation lists 262,144 tokens shared by input and generated output. Use the hosted model’s current metadata for application planning.

Does Beam support images, audio or files?

The documented API accepts text. Native image, audio and file inputs are unsupported. An agent can use separate tools that convert other inputs to text; that capability belongs to the surrounding system.

Can I run Beam on a gaming PC?

A uniform 4-bit raw weight payload is approximately 250.5 GB before overhead. The 23B-active count reduces the computation used per token, while storage still depends on the full weights. Supported serving configurations have to be assessed after release.

Is Beam better than Qwen, Kimi or GLM?

Reflection’s reported results are mixed against GLM 5.2 and trail the stronger competitors on several displayed tasks. Beam leads Inkling and Nemotron 3 Ultra on the four comparison rows above. There is no universal winner across workloads.

What does the Beam API cost?

The launch announcement and developer pages reviewed for this guide do not publish per-million-token rates. Confirm platform pricing and your account’s allowances. Estimated generation compute does not establish an API price.

Sources and reporting notes

Checked October 5, 2026. This article is launch analysis. Benchmark values and training details are vendor reports; parameter ratios, memory payloads and illustrative economics are identified as Kingy calculations or hypothetical examples. No hands-on Beam performance trial was conducted. The API example was syntax-checked without making a paid or authenticated request.

Explore more on Kingy: AI models, AI guides and reviews and tests.