Reflection Beam: 501B Open-Weight MoE, Weights Due Soon
Reflection previews Beam, a 501B MoE with 23B active parameters. Apache 2.0 weights are promised later this month; benchmarks are unverified.
Reflection previews Beam, a 501B MoE with 23B active parameters. Apache 2.0 weights are promised later this month; benchmarks are unverified.
Introduction
On October 5, 2026, Reflection announced Beam, a 501B-parameter open-weight model, in a post titled "Introducing Beam: Reflection's 501B open-weight model." Beam is a sparse mixture-of-experts (MoE) text model with 23B active parameters, built for coding, reasoning, and agentic workloads. One point matters before anything else: the weights are not out yet. Reflection says "Beam is undergoing final red-teaming and evaluations," offers early access through a waitlist at platform.reflection.ai, and says it will release the weights, technical report, model card, and developer artifacts "later this month" under Apache 2.0. This article therefore treats Beam as an announced preview, and every performance figure below is Reflection's own.
Architecture and training
| Item | What Reflection reports |
|---|---|
| Design | Sparse MoE, 501B total and 23B active parameters, text-only |
| Structure | 52 layers, interleaved local and global attention, fine-grained routed experts |
| Pretraining | 23.8T tokens from web, public sources, and proprietary licensed datasets |
| Pretraining run | Under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, 92.3% goodput toward the end, nine semi-automatic rewinds |
| Context | Midtraining extends effective context length to 1M tokens |
Sparse MoE means each token activates only a small share of the model's experts. With 23B of 501B parameters active, per-token compute is closer to a mid-sized dense model than to a 500B one, while total capacity stays large.
Reinforcement learning at scale
Reflection describes a high-compute RL phase: 10.5K NVIDIA GB300 GPUs for four weeks, over 100 million rollouts, a maximum context of 256K, roughly 1.3 billion sandboxes for training and grading, and nearly one million environments, with an average of 110K concurrent rollouts. Reflection says "we believe this is one of the largest scale RL runs conducted by any open lab to date," which is a statement of belief and not an independently established ranking.
Its fully asynchronous RL stayed stable even with one-day staleness, meaning rollouts were generated by a policy 107 weight versions behind the current one. A reasoning-effort parameter lets users trade response length against performance: lower settings favor shorter answers, and higher settings allow longer reasoning on demanding tasks.
Benchmarks: Reflection's table
Reflection's coding and agentic table uses NR for scores a model did not report. The figures below are Reflection's, and TechCrunch notes the company's performance claims "haven't been independently verified."
| Benchmark | Beam | Inkling | Nemotron 3 Ultra | GLM 5.2 | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | NR | NR | 44.0 | 61.0 | 68.0 | 51.0 | 74.2 |
| Terminal Bench v2.1 | 80.1 | 63.8 | 56.4 | 81.0 | 88.2 | 88.3 | 86.6 | 90.6 |
| SWE Bench Pro v1 | 65.5 | 54.3 | 46.4 | 62.1 | NR | NR | 67.7 | NR |
| SWEBench Verified | 80.9 | 77.6 | 70.7 | NR | NR | NR | NR | NR |
| SWE Bench Pro v2-Hard | 77.2 | 56.9 | NR | NR | 84.3 | 88.2 | NR | NR |
Reading the table honestly gives a mixed picture. Beam is ahead of Inkling and Nemotron 3 Ultra wherever they report, level with GLM 5.2 on DeepSWE (44.4 versus 44.0) and ahead on SWE Bench Pro v1, but slightly behind it on Terminal Bench (80.1 versus 81.0). It trails Qwen 3.8 Max on all three shared tests, and on DeepSWE and Terminal Bench it trails several listed models by wide margins. Reflection itself says Beam is "competitive" with GLM 5.2, "approaching" Qwen 3.8-Max on coding and agentic tasks, and that Kimi K3 "remains ahead on raw capability."
The efficiency claim
Reflection's main pitch is efficiency: on advanced reasoning benchmarks it says Beam scores comparably to GLM-5.2 while using 3-4x less inference compute. Its method estimates generation compute as 2 x active parameters x mean generated tokens, and Reflection states this excludes prompt prefill, context-dependent attention, and serving overhead, so it is "an approximate compute comparison rather than measured inference cost." TechCrunch reports GLM-5.2 has about 744B total and 40B active parameters, so Beam's smaller active count is part of the arithmetic.
Safety approach
Reflection says it trained a second model for safety and alignment and merged it with the main RL model through multi-teacher on-policy distillation, using deliberative alignment. Safety evaluation results are promised in the technical report, which is not yet published, so there is nothing to check yet.
Company context
TechCrunch describes Reflection as Brooklyn-based, founded in 2024 by two former Google DeepMind researchers, and reports it has raised roughly $4.7B per PitchBook from backers including Nvidia, Sequoia, and Lightspeed, with a last round at a $25B pre-money valuation. Compute deals with SpaceX and Nebius for GB300 access through 2029 reportedly total over $7B. Reflection is aiming at enterprises and sovereign "AI factories," and TechCrunch says it is testing a sovereign AI factory partnership with South Korea's Shinsegae Group. Inkling from Thinking Machines Lab is multimodal while Beam is text-only, and Beam outscores Inkling on the four coding tests where both report.
Outlook and assessment
The real test comes when the weights, model card, and technical report ship. Independent runs of the coding and agentic benchmarks, plus a measured cost comparison instead of a FLOPs estimate, will show whether the efficiency story holds in deployment. Teams that need an Apache 2.0 licensed model with long context and agentic tuning should watch for the release Reflection has promised for this month, but should not plan around Beam until the artifacts exist and third parties have tested them.
Editor's Verdict
Reflection Beam: 501B Open-Weight MoE, Weights Due Soon is a workable proposition that fills a clear gap, even if it doesn't fundamentally change the landscape.
The strongest case for paying attention: the planned Apache 2.0 license, if delivered as stated, would permit broad commercial use and modification of the weights. That alone raises the bar for what readers should expect in this space. Reinforcing that, a sparse design with 23B active parameters keeps per-token compute lower than the 501B total size suggests — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the announcement is a preview, since Reflection says Beam is in final red-teaming and the weights, model card, and technical report arrive later this month. On the other side of the ledger, one constraint is real rather than a marketing footnote: the weights, technical report, and model card are not yet released, so none of the benchmark or safety claims can be independently reproduced today. It should factor into any serious decision. Layered on top of that, the coding table shows Beam trailing several listed models, including Qwen 3.8 Max, Kimi K3, and DeepSeek V4.1 Flash on DeepSWE and Terminal Bench — which narrows the set of teams for whom this is an obvious yes.
For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, the smart move is to track its trajectory and revisit once the rough edges are filed down. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- The planned Apache 2.0 license, if delivered as stated, would permit broad commercial use and modification of the weights.
- A sparse design with 23B active parameters keeps per-token compute lower than the 501B total size suggests.
- The blog discloses training scale, hardware, goodput, and RL staleness details that outside engineers can compare against other labs.
- A configurable reasoning-effort setting lets users trade answer length against performance on a per-task basis.
Cons
- The weights, technical report, and model card are not yet released, so none of the benchmark or safety claims can be independently reproduced today
- The coding table shows Beam trailing several listed models, including Qwen 3.8 Max, Kimi K3, and DeepSeek V4.1 Flash on DeepSWE and Terminal Bench.
- The efficiency comparison is a FLOPs approximation and not a measured inference cost, so real serving savings may differ.
- The model is text-only, so it cannot take image or audio input directly the way multimodal open models can.
References
Comments0
Key Features
1. Architecture: sparse MoE with 501B total and 23B active parameters, 52 layers, interleaved local/global attention; text-only. 2. Training (Reflection): 23.8T-token pretraining in under four weeks on 6,144 GB300 NVL72 GPUs; RL on 10.5K GB300 GPUs for four weeks with over 100 million rollouts. 3. Context: midtraining extends effective context length to 1M tokens; RL used a 256K maximum context. 4. Controls: a reasoning-effort parameter trades response length against performance. 5. Availability: waitlist early access now; weights, technical report, model card, and developer artifacts promised later in October under Apache 2.0 (not yet released).
Key Insights
- The announcement is a preview, since Reflection says Beam is in final red-teaming and the weights, model card, and technical report arrive later this month.
- Reflection's own table shows a mixed result, with Beam level with GLM 5.2 on DeepSWE but behind Qwen 3.8 Max, Kimi K3, and DeepSeek V4.1 Flash on that test.
- The 3-4x inference-compute claim rests on a FLOPs estimate of 2 x active parameters x mean generated tokens, which Reflection says excludes prefill, attention, and serving overhead.
- Beam's 23B active parameters out of 501B explain much of the efficiency pitch, since per-token compute scales with the active share.
- Reflection's RL scale statement is framed as belief, and TechCrunch notes the performance claims have not been independently verified.
- A text-only design separates Beam from the multimodal Inkling, which Beam outscores on the four coding tests where both report.
- The reasoning-effort parameter gives deployers a direct knob for latency and cost, which matters for agentic loops that run many steps.
Was this review helpful?
Share
Related AI Reviews
Mistral Large 4 Preview: 1T MoE, Weights Due in October
Mistral Large 4 enters public preview as a ~1T-parameter multimodal MoE with 1M context. Weights are promised for October; scores are vendor-run.
Meta Muse Gadget SDK: Build Your Own ESP32 and Pi Devices
Meta open-sourced Apache 2.0 ESP32 and Linux SDKs for DIY Muse gadgets, with candid security caveats and a Home Link giveaway of 5,000 units.
Meta Expands Muse Agent With Video Chat, Glasses, Charm
Meta expanded its Muse AI agent at Connect 2026 with video chat, smart glasses, Mac control, and a Muse Charm wearable shipping in December.
Grok 4.7 Review: Same Price, Trails Fable 5.1 and GPT-6
xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing on a larger base model, but Artificial Analysis scores it 46 versus 53 for Fable 5.1 and GPT-6.
