TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Reflection has introduced Beam, its first open-weight model, a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active per token. The company reports strong coding and reasoning results alongside lower inference compute on some comparisons, but Beam is still undergoing final red-teaming and evaluation; its weights and technical materials have not yet been released.
Reflection has introduced Beam, its first open-weight model, describing it as a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active per token. The company says Beam is designed for coding, reasoning and agentic workloads; the model is still undergoing final red-teaming and evaluation, and its weights are not yet available.
Reflection says it pretrained Beam on 23.8 trillion curated tokens from web sources and proprietary licensed datasets. The company reports that the model matches or outperforms available open base models of similar size, though its announcement does not provide an independent assessment of that comparison. Reflection says it will publish the model weights, technical report, model card and developer artifacts later this month, and is taking sign-ups for early access.
The company says Beam’s training included a large reinforcement-learning campaign: more than 100 million rollouts generated over four weeks using 10,500 NVIDIA GB300 GPUs. Reflection also reports using roughly 1.3 billion sandboxes for training and grading, and sourcing one million coding, agentic and STEM environments. These are figures supplied by the company; the technical report has not yet been released to provide further details.
Reflection’s published benchmark table shows Beam scoring 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. The company characterizes its performance as competitive with larger open models on coding and agentic tasks, while acknowledging that some models remain ahead on raw capability. It says Beam’s advantage is more efficient inference, including reasoning benchmark results comparable to GLM-5.2 with three to four times less inference compute. The comparisons are company-reported and depend on the benchmarks and estimation methods used.
Lower Compute Could Broaden Model Use
Beam’s central pitch is not simply its 501-billion-parameter total size, but that only 23 billion parameters are active per token and that the model may deliver useful performance with less inference compute. If that holds up in independent testing and real deployments, it could lower the compute burden for organizations running coding or agentic systems at scale. That matters to companies weighing both model capability and the cost of serving repeated requests.
The announcement also puts a spotlight on reinforcement learning at scale as a way to improve models that act across multiple steps, use tools or respond to feedback from an environment. But the reported training resources are substantial, and compute-efficiency claims do not by themselves establish lower total operating costs. Deployment hardware, software, latency, reliability and task-specific results will all affect the practical value.
As an affiliate, we earn on qualifying purchases.
From Training Claims to Public Weights
Reflection presents Beam as its first open-weight release, rather than a model whose parameters are already available to download. The company has announced a planned release of weights and accompanying documentation, but the present announcement is an introduction and early-access sign-up, not a completed public launch.
In its benchmark comparisons, Reflection says Beam is competitive with models including GLM-5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, while conceding that Kimi K3 remains ahead on raw capability. For compute estimates, the company says it counts approximately two operations per active parameter per generated token, includes reasoning and final-answer tokens, and excludes prompt prefill, some attention costs and serving overhead. Those caveats mean its efficiency figures are approximate compute comparisons, not direct measurements of the cost or speed of operating the models.
““Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.””
— Reflection
As an affiliate, we earn on qualifying purchases.
Independent Tests Await the Release
Beam’s weights and technical report are not yet public, so outside researchers and developers cannot yet reproduce the company’s results from the announced materials. Reflection says the model remains in final red-teaming and evaluation, but has not specified the findings, the evaluation scope or whether those processes will change the release schedule.
It is also not yet clear how Beam performs across a broad range of real-world tasks, how its results compare under identical testing conditions, or what hardware and serving costs users will face. The company’s inference estimates exclude several factors that can affect actual usage. The announcement does not provide the full benchmark methodology or independent verification of its training-scale and performance claims.
As an affiliate, we earn on qualifying purchases.
Release Materials Set the Next Test
Reflection says it plans to publish Beam’s weights, technical report, model card and developer artifacts later this month, after its current red-teaming and evaluations. Those materials, if released as planned, should clarify how the company conducted its tests and what restrictions or deployment guidance apply. The announcement does not give a specific release date.
Developers and evaluators will then be able to examine the model directly, check the reported coding and agentic results, and compare inference requirements using their own workloads. Until that release, Beam’s benchmark standing and efficiency remain claims made by Reflection rather than independently established findings.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Beam?
Beam is Reflection’s first open-weight model, a sparse mixture-of-experts system the company says has 501 billion total parameters, with 23 billion active per token. It is intended for coding, reasoning and agentic workloads.
Can developers download Beam now?
No. Reflection says Beam is undergoing final red-teaming and evaluation. It plans to release the weights and related technical materials later this month and is offering an early-access sign-up.
What performance has Reflection reported?
Reflection lists scores including 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. It says Beam is competitive with larger open models on coding and agentic tasks, but these results have not yet been independently verified through the materials announced.
What does Reflection mean by inference efficiency?
The company argues that Beam can deliver strong results with less compute during generation. Its published estimates use active parameters and generated tokens, but exclude prompt prefill, some attention operations and serving overhead, so they are not a complete measure of real-world operating costs.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
