AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Reflection has introduced Beam, its first open-weight model, a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active per token. The company reports strong coding and reasoning results alongside lower inference compute on some comparisons, but Beam is still undergoing final red-teaming and evaluation; its weights and technical materials have not yet been released.

Reflection has introduced Beam, its first open-weight model, describing it as a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active per token. The company says Beam is designed for coding, reasoning and agentic workloads; the model is still undergoing final red-teaming and evaluation, and its weights are not yet available.

Reflection says it pretrained Beam on 23.8 trillion curated tokens from web sources and proprietary licensed datasets. The company reports that the model matches or outperforms available open base models of similar size, though its announcement does not provide an independent assessment of that comparison. Reflection says it will publish the model weights, technical report, model card and developer artifacts later this month, and is taking sign-ups for early access.

The company says Beam’s training included a large reinforcement-learning campaign: more than 100 million rollouts generated over four weeks using 10,500 NVIDIA GB300 GPUs. Reflection also reports using roughly 1.3 billion sandboxes for training and grading, and sourcing one million coding, agentic and STEM environments. These are figures supplied by the company; the technical report has not yet been released to provide further details.

Reflection’s published benchmark table shows Beam scoring 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. The company characterizes its performance as competitive with larger open models on coding and agentic tasks, while acknowledging that some models remain ahead on raw capability. It says Beam’s advantage is more efficient inference, including reasoning benchmark results comparable to GLM-5.2 with three to four times less inference compute. The comparisons are company-reported and depend on the benchmarks and estimation methods used.

At a glance
announcementWhen: Announced in Reflection’s introduction;…
The developmentReflection announced Beam, a 501-billion-parameter open-weight model focused on coding, reasoning and agentic workloads, with a public release planned later this month.

Lower Compute Could Broaden Model Use

Beam’s central pitch is not simply its 501-billion-parameter total size, but that only 23 billion parameters are active per token and that the model may deliver useful performance with less inference compute. If that holds up in independent testing and real deployments, it could lower the compute burden for organizations running coding or agentic systems at scale. That matters to companies weighing both model capability and the cost of serving repeated requests.

The announcement also puts a spotlight on reinforcement learning at scale as a way to improve models that act across multiple steps, use tools or respond to feedback from an environment. But the reported training resources are substantial, and compute-efficiency claims do not by themselves establish lower total operating costs. Deployment hardware, software, latency, reliability and task-specific results will all affect the practical value.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Training Claims to Public Weights

Reflection presents Beam as its first open-weight release, rather than a model whose parameters are already available to download. The company has announced a planned release of weights and accompanying documentation, but the present announcement is an introduction and early-access sign-up, not a completed public launch.

In its benchmark comparisons, Reflection says Beam is competitive with models including GLM-5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, while conceding that Kimi K3 remains ahead on raw capability. For compute estimates, the company says it counts approximately two operations per active parameter per generated token, includes reasoning and final-answer tokens, and excludes prompt prefill, some attention costs and serving overhead. Those caveats mean its efficiency figures are approximate compute comparisons, not direct measurements of the cost or speed of operating the models.

““Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.””

— Reflection

Amazon

GPU for machine learning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Await the Release

Beam’s weights and technical report are not yet public, so outside researchers and developers cannot yet reproduce the company’s results from the announced materials. Reflection says the model remains in final red-teaming and evaluation, but has not specified the findings, the evaluation scope or whether those processes will change the release schedule.

It is also not yet clear how Beam performs across a broad range of real-world tasks, how its results compare under identical testing conditions, or what hardware and serving costs users will face. The company’s inference estimates exclude several factors that can affect actual usage. The announcement does not provide the full benchmark methodology or independent verification of its training-scale and performance claims.

Amazon

AI development software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Release Materials Set the Next Test

Reflection says it plans to publish Beam’s weights, technical report, model card and developer artifacts later this month, after its current red-teaming and evaluations. Those materials, if released as planned, should clarify how the company conducted its tests and what restrictions or deployment guidance apply. The announcement does not give a specific release date.

Developers and evaluators will then be able to examine the model directly, check the reported coding and agentic results, and compare inference requirements using their own workloads. Until that release, Beam’s benchmark standing and efficiency remain claims made by Reflection rather than independently established findings.

Amazon

AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Beam?

Beam is Reflection’s first open-weight model, a sparse mixture-of-experts system the company says has 501 billion total parameters, with 23 billion active per token. It is intended for coding, reasoning and agentic workloads.

Can developers download Beam now?

No. Reflection says Beam is undergoing final red-teaming and evaluation. It plans to release the weights and related technical materials later this month and is offering an early-access sign-up.

What performance has Reflection reported?

Reflection lists scores including 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. It says Beam is competitive with larger open models on coding and agentic tasks, but these results have not yet been independently verified through the materials announced.

What does Reflection mean by inference efficiency?

The company argues that Beam can deliver strong results with less compute during generation. Its published estimates use active parameters and generated tokens, but exclude prompt prefill, some attention operations and serving overhead, so they are not a complete measure of real-world operating costs.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Sovereignty Paradox: Mistral’s Impact On European AI

Mistral’s rapid growth and European focus reveal strategic risks and challenges in AI sovereignty amid global competition and internal limitations.

Best AI-Driven Routers For A Faster, Smarter Home Network In 2026

Discover the best AI-powered routers for faster, more reliable home Wi-Fi in 2026, with expert rankings and key features explained.

Opus 5 Is Currently #1 On Artificial Analysis Intelligence Leaderboard

Opus 5 has been officially ranked #1 on the Artificial Analysis Intelligence Leaderboard, marking a significant milestone in AI development.

Your Intellectual Fly Is Open When You Use An LLM To Author A Post (2025)

Experts warn that employing large language models for content creation can expose users’ intellectual oversights, highlighting potential privacy and credibility issues.