Reflection unveils Beam, its first major open-weight model
Reflection AI's open-weight Beam model activates only 23 billion parameters. An architecture designed to reduce the inference cost of code.
Reflection AI enters the race for large open models with Beam, its first open-weight model. Introduced on October 5, 2026, Beam is based on a Mixture-of-Experts architecture featuring 501 billion total parameters, but only 23 billion activated during inference. Reflection primarily targets Beam for code, reasoning, and agentic tasks.
This ratio between total size and active parameters is a central part of Beam's positioning. Reflection is not just looking to compete on raw scores, but on the amount of resources required to achieve those performances. Notably, the company claims to achieve results comparable to GLM-5.2 on several reasoning benchmarks while using three to four times less compute at inference. Compared to models exceeding 2 trillion parameters like Qwen 3.8-Max, the resource gap per token would be even more significant. However, these measurements are based on evaluations and estimates published by Reflection, rather than a uniform independent campaign.
On development tasks, Beam scores 80.9 on SWE-bench Verified, 80.1 on Terminal Bench 2.1, and 78 on SWE-bench Multilingual, according to Reflection. On DeepSWE v1.1, its score of 44.4 places it close to GLM 5.2, but behind several newer models such as Qwen 3.8-Max, Kimi K3, or DeepSeek V4.1 Flash. Reflection acknowledges that Kimi K3 retains an advantage in raw capability, presenting Beam instead as an efficiency-focused proposition.
To train this model, Reflection states it began with pre-training on 23.8 trillion tokens from the web, public sources, and licensed proprietary datasets. The company explains that it eliminated approximately 95% of raw internet data during its various processing and selection stages, with a particular focus on code, technical content, mathematics, and science.
Beam's pre-training was reportedly completed in less than four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Reflection then dedicated an even larger infrastructure to reinforcement learning: 10,500 GB300 GPUs for four weeks, over 100 million rollouts, and approximately 1.3 billion execution environments used to train and evaluate the model's interactions. The company indicates it maintained up to 170,000 concurrent environments during this phase.
Reflection also sought to control the amount of reasoning used by the model. Beam features a reasoning effort parameter, allowing users to prioritize shorter, less expensive responses or, conversely, to allocate more compute to complex tasks. During training, Reflection says it rewarded correct solutions while penalizing unnecessarily long reasoning to improve the ratio between performance and token consumption.
This direction is also reflected in agentic tasks. Reflection claims to have observed progress in web navigation, even though some training phases did not directly include this type of task. With access to the web and tools, Beam can search for documentation, query other models, manipulate terminal environments, or rely on external services to process documents. However, Beam remains a text-based model: other modalities must be represented in a form it can exploit.
For example, Reflection shows Beam building a real-time map of the New York subway after searching MTA documentation, generating a 3D game with p5.js, or automatically preparing a fine-tuning notebook for a Gemma model. These demonstrations are provided by Reflection and should therefore be considered selected examples rather than independent evaluations of its general behavior.
The model also offers an announced effective context length of 1 million tokens. Reflection explains that this was achieved during an intermediate training phase using large code repositories, long documents, and tasks requiring information retention over an extended period.
However, Beam is not yet fully available for download. As of October 8, Reflection's official organization on Hugging Face does not yet host any public models. Reflection is currently offering early access to selected users and plans to release the weights, technical report, model card, and developer tools later this month. The weights are to be distributed under the Apache 2.0 license, along with the resources needed to run, evaluate, and fine-tune Beam.
This openness is part of a broader positioning for Reflection. The company wants to offer open models, the software needed to customize them, and an infrastructure to deploy them via its API, in a private cloud, on-premises, or in isolated environments.
Beam thus arrives in a landscape where some of the best-performing open models now come from Chinese laboratories. Reflection does not systematically claim to outperform them: its own charts still place Kimi, Qwen, GLM, or DeepSeek ahead of Beam on several evaluations. Its bet is different. With only 23 billion active parameters out of 501 billion in total, the company seeks to bring performance closer to this frontier while reducing the cost required to run them. The actual release of the weights and independent evaluations will now measure to what extent this advantage is confirmed.