zgba 站群
Beam: Reflections 501B open-weight model

Beam: Reflection’s 501B open-weight model

We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Together, these efforts produced competitive open-weight performance with frontier inference compute efficiency.

Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month.

We trained Beam with a particular focus on coding and agentic performance. Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time.

The below figure shows Beam’s performance across a range of coding, agentic, reasoning, and STEM benchmarks. NR denotes scores that have not been reported.

Beam pairs coding and agentic capabilities with highly efficient reasoning. On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute. Efficiency gains are even more pronounced when comparing to models in the 2T+ parameter family like Qwen 3.8-Max, which require significantly more inference compute per token.

These results translate into more intelligence per token, delivering strong model capabilities at lower cost, making Beam a powerful workhorse model for enterprise coding and agentic workloads.

We made high-compute reinforcement learning a central scaling axis for Beam, investing in RL science, data, and infrastructure to turn more compute into stronger capabilities. Scaling RL enables more extensive exploration of problem-solving strategies, while longer rollouts support multi-step reasoning, tool use, and adaptation to environment feedback.

To scale reinforcement learning, we deployed 10.5K NVIDIA GB300 GPUs for four weeks generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes. To sustain a run of this magnitude, we sourced one million high-quality coding, agentic, and STEM environments. We believe this is one of the largest scale RL runs conducted by any open lab to date. Across our evaluation suite, capabilities continued to improve as we increased RL compute, with no sign of a plateau.

We trained Beam with asynchronous policy gradients. At scale, policy staleness becomes a major source of instability for these methods. Long running rollouts have tokens that are generated by multiple model checkpoints, with earlier tokens becoming increasingly stale relative to the current policy. Numerical mismatch between training and inference engines further compounds this challenge.

We developed new algorithms to maintain stable learning under these conditions while systematically reducing training–inference mismatch throughout our pipeline. These advances enable fully asynchronous RL at scale that remains stable, even when learning from interactions generated more than a day earlier.

We trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning. Later, as Beam developed stronger agentic capabilities, completion lengths grew again, but those additional tokens supported further gains in performance. Throughout training, RL improved the tradeoff between capability and token usage.

Users can control this tradeoff through Beam’s reasoning effort parameter: lower settings favor shorter responses, while higher settings allow longer reasoning to improve performance on demanding tasks. This gives users the flexibility to match reasoning effort to their task and compute budget.

We designed Beam’s RL training to develop reasoning and agentic capabilities that generalize beyond its training tasks. During a phase of training on reasoning, software engineering, and terminal tasks, we saw consistent gains in browsing despite the absence of browsing tasks from the RL mixture. This transfer suggests that Beam was learning broader agentic capabilities that generalize across domains. When given web access, it organically learned to search for and query other large language models, and to use OCR APIs to read documents.

The demos below showcase Beam applying these capabilities across research, application development, gameplay, and machine learning workflows. The examples range from building a live NYC subway dashboard using public data to creating interactive applications and preparing model fine-tuning notebooks. Although Beam is text-only, it can work with information from other modalities when represented as text. In another out of distribution domain, Beam also created a fine-tuning notebook for the latest and smallest Gemma-4 model on a Text2SQL task.

Together, these demos illustrate the breadth of tasks Beam can tackle by combining reasoning, coding, and tool use. Each example includes the initial request and resulting output, along with relevant setup and user iterations.

Frontier-scale reinforcement learning requires a large volume of difficult, high-quality tasks. We built a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.

We relied heavily on an iterative curation process. First, we synthesized or sourced environments across a broad set of domains including software engineering, terminal use, competitive coding, STEM, web search, tool use, and general knowledge work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor impossible for the model) and quality (e.g., not underspecified, misleading, guessable, hackable, or otherwise broken or noisy). Third, we tested the tasks through RL, which allowed us to identify further quality or difficulty issues and inform the next iteration of sourcing and filtering.

Throughout Beam’s development, we found that compromises in data quality led to capability plateaus and other training issues. Systematic improvements to task quality were essential to sustaining capability gains throughout the run, which ended with no sign of saturation.

High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. We built an asynchronous platform that lets these processes run independently while coordinating the flow of experience and model updates.

During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:

Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it, allowing the training algorithm to account for policy staleness as completed rollouts flow into training.

Flexible compute allocation: We adjusted the balance between inference and training as the workload evolved, operating at inference-to-training GPU ratios from 3.9:1 to 5.4:1. We also resized the trainer across five GPU mesh configurations within the same training lineage without losing training state.

Fast model updates: New weights reached the inference fleet in a median of approximately 12 seconds. Hierarchical distribution transfers weights across racks over RoCE, then shares them locally over NVLink. Compared with every replica pulling weights directly, this reduced cross-rack traffic by 75% and made fleet-wide adoption of new weights 2.2× faster.

Resilience to inference failures: During the run, 71 inference incidents were handled without terminating the training job. Inference capacity recovered in a median of eight minutes, with lost capacity accounting for just 0.02% of elapsed serving GPU-minutes.

Environments at scale: We supported up to 170K concurrent sandboxes during the run. Across the platform, we processed more than one billion sandbox creation requests, spanning over 20 clusters, two clouds, and four regions. 90% of new sandboxes were ready in under 10 seconds.

Efficient Trainer Packing: Dynamic packing kept training batches 99.99% full on average, holding per-GPU trainer throughput within 1.5% as mean rollout length grew almost 70%.

Observability and reward integrity: Per-token records enabled numerical consistency checks between training and inference at every step. Independent judges re-screened passing solutions for verifier exploits, while replayable records made rewards and their use in training inspectable.

Together, these capabilities enabled us to train on longer interactions and more demanding environments while maintaining throughput, recovering from failures, and checking the integrity of the learning process.

Reinforcement learning builds on top of a robust base model. To facilitate reasoning, we ensured Beam’s foundation had rich knowledge in coding domains, innate agentic capabilities that could be amplified, and stable MoE optimization dynamics.

While developing Beam, we pretrained a series of iteratively bigger models to establish and verify our scaling recipe. Making their performance predictable required carefully designing model tiers, curating diverse in-house code and web validation sets, and aggressively decontaminating all training data against them. The scaling held; the final Beam Base matches its predicted performance, and also matches or outperforms accessible similar-sized open-source base models.

The Beam architecture and optimization recipe emphasizes a numerically healthy foundation for sustained downstream RL. It combines interleaved local and global attention, fine-grained routed experts, a controlled residual stream, and multiple forms of load balancing to ensure stable, balanced expert utilization and healthy signal propagation through residuals.

For expert utilization, we built on auxiliary-loss-free load balancing (DeepSeek-AI et al., 2024), introducing cosine decay of expert-bias updates to reduce routing perturbations later in training. Sequence-level balancing further encourages balanced expert utilization on data outside the pretraining distribution, preparing the model for the changing distribution of downstream RL. As a result, the final pretrained base has almost-perfect uniform utilization, ensuring all experts can be used for learned reasoning.

For the residual stream, we developed a depth-based scaling approach that counteracts activation growth as sublayer outputs accumulate, helping keep residual norms stable as model depth increases. Combined with SandwichNorm, elementwise attention gating, and FP32 residual accumulation – which reduces rounding error when adding small updates to the stream – this recipe controls activation growth and outliers, thereby ensuring healthy signal propagation throughout all layers of Beam. This stability persists throughout pretraining, reinforcement learning, and alignment.

Beam was pretrained on 23.8 trillion diverse high-quality tokens from the web, public sources, and proprietary licensed datasets. Our data pipeline was designed to give Beam a foundation for downstream agentic coding: source code, technical explanations, and mathematical and scientific knowledge, preserved through every stage of curation. We train on almost all publicly accessible and unrestrictively-licensed code and code documentation on the web.

We trained our own quality classifiers for web, code, and STEM content, divided data into fine-grained quality tiers, and weighed training toward stronger material. After extensive scientific iteration, we optimized both precision and recall of data curation significantly beyond conventional web filters used in state-of-the-art OSS data frameworks. On one hand, about 95% of raw Internet tokens are eliminated through parsing, deduplication, and curation. On the other hand, we found that conventional techniques would have missed roughly 1.8 trillion high-quality tokens we retain, including 87% of our curated web-code tokens.

Code modeling requires its own curation for the highest performance. For each language, we applied individually tuned filters

View original article