Reflection’s Beam: Coding Test Performance vs. Lower Inference Compute in Open Models
Reflection AI has developed Beam, a 501-billion-parameter open-weight model designed for coding and agent tasks, with plans to release the weights under an Apache 2.0 license later this month.
Performance Metrics
Terminal Bench v2.1
On Terminal Bench v2.1, a benchmark for command-line operations, Beam achieved a score of 80.1, trailing GLM 5.2 at 81.0, Kimi K3 at 88.3, and DeepSeek V4.1 Flash at 90.6.
DeepSWE v1.1
In the DeepSWE v1.1 test, Beam scored 44.4, surpassing GLM 5.2’s 44.0 but lagging 29.8 points behind DeepSeek V4.1 Flash.
Compute Efficiency
Training Details
The training phase for Beam spanned four weeks of reinforcement learning, during which the model improved by completing tasks and receiving evaluations. This process utilized approximately 10,500 NVIDIA GB300 GPUs, generated over 100 million attempts, leveraged nearly one million training environments, and employed roughly 1.3 billion sandboxes for training and grading. The company noted that performance metrics continued to rise as the training period concluded.
Extended Capabilities
Some capabilities acquired during training extended beyond designated tasks. For instance, browsing scores improved during a phase where browsing was not explicitly part of the task mix. When granted web access, Beam began querying external large language models and utilizing OCR services to process documents. This behavior enhances agent functionality but warrants careful monitoring before granting network access.
Safety Measures
Safety evaluations remain ongoing. Reflection developed a separate safety and alignment model from the same initial checkpoint, then merged the two through distillation, a technique where one model mimics another’s outputs. The company intends to publish safety results in a technical report alongside the model weights and open-source the internal safety tests it developed.
Additional Information
Early access to Beam is available via a waitlist for a restricted user group. Additional details on AI advancements include updates on NIS2 compliance strategies, newly disclosed vulnerabilities in NetScaler and Exchange Server, and developments in law enforcement efforts targeting cybercriminal networks.
