Reflection’s Beam: Coding Test Performance vs. Lower Inference Compute in Open Models

www.news4hackers.com-reflection-s-beam-coding-test-performance-vs-lower-inference-compute-in-open-models-reflection-s-beam-coding-test-performance-vs-lower-inference-compute-in-open-models

Reflection AI has developed Beam, a 501-billion-parameter open-weight model designed for coding and agent tasks, with plans to release the weights under an Apache 2.0 license later this month.

Performance Metrics

Terminal Bench v2.1

On Terminal Bench v2.1, a benchmark for command-line operations, Beam achieved a score of 80.1, trailing GLM 5.2 at 81.0, Kimi K3 at 88.3, and DeepSeek V4.1 Flash at 90.6.

DeepSWE v1.1

In the DeepSWE v1.1 test, Beam scored 44.4, surpassing GLM 5.2’s 44.0 but lagging 29.8 points behind DeepSeek V4.1 Flash.

Compute Efficiency

Reflection claims Beam matches GLM 5.2’s performance on advanced reasoning tasks while requiring three to four times less inference compute. However, this compute estimate is based on a formula that multiplies twice the active parameter count by the average token generation rate, serving as a rough comparative metric rather than an exact inference cost measurement. The calculation excludes prompt processing, attention mechanisms, and serving overhead.

Training Details

The training phase for Beam spanned four weeks of reinforcement learning, during which the model improved by completing tasks and receiving evaluations. This process utilized approximately 10,500 NVIDIA GB300 GPUs, generated over 100 million attempts, leveraged nearly one million training environments, and employed roughly 1.3 billion sandboxes for training and grading. The company noted that performance metrics continued to rise as the training period concluded.

Extended Capabilities

Some capabilities acquired during training extended beyond designated tasks. For instance, browsing scores improved during a phase where browsing was not explicitly part of the task mix. When granted web access, Beam began querying external large language models and utilizing OCR services to process documents. This behavior enhances agent functionality but warrants careful monitoring before granting network access.

Safety Measures

Safety evaluations remain ongoing. Reflection developed a separate safety and alignment model from the same initial checkpoint, then merged the two through distillation, a technique where one model mimics another’s outputs. The company intends to publish safety results in a technical report alongside the model weights and open-source the internal safety tests it developed.

Additional Information

Early access to Beam is available via a waitlist for a restricted user group. Additional details on AI advancements include updates on NIS2 compliance strategies, newly disclosed vulnerabilities in NetScaler and Exchange Server, and developments in law enforcement efforts targeting cybercriminal networks.



About Author

en_USEnglish