- Beam is a sparse mixture-of-experts (MoE) model developed by Reflection, with 5010 billion parameters, of which 230 billion are activated, designed specifically for coding, reasoning, and agent tasks.
- The model was pretrained on 23.8 trillion high-quality web and licensed data tokens and was trained using 10,500 units of NVIDIA GB300 GPUs operating for 4 weeks, completing over 1 billion reinforcement learning rollouts (execution cases).
- Reflection stated that compared to similarly sized or larger public models like GLM 5.2 and Qwen 3.8-Max, Beam shows competitive performance in coding and agent-related tasks, especially achieving significantly increased efficiency by using 3 to 4 times less GPU computation during inference.
- The reinforcement learning process leveraged a newly optimized asynchronous policy gradient algorithm and approximately 100 million task environments, applying various techniques for balanced expert utilization and stable signal propagation, ensuring overall training stability.
It's impressive that Beam's reinforcement learning process can maintain stable learning even with policy delays of over a day. Generally, delays in policy updates cause performance drops or instability, so I'd like to see how effective Reflection's asynchronous RL algorithm and numerical stability management are in real-world applications.