The two-year old startup Reflection AI has launched Beam, its first open-weight model, as the NVIDIA-backed startup enters competition with leading Chinese open-weight AI models.
Beam is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active parameters, built for coding, reasoning and agentic workloads.
In its announcement, Reflection said Beam is competitive with larger open models such as Zhipu AI (Z.ai)’s GLM-5.2 and approaches Alibaba’s Qwen 3.8-Max on coding and agentic tasks, while Moonshot AI’s Kimi K3 remains ahead on raw capability. The company said Beam’s advantage is its efficiency at inference time.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
– Frontier reasoning efficiency
– Advances the Western open frontier on coding & agentic tasks
– Trained end-to-end from scratchFull weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
— Reflection (@reflection_ai) October 5, 2026
On advanced reasoning benchmarks, Reflection AI said, “it achieves scores comparable to Zhipu AI (Z.ai) GLM-5.2 while using 3-4× less inference compute.” The company said the efficiency gains are even more pronounced when compared with models in the 2-trillion-plus parameter class, such as Qwen 3.8-Max, which require significantly more inference compute per token.
The Brooklyn-based startup said these results translate into “more intelligence per token,” delivering strong model capabilities at “lower cost” and making Beam a “powerful workhorse model” for enterprise coding and agentic workloads.
Founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, Reflection AI has been developing its own large-scale training and reinforcement-learning (RL) infrastructure.
The company pretrained Beam on 23.8 trillion tokens from the web, public sources and proprietary licensed datasets, followed by a large-scale reinforcement-learning campaign. The four-week RL run generated more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs, with training and grading using approximately 1.3 billion sandboxes.
Reflection said the RL campaign was designed to improve both reasoning capability and token efficiency. Users can control the trade-off through a reasoning-effort parameter, with lower settings favoring shorter responses and higher settings allowing longer reasoning for demanding tasks.
Beam was pretrained end-to-end in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Reflection said it built nearly the entire training infrastructure in-house and achieved 92.3% goodput toward the end of the pretraining run.
The launch comes as US technology companies face growing competition from Chinese open-weight models, which have gained attention for their coding capabilities, customizability and lower costs.
To improve expert utilization, Reflection built on DeepSeek-AI’s auxiliary-loss-free load balancing and used cosine decay for expert-bias updates to limit routing changes later in training.
The company is positioning Beam as part of the Western open-weight AI push, with the model currently undergoing final red-teaming and evaluations.
Reflection AI plans to release the model weights later this month, alongside its technical report, model card and developer artifacts.
The company is making the early version available to a select group of users and said Beam is the first model in a planned series.
Also Read: China Slams US Accusations of ‘Systematic’ Copying of Frontier AI Models





