For companies evaluating AI models, benchmark scores reveal only part of the picture. Deployment location and operating cost can be just as important as performance, especially when organizations want more control over their data and infrastructure.
Reflection AI is building its business around companies that want capable models they can run themselves. On October 5, it unveiled Beam, its first model intended for open-weight release. CEO Misha Laskin described it as a “workhorse” for coding and agentic tasks, with Chinese open models among its main competitors.
Reflection bets that organizations want a Western model they can deploy on their own infrastructure. Beam will have to show that it can deliver enough capability at a competitive operating cost to make that option viable.
What Beam offers, and what is still pending
According to Reflection’s announcement, Beam has 501 billion total parameters, with 23 billion active per token, and was pretrained on 23.8 trillion tokens. Its sparse mixture-of-experts architecture uses only a fraction of those parameters per token, which can reduce inference computation.
Beam is currently text-only, so it covers a narrower range of tasks than multimodal models. It also remains in early access while Reflection completes red-team testing and final evaluations. The company plans to release the weights later this month under Apache 2.0, along with documentation and software for deployment, evaluation, and fine-tuning.
Beam’s planned release also raises an important distinction between open-weight and open-source AI. Open-weight refers to access to a model’s learned parameters, while the Open Source AI Definition also addresses supporting code and information about training data. Until Beam’s full release arrives, the accompanying materials will help determine whether it meets the open-source AI definition.
Why Chinese models are the comparison
Reflection says Beam performs comparably to Z.ai’s GLM-5.2 and approaches Alibaba’s Qwen3.8-Max on coding and agentic tasks. Those models are more direct comparisons than closed systems from OpenAI or Anthropic because organizations can deploy their weights themselves instead of relying only on an external API.
Independent testing offers context for Reflection’s GLM-5.2 comparison. In its July assessment, the U.S. Center for AI Standards and Innovation judged GLM-5.2 probably the most capable open-weight model available at release, with overall capabilities similar to GPT-5.2.
NIST evaluated GLM-5.2 rather than Beam, so the assessment says nothing about whether Reflection’s own performance claims are accurate. It does show that matching GLM-5.2 would put Beam against one of today's leading open-weight models.
Alibaba’s Qwen3.8-Max announcement describes a multimodal model with 2.4 trillion parameters and a context window of up to 1 million tokens. Even if Beam reaches similar coding performance, the comparison doesn't mean the two models offer equivalent overall capabilities because Beam is limited to text.
The real test is cost per completed task
Reflection claims Beam uses three to four times less inference compute than comparable open models. TechCrunch reports that those performance claims have not yet been independently verified.
Reflection’s compute estimates also exclude prompt processing, some attention calculations, and serving overhead. As a result, the figures describe approximate computational work rather than the full cost of running Beam in production.
For companies operating AI at high volume, lower inference requirements could still have a substantial effect. If Beam can complete comparable tasks with less computation, organizations may be able to serve more requests from the same infrastructure or reduce the capacity needed to support a workload.
The more relevant metric for buyers, however, is cost per completed task. A cheaper response provides little benefit if the model needs repeated attempts or extensive human correction before the work is usable.
Sovereign AI still comes with an infrastructure bill
Laskin sees demand from governments and businesses that cannot or prefer not to use Chinese models but still want greater control over their AI systems. His assessment of the current Western alternatives was blunt: “They don’t really have very good options today.”
For those organizations, Beam’s appeal would be the ability to deploy the model in their own environment, customize it, and keep proprietary data under their control. Reflection appears to see two related markets here: governments pursuing sovereign AI and enterprises that want more authority over where their models run and how they are modified.
Building that alternative requires substantial capital. Reflection’s SpaceX compute agreement could reach $6.3 billion, subject to termination provisions, while its Nebius agreement is worth $1 billion.
Reflection also says Beam’s reinforcement-learning run used 10,500 Nvidia GB300 GPUs for four weeks and generated more than 100 million rollouts. Those figures illustrate the cost of developing models intended to give customers more control over deployment.
Reflection is betting that a Western open-weight model can compete by combining self-hosted deployment with lower inference costs. The public release and independent testing will show whether Beam can deliver those performance and efficiency claims outside Reflection’s own evaluations.