Reflection AI's 501B 'Beam' model sounds huge - but you can't actually use it yet

A new open-weight contender claims to rival Chinese heavyweight models on coding and agentic tasks, except nobody outside Reflection AI has been able to test it.

Reflection AI has announced Beam, a enormous open-weight language model it’s pitching as a serious rival to the current crop of large Chinese open models. The headline numbers are certainly big: 501 billion total parameters, though only 23 billion are “active” at any one time thanks to a sparse Mixture-of-Experts design, trained on 23.8 trillion tokens of web and licensed data. There’s just one wrinkle worth flagging before anyone gets carried away - the weights haven’t actually been released.

What’s actually being claimed

Per Reflection’s own blog post, Beam is built primarily for coding, reasoning and “agentic” tasks - the kind of work where a model has to use tools, browse, or act semi-autonomously rather than just answer questions. Alongside pretraining, the company ran a large reinforcement learning programme: over 100 million rollouts across 10,500 Nvidia GB300 GPUs, running continuously for four weeks. That’s a genuinely significant amount of compute, and it’s the sort of training regime usually reserved for frontier closed models from the likes of OpenAI or Google.

Reflection says the result is a model that’s competitive with GLM 5.2 on coding and agentic benchmarks, closing in on Qwen 3.8-Max, though it concedes that Kimi K3 - currently one of the strongest open models around - still leads on raw capability. Reflection’s pitch instead is efficiency: Beam is framed as getting near-frontier results while being cheaper to actually run, thanks to that MoE architecture only activating a fraction of its total parameters per query.

The catch

All of the above comes from a single source: Reflection AI itself. The benchmark tables in the blog post - covering things like SWE-Bench, Terminal-Bench, GPQA Diamond and AIME - are the company’s own reported numbers, measured against rival models using methods nobody outside Reflection has verified. The post itself admits Beam is “undergoing final red-teaming and evaluations,” and that weights, a technical report, model card and developer tools will only arrive “later this month.” Right now, interested parties can sign up for early access - that’s it. There’s no independently reproducible benchmark, no public demo, and no way for outside researchers to check whether Beam performs as described once it’s out in the wild. Big self-reported numbers from a lab with an obvious incentive to look good are a reasonable basis for interest, not for conclusions.

Who this actually matters to

This isn’t a consumer product announcement - nobody’s chatbot app is about to get smarter overnight. Beam is aimed squarely at developers and companies building coding assistants, agents and tooling who want an open-weight alternative that isn’t Chinese-developed (Reflection explicitly frames it as advancing the “Western open-weight frontier”). If you’re not fine-tuning or self-hosting large language models, this release - whenever it lands - won’t touch your day-to-day use of AI tools at all.

The takeaway

Beam looks like a serious, well-resourced effort at an open alternative to models like GLM, Qwen and Kimi, and the training compute involved is real and substantial. But “open-weight” doesn’t mean open yet - until the weights, technical report and model card actually surface and independent testers get their hands on it, the benchmark claims are a sales pitch, not a result. Worth watching, not worth rearranging your stack for just yet.

Sources