Reflection AI shipped Beam on October 5, a 501-billion-parameter open-weight model pitched as America’s answer to DeepSeek. You cannot download it. You cannot look up what it costs. And on the agentic-coding benchmarks its own launch leans on, it trails the Chinese models it was built to beat.
That is not a dismissal. Beam is a genuine piece of frontier engineering from a lab that raised roughly $4.7 billion and carries a $25 billion pre-money valuation, with Nvidia, Eric Schmidt, Sequoia and Lightspeed on the cap table and more than $7 billion in compute commitments from SpaceX and Nebius through 2029, per TechCrunch’s launch coverage. But if you are deciding whether to build on it, the launch-day excitement points at the wrong things. The two properties that make an open-weight model worth choosing over a closed API, the fact that you can run the weights yourself and the fact that you can model the cost, are exactly the two you cannot evaluate this week. So the useful question is not “is Beam good.” It is “what are you actually getting, and when.”
What shipped, and what didn’t
Beam is a sparse mixture-of-experts model: 501 billion total parameters with 23 billion active per token, 52 layers, pretrained on 23.8 trillion tokens across 6,144 Nvidia GB300 GPUs, then tuned with more than 100 million reinforcement-learning rollouts, according to the detailed spec writeup at DataCamp. It is text-only, with no image or audio input. The context window is a claimed 1 million tokens, though the beta API currently caps a request at 262,144 tokens of input plus output combined. It carries the now-standard reasoning-effort dial (low, medium, high, xhigh, max), and the model ID is Beam-501B-A23B.
Here is what did not ship. There are no weights. Reflection says it will release them under an Apache 2.0 license “later in October,” together with the technical report, model card and the tooling to run, evaluate and fine-tune the model, as The Next Web reported. Until that happens you cannot self-host, quantize or fine-tune anything. As of launch week the model was not on Hugging Face, Ollama, OpenRouter or models.dev. The only way to touch Beam is a gradual API beta behind a waitlist at platform.reflection.ai, running an OpenAI-compatible Chat Completions endpoint. And there is no published per-token price. An open-weight model is being evaluated, right now, as a closed, unpriced, waitlisted API. That inversion is the whole story.
The benchmark read
Reflection’s own numbers are specific, and they tell a sharper story than “rivals DeepSeek.” Where Beam leads, it leads clearly: 80.9 on SWE-Bench Verified, 77.2 on SWE-bench Pro v2-Hard, and a near-saturated 97.8 on AIME 2026. Those are frontier-class results on classic software-engineering and math reasoning.
Where it was supposedly built to compete, it is behind. On Terminal-Bench v2.1, the agentic-coding benchmark that most closely tracks “can this thing drive a real toolchain,” Beam posts 80.1 against Moonshot’s Kimi K3 at 88.3 and DeepSeek V4.1 Flash at 90.6. On DeepSWE v1.1 the gap is wider: 44.4 versus Kimi K3’s 68.0 and DeepSeek’s 74.2. On BrowseComp, a web-search agent test, Beam scores 77.4 to Kimi K3’s 91.2. On GPQA Diamond it trails both GLM 5.3 and Kimi K3. Reflection does not hide this. The company concedes Kimi K3 remains ahead on raw capability and frames Beam instead as an efficiency play: comparable reasoning to China’s GLM-5.2 while using, in its words, “three to four times less inference compute.”
Read that claim carefully, because Reflection itself did. The company labels the compute comparison an “approximate comparison, not a measured cost,” as The Next Web noted. The 3-to-4x figure is derived from active parameters times generated tokens, not a metered bill, and it excludes prompt processing and serving overhead. Every benchmark above is self-reported, run by Reflection, with no independent verification. None of that makes the numbers wrong. It makes them a vendor’s opening bid, and the efficiency pitch, the one real differentiator, is the number least anchored to anything you could reproduce. This is the same discipline that applies to every leaderboard: the score is a hypothesis about your workload, not a result, until you run it on your own tasks. I have made that argument about trusting benchmarks before, and a launch where the headline metric is self-labeled “approximate” is the cleanest possible example of why.
“Open” is a property of the license, not your data center
The appeal of Beam is right there in co-founder and CEO Misha Laskin’s framing: “The only way to own intelligence is, by definition, if it’s open.” For a specific kind of buyer, that is exactly correct. If you are in a regulated sector, running sovereign-AI infrastructure, or under data-residency rules that make a US or EU hosted API a non-starter, an open-weight frontier model you can stand up inside your own perimeter is not a preference. It is the only architecture that qualifies. A US-governed open frontier model, as an alternative to depending on Chinese weights or on a closed American API, is a genuinely useful thing to exist.
But open is a property of the license. Runnable is a property of your hardware. A 501-billion-parameter mixture-of-experts model is not something you spin up on a spare GPU. Even at 23 billion active parameters, serving the full weight set at a production SLA means a multi-accelerator node and the memory, interconnect and ops maturity to keep it fed. Spending two decades in IT operations, including building sovereign-AI infrastructure at a large telecom, taught me that the license was never the hard part. The hard part was whether we could actually stand the thing up, serve it at the latency our own systems demanded, and keep it running at 3 a.m. “You can have the weights” and “you can operate the model” are different sentences, and the distance between them is a capital-expenditure line, a staffing plan and a quarter of work. For most of the enterprises and public-sector teams Beam is pitched at, the honest near-term path is the hosted API, which quietly reintroduces exactly the vendor dependency that open weights were supposed to remove. Self-hosting an open-source model is viable, but it is an infrastructure project, not a download.
A third access regime, not just a third model
Zoom out and Beam fits a pattern this beat has been tracking all autumn. Google gated its strongest cyber model behind a membership program. OpenAI benched a finished flagship over scope and deception concerns. Now Reflection ships a frontier model as open-weight-in-principle but API-waitlist-in-practice, with the real weights trailing the announcement by weeks. Buyers are no longer just picking a model. They are picking an access and governance regime, and Beam’s regime is “open license, phased release, unpriced beta first.” Co-founder Ioannis Antonoglou gestured at why the phasing exists when he said “it is possible that you get to a level of capability that you want to just be more careful with how you deploy it.” That is a reasonable instinct. It is also a reminder that even the lab whose entire thesis is openness is metering the on-ramp.
Where it fits, and what to do when the weights land
Beam earns a place on your evaluation list if you need a US-governed open-weight frontier model for sovereignty, data-residency or provenance reasons, and you have, or are willing to build, the infrastructure to serve a 501B model. It is worth the waitlist for a serious proof of concept. It does not earn a production commitment today, for a simple reason: you cannot commit to a model whose weights you have not seen and whose price does not exist. And if your decision rests on agentic-coding or web-search accuracy right now, the current open-weight leaders on those specific benchmarks are Chinese models, not Beam.
So treat the launch as a calendar entry, not a purchase. When the Apache 2.0 weights actually drop, do three things before you build on them. Run 50 to 200 of your own representative tasks at two or three effort levels, rather than trusting the self-reported board. Reproduce the efficiency claim on your own serving stack, since “three to four times less compute” is Reflection’s approximation, not your invoice; the only number that matters is cost per completed task on your workload. And price the full standing-up: the GPUs, the serving layer, the people, measured against whatever hosted frontier API you would otherwise call. An open-weight frontier model controlled from the United States is a real and welcome development, and owning your own intelligence is a real strategic goal. Just remember that this week you can own neither the weights nor the price, and plan the work for the week you can.
