Reflection AI Beam — America's open-weight AI model for coding and agents

America's Open-Weight Counterpunch Has Arrived

For the better part of two years, the open-source AI race has had a clear pattern: Chinese labs ship efficient open-weight models, and the rest of the world reacts. DeepSeek's releases, Alibaba's Qwen line, Moonshot's Kimi, MiniMax, and Z.ai's GLM series have dominated open-model leaderboards and undercut Western API pricing. On October 5, 2026, a New York startup said enough: Reflection AI unveiled Beam, its first open-weight model — a 501-billion-parameter sparse mixture-of-experts system purpose-built for coding and agentic workloads.

Beam is not a chatbot novelty. It is a direct bid to break the assumption that the best cheap open models will keep coming out of China. Here is what the model actually is, what Reflection claims, what the claims do not yet prove, and whether you should care.

What Exactly Is Beam?

Reflection AI is a Brooklyn, New York startup founded in March 2024 by former Google DeepMind researchers Misha Laskin (now CEO) and Ioannis Antonoglou (president and CTO, and a co-creator of AlphaGo). The company focuses on AI that automates software development — one of the fastest-growing uses of large models — and has attracted heavy backing: according to Reuters, Reflection raised $2 billion in October 2025 at an $8 billion valuation, with Nvidia, former Google CEO Eric Schmidt, Citi, 1789 Capital, Lightspeed, and Sequoia participating. By spring 2026, CEO Misha Laskin put the company's valuation at roughly $25 billion.

Beam is the company's first model. In Reflection's own words from its launch materials, it is "a highly efficient agentic open model with 501B total parameters and 23B active" — built for coding, reasoning, and agentic workloads.

Key Specs at a Glance

  • Architecture: Sparse mixture-of-experts (MoE)
  • Parameters: 501 billion total, only 23 billion active per task
  • Purpose: Coding, reasoning, and agentic AI workloads
  • Benchmarks (self-reported): 80.9 on SWE-Bench Verified, 80.1 on Terminal-Bench 2.1, 97.8 on AIME 2026
  • License: Full weights planned under Apache 2.0 (commercial use allowed)
  • Availability: Early version through a waitlist; full weights, technical report, and model card expected later in October 2026

How the Mixture-of-Experts Design Works

A mixture-of-experts model does not fire its entire network for every input. Instead, each request is routed through a small subset of "expert" sub-networks. That is how Beam can carry 501 billion parameters while only activating 23 billion for any given task — the headline capability without the headline inference bill.

For buyers, this distinction matters more than the raw parameter number. Inference cost — what you pay per query once a model is in production — is what shows up on the cloud bill. A model that scores a point or two below a rival but runs at a fraction of the token cost can win the deployment decision for any company running millions of agent calls a day. That is the heart of Reflection's pitch: not "we beat everyone outright," but "you get roughly comparable performance at a fraction of the compute."

The Benchmark Numbers — and the Caveat

Reflection published specific self-reported scores: 80.9 on SWE-Bench Verified, 80.1 on Terminal-Bench 2.1, and 97.8 on AIME 2026. TechCrunch's launch coverage notes the standard caveat: these figures come from Reflection itself and had not been independently verified at launch. Treat them as the company's opening bid, not a leaderboard result.

The company's own comparison frames Beam against Z.ai's GLM-5.2 (about 744 billion total parameters with 40 billion active), where tracked comparisons put GLM-5.2 slightly ahead on Terminal-Bench 2.1 at 81.0 versus Beam's 80.1. Reflection's counter-argument is efficiency: it claims Beam needs three to four times less inference compute than GLM-5.2 to reach a comparable score. Reflection also says Beam is closing in on Alibaba's Qwen3.8-Max on coding and agentic tasks, per Reuters.

Beam vs. the Chinese Open-Weight Leaders

The realistic competitive set is DeepSeek, Qwen, Kimi K2, MiniMax, and GLM — the Chinese model families that have dominated open-weight downloads over the past year. What the launch coverage makes clear is that full side-by-side scorecards against most of these rivals were not published at launch; independent apples-to-apples numbers simply were not available, and any outlet claiming otherwise would be guessing.

What is documented: the Beam-vs-GLM-5.2 efficiency comparison, and Reflection's positioning of Beam as the US-domiciled alternative. For enterprises and government teams in regulated environments, that matters. A credible Apache 2.0-licensed model from a US company sidesteps the data-sovereignty and export-control questions that can come with deploying a Chinese open model — an audience Reflection explicitly name-checked: enterprises, governments, and developers.

Availability: How to Try Beam

Beam is not fully public yet. Reflection says the model is undergoing final red-teaming and evaluation, with an early version offered to a select group through a waitlist. The full weights, technical report, model card, and developer artifacts are expected to ship under an Apache 2.0 license later in October 2026 — a license that permits commercial use, modification, and redistribution.

No pricing has been announced, and Reflection has not disclosed API access terms. The honest status: if you are not on the waitlist, you cannot run Beam today. The thing to watch is the full weights release — once they land under Apache 2.0, developers will be able to download, fine-tune, and self-host the model.

Pros and Cons

Strengths: open weights with a commercial-friendly Apache 2.0 license coming; purpose-built for coding and agents rather than a generalist chatbot; an efficiency story (3–4x less inference compute than GLM-5.2, per Reflection) that could reshape deployment economics; US-domiciled with serious compute backing.

Weaknesses: benchmark scores are self-reported and unverified; not fully public yet — waitlist only; head-to-head data against most Chinese rivals is missing; competing in a lane where rivals already have established ecosystems, tooling, and community fine-tunes.

Who Beam Is For

Beam is aimed at developers and teams building coding agents and agentic workflows — especially those running open models at scale where inference cost is the deciding factor. Enterprises and public-sector teams that need a US-based open-weight option with a permissive license are the explicit target audience. It is not yet a product for casual chat users, and it is not competing with closed APIs like ChatGPT or Claude on general conversation.

The Verdict

Beam is the most credible American answer to Chinese open-weight dominance to date: serious architecture, serious funding, serious compute, and a license developers can actually build businesses on. But credibility is not the same as proof. The self-reported benchmarks need independent verification, the weights are not out yet, and the efficiency claims need third-party measurement before procurement teams should bank on them. Watch the Apache 2.0 weights release later this month — that is when the real evaluation begins.

FAQ

Is Reflection AI Beam open source?

It will be open-weight: Reflection plans to release the full weights, technical report, and model card under the Apache 2.0 license later in October 2026, which permits commercial use and modification. The release has not happened yet.

How big is the Beam model?

Beam has 501 billion total parameters but activates only 23 billion per task, using a sparse mixture-of-experts design that keeps inference costs down.

Can I try Beam right now?

Only through Reflection's early-access waitlist, as the model is still in final red-teaming. General availability of the weights is expected later in October 2026.

Is Beam better than DeepSeek or GLM-5.2?

Reflection claims Beam matches Z.ai's GLM-5.2 while using 3–4x less inference compute, but the benchmarks are self-reported and unverified, and no full independent comparison against DeepSeek or other rivals has been published. It is too early for a definitive answer.