Nvidia-Backed Reflection Debuts Beam, Its First Open-Weight Frontier Model: 501B Params, 23B Active
On Monday, October 5, two-year-old US startup Reflection AI officially launched Beam, its first frontier open-weight model — a text-only MoE: 501B total parameters, only 23B active per token; pretrained on 23.8 trillion tokens, with a 1-million-token context window. The company says Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks at "a third to a quarter" of the inference compute. The launch confirms Axios's weekend report that a release was imminent.

Related
- 01
Mistral Debuts “Le Chonk”: Open-Weight Large 4 in Public Preview, Weights Drop Oct 27
- 02
Nvidia-Backed Upscale AI Launches Token Fabric for Cross-Vendor AI Chip Networking
- 03
Nvidia in Talks to Deepen Investment in Reflection AI or Acquire It
- 04
Super Micro Contractor Pleads Guilty in $2.5B AI-Chip Smuggling Case
What Happened
Reflection AI — founded in 2024 in Brooklyn by two ex-Google DeepMind researchers, Misha Laskin and Ioannis Antonoglou, to automate software development — launched Beam via an official blog post on Monday, with TechCrunch and Reuters covering the same day. The company pitches Beam as a "workhorse model" for enterprises, the public sector and developers, aimed at reasoning, coding and agentic tasks at "a fraction of the token cost and inference time compute."
The architecture: 501B total parameters with only 23B used per token; a text-only MoE trained with high-compute reinforcement learning on 23.8 trillion tokens. For comparison, the rival it names — Z.ai's GLM-5.2 — runs about 744B total parameters with 40B active. The company's comparison list also includes closed labs (Anthropic, OpenAI), China's open camp (DeepSeek, Qwen, Z.ai), Western open players (Mistral, Meta, Cohere), and the US rival Inkling, the model Mira Murati's Thinking Machines Lab released in July.
Key Facts
- Scores: a company scorecard, not yet independently verified: Reflection's published claims: parity with GLM-5.2 on advanced reasoning benchmarks at 3–4x less inference compute; beating today's leading Western open models; outscoring Inkling on four coding tests where both report results (though Inkling is multimodal and Beam is text-only — not quite the same race). TechCrunch explicitly notes these performance claims "haven't been independently verified."
- Reading the scorecard: wins and losses: In the company's own numbers, Beam scores 80.1 on Terminal-Bench 2.1 against GLM-5.2's 81.0 — behind; and 65.5 against 62.1 on SWE-bench Pro v1 — ahead.
- What the "3–4x cheaper" figure covers: That multiple is a model-compute comparison: estimated from benchmark data by Artificial Analysis and DataCurve, excluding prompt prefill, context-dependent attention and serving overhead. It compares the model's own compute, not the actual per-token bill.
- The weights aren't out yet: The company says it will release Beam's weights and the full technical report later this month under Apache 2.0, distributed through hyperscalers and neoclouds with open-source library integrations. What shipped on launch day was a blog post and a scorecard — the weights aren't downloadable yet.
- Funding and compute: Reflection has raised roughly $4.7B (Nvidia, Sequoia, Lightspeed and others) at a $25B pre-money valuation in its last round; this summer it signed compute deals worth $7B+ with SpaceX and Nebius, locking in GB300 supply through 2029. Its main pitch is "AI factories" — helping enterprises and sovereign nations train customized, local AI systems on their own data — and it has begun testing a sovereign AI-factory partnership with South Korea's Shinsegae Group.
- The full scorecard: newer Chinese models rank ahead (updated late Oct 5): In the company's full comparison table, Terminal-Bench 2.1 reads: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM-5.3 88.2, GLM-5.2 81.0, Beam 80.1. Humanity's Last Exam (no tools): Beam 36.2 vs Kimi K3's 46.9. GPQA Diamond: Beam 90.5 vs 93.5. The announcement itself concedes: "where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time." Training scale: 10,500 Nvidia GB300 GPUs ran a four-week RL run producing 100M+ rollouts. Independent benchmarker Artificial Analysis, given early access, says early indicators suggest Beam "will be one of the most token-efficient open models we've seen for its level of intelligence" — but that is a preliminary read on vendor-provided hardware, with no published independent evaluation yet.
Context
Over the past two years, Chinese open models such as DeepSeek, Qwen, Kimi and Z.ai have ranked near the top among open models. Reflection, two years old, has raised $4.7B and signed $7B in compute deals; the weights of the model it just launched are not yet out. Nvidia CEO Jensen Huang has long promoted the "AI factory" concept, which Reflection is also pitching to enterprises and sovereign states — and Nvidia is both its investor and its compute supplier.
Why it matters
Beam has 501B total parameters with 23B active, and the company says it needs a third to a quarter of GLM-5.2's inference compute; scores are self-reported and weights open later in October.
Comments
Today
Oct 12 Monday- BriefAmazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
10 stories · Oct 11
- Quiz
- Call
Latest news
All →- Oct 11Amazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
- Oct 11GSK Expands Chai Deal After Wet-Lab Validation of AI Designs
- Oct 11PPT Master Hits GitHub Trending: Documents Become Native PowerPoint
- Oct 11context-mode Hits GitHub Trending: Tool Output, Sandboxed First
- Oct 11Cloudflare Acquires Deno: Deploy Shuts Down in Six Months
- Oct 11Anthropic Updates Claude Usage Policy: Armed Drones Named and Banned, Effective Nov 12
