— reading now
Flash

TopicNvidia and Reflection AI

Story·Models & Products·2026-10-06 06:40

Nvidia-Backed Reflection Debuts Beam, Its First Open-Weight Frontier Model: 501B Params, 23B Active

On Monday, October 5, two-year-old US startup Reflection AI officially launched Beam, its first frontier open-weight model — a text-only MoE: 501B total parameters, only 23B active per token; pretrained on 23.8 trillion tokens, with a 1-million-token context window. The company says Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks at "a third to a quarter" of the inference compute. The launch confirms Axios's weekend report that a release was imminent.

What Happened

Reflection AI — founded in 2024 in Brooklyn by two ex-Google DeepMind researchers, Misha Laskin and Ioannis Antonoglou, to automate software development — launched Beam via an official blog post on Monday, with TechCrunch and Reuters covering the same day. The company pitches Beam as a "workhorse model" for enterprises, the public sector and developers, aimed at reasoning, coding and agentic tasks at "a fraction of the token cost and inference time compute."

The architecture: 501B total parameters with only 23B used per token; a text-only MoE trained with high-compute reinforcement learning on 23.8 trillion tokens. For comparison, the rival it names — Z.ai's GLM-5.2 — runs about 744B total parameters with 40B active. The company's comparison list also includes closed labs (Anthropic, OpenAI), China's open camp (DeepSeek, Qwen, Z.ai), Western open players (Mistral, Meta, Cohere), and the US rival Inkling, the model Mira Murati's Thinking Machines Lab released in July.

Key Facts

  1. Scores: a company scorecard, not yet independently verified: Reflection's published claims: parity with GLM-5.2 on advanced reasoning benchmarks at 3–4x less inference compute; beating today's leading Western open models; outscoring Inkling on four coding tests where both report results (though Inkling is multimodal and Beam is text-only — not quite the same race). TechCrunch explicitly notes these performance claims "haven't been independently verified."
  2. Reading the scorecard: wins and losses: In the company's own numbers, Beam scores 80.1 on Terminal-Bench 2.1 against GLM-5.2's 81.0 — behind; and 65.5 against 62.1 on SWE-bench Pro v1 — ahead.
  3. What the "3–4x cheaper" figure covers: That multiple is a model-compute comparison: estimated from benchmark data by Artificial Analysis and DataCurve, excluding prompt prefill, context-dependent attention and serving overhead. It compares the model's own compute, not the actual per-token bill.
  4. The weights aren't out yet: The company says it will release Beam's weights and the full technical report later this month under Apache 2.0, distributed through hyperscalers and neoclouds with open-source library integrations. What shipped on launch day was a blog post and a scorecard — the weights aren't downloadable yet.
  5. Funding and compute: Reflection has raised roughly $4.7B (Nvidia, Sequoia, Lightspeed and others) at a $25B pre-money valuation in its last round; this summer it signed compute deals worth $7B+ with SpaceX and Nebius, locking in GB300 supply through 2029. Its main pitch is "AI factories" — helping enterprises and sovereign nations train customized, local AI systems on their own data — and it has begun testing a sovereign AI-factory partnership with South Korea's Shinsegae Group.
  6. The full scorecard: newer Chinese models rank ahead (updated late Oct 5): In the company's full comparison table, Terminal-Bench 2.1 reads: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM-5.3 88.2, GLM-5.2 81.0, Beam 80.1. Humanity's Last Exam (no tools): Beam 36.2 vs Kimi K3's 46.9. GPQA Diamond: Beam 90.5 vs 93.5. The announcement itself concedes: "where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time." Training scale: 10,500 Nvidia GB300 GPUs ran a four-week RL run producing 100M+ rollouts. Independent benchmarker Artificial Analysis, given early access, says early indicators suggest Beam "will be one of the most token-efficient open models we've seen for its level of intelligence" — but that is a preliminary read on vendor-provided hardware, with no published independent evaluation yet.

Context

Over the past two years, Chinese open models such as DeepSeek, Qwen, Kimi and Z.ai have ranked near the top among open models. Reflection, two years old, has raised $4.7B and signed $7B in compute deals; the weights of the model it just launched are not yet out. Nvidia CEO Jensen Huang has long promoted the "AI factory" concept, which Reflection is also pitching to enterprises and sovereign states — and Nvidia is both its investor and its compute supplier.

Why it matters

Beam has 501B total parameters with 23B active, and the company says it needs a third to a quarter of GLM-5.2's inference compute; scores are self-reported and weights open later in October.
Useful Tap if this story helped you

SourcesReflection official blog (2026-10-05), TechCrunch (2026-10-05), Reuters (2026-10-05), Axios (2026-10-04 weekend scoop), RuntimeWire (2026-10-05 scorecard breakdown), Startup Fortune (2026-10-05 company background), implicator.ai (2026-10-05 full scorecard). Compiled from public reporting; not investment advice.

Comments

  1. Loading comments…
Ask the cat