— reading now
Flash

TopicTrending open-source AI on GitHub

Story·Apps & Agents·2026-10-05 20:15

Open-Source Strata Runs a 125-Billion-Parameter Model on a 12GB Gaming PC

On September 24, developer Niko1221 released Strata, an open-source inference engine (MIT license) on GitHub, with a single goal: run the 125-billion-parameter Qwen3.8-Flash-Next on an ordinary gaming PC — 12GB of VRAM and 32GB of RAM is enough. In 12 days the project passed 12,000 stars and 1,000 forks, with Hackaday and others covering it.

Download on GitHub →

What Happened

Strata is not a model — it's an inference engine. Author Niko1221 targets exactly one model: Qwen3.8-Flash-Next from Alibaba's Qwen team — a 125-billion-parameter mixture-of-experts architecture with 24,576 "expert" sub-networks, of which only about 10 are needed per generated token, roughly 6 billion active parameters. All 125 billion parameters must stay available, but only a small fraction compute at any moment.

The conventional approach is layer offloading — shuffling whole transformer layers between VRAM and RAM, with PCIe as the bottleneck. Strata tiers by expert popularity: the hottest few thousand experts stay cached in VRAM, all 24,576 experts are pinned in system RAM, the CPU evaluates the ones that missed the VRAM cut in parallel, and a ~29GB n-gram lookup table on the SSD accelerates prompt processing. Add MTP (multi-token prediction) speculative decoding — a small layer drafts a few tokens, the big model verifies them in one pass — claimed at 1.6–1.8x faster. Installation: double-click a .bat on Windows, run setup.sh on Linux, with OpenAI/Anthropic-compatible APIs exposed for Cursor, Claude Code and others, plus an MCP server.

Key Facts

  1. Requirements: 12GB+ VRAM (NVIDIA RTX 20/30/40/50 series, AMD RX 7900/7800/7700/9070 etc.) + 32GB+ RAM + ~80GB disk. Windows 10/11 or Linux, one-click install; open localhost:8080 in a browser, MCP server supported.
  2. Speed: Project's own benchmarks (RTX 5070 12GB): 94 tokens/s at Q2_0, 53 tokens/s at IQ3_S, prompt ingestion at 1,600–2,650 tokens/s. Independent community verification: 120+ tokens/s on 3x RTX 3090 (llama.cpp manages ~20 on the same rig); a rigorous A/B verification on an old RTX 2070 as well.
  3. Marketing numbers vs measurements: "6x faster than llama.cpp" was measured by the community at ~2x like-for-like; the main hardware requirement is RAM rather than VRAM — 64GB of DDR5 currently costs about $913, more than an RTX 5070.
  4. Note for Chinese users: The 1-bit Coder quant systematically breaks Chinese output (issue #438: 0/10 correct in Chinese, 10/10 in English); the README itself admits the Coder build is weak on Chinese/CJK. For Chinese, pick Q2_0 / IQ2_XS and other quants that keep all experts.

Context

In local inference, llama.cpp is the general-purpose engine, vLLM targets servers, and Ollama focuses on ease of use. Strata optimizes deeply for exactly one model, using its MoE sparse activation. On Qwen's own benchmarks, Qwen3.8-Flash-Next matches Qwen3-27B on agentic coding and approaches Claude Opus 4.6 on reasoning (91.7 on GPQA Diamond). v0.1.39 shipped Oct 4: byte-identical outputs, faster decoding, longer contexts, more concurrency. 40 releases and 845 commits in 12 days.

Why it matters

A 125-billion-parameter model runs on a gaming PC with 12GB of VRAM and 32GB of RAM; community tests show about 2x llama.cpp speed, versus the advertised 6x.
Useful Tap if this story helped you

SourcesGitHub (Niko1221/Strata, released 2026-09-24), Hackaday (2026-10-04), Startup Fortune, linuxcompatible.org, AlphaSignal. Compiled from public reporting; not investment advice.

Comments

  1. Loading comments…
Ask the cat