TopicTrending open-source AI on GitHub
Open-Source Strata Runs a 125-Billion-Parameter Model on a 12GB Gaming PC
On September 24, developer Niko1221 released Strata, an open-source inference engine (MIT license) on GitHub, with a single goal: run the 125-billion-parameter Qwen3.8-Flash-Next on an ordinary gaming PC — 12GB of VRAM and 32GB of RAM is enough. In 12 days the project passed 12,000 stars and 1,000 forks, with Hackaday and others covering it.

What Happened
Strata is not a model — it's an inference engine. Author Niko1221 targets exactly one model: Qwen3.8-Flash-Next from Alibaba's Qwen team — a 125-billion-parameter mixture-of-experts architecture with 24,576 "expert" sub-networks, of which only about 10 are needed per generated token, roughly 6 billion active parameters. All 125 billion parameters must stay available, but only a small fraction compute at any moment.
The conventional approach is layer offloading — shuffling whole transformer layers between VRAM and RAM, with PCIe as the bottleneck. Strata tiers by expert popularity: the hottest few thousand experts stay cached in VRAM, all 24,576 experts are pinned in system RAM, the CPU evaluates the ones that missed the VRAM cut in parallel, and a ~29GB n-gram lookup table on the SSD accelerates prompt processing. Add MTP (multi-token prediction) speculative decoding — a small layer drafts a few tokens, the big model verifies them in one pass — claimed at 1.6–1.8x faster. Installation: double-click a .bat on Windows, run setup.sh on Linux, with OpenAI/Anthropic-compatible APIs exposed for Cursor, Claude Code and others, plus an MCP server.
Key Facts
- Requirements: 12GB+ VRAM (NVIDIA RTX 20/30/40/50 series, AMD RX 7900/7800/7700/9070 etc.) + 32GB+ RAM + ~80GB disk. Windows 10/11 or Linux, one-click install; open localhost:8080 in a browser, MCP server supported.
- Speed: Project's own benchmarks (RTX 5070 12GB): 94 tokens/s at Q2_0, 53 tokens/s at IQ3_S, prompt ingestion at 1,600–2,650 tokens/s. Independent community verification: 120+ tokens/s on 3x RTX 3090 (llama.cpp manages ~20 on the same rig); a rigorous A/B verification on an old RTX 2070 as well.
- Marketing numbers vs measurements: "6x faster than llama.cpp" was measured by the community at ~2x like-for-like; the main hardware requirement is RAM rather than VRAM — 64GB of DDR5 currently costs about $913, more than an RTX 5070.
- Note for Chinese users: The 1-bit Coder quant systematically breaks Chinese output (issue #438: 0/10 correct in Chinese, 10/10 in English); the README itself admits the Coder build is weak on Chinese/CJK. For Chinese, pick Q2_0 / IQ2_XS and other quants that keep all experts.
Context
In local inference, llama.cpp is the general-purpose engine, vLLM targets servers, and Ollama focuses on ease of use. Strata optimizes deeply for exactly one model, using its MoE sparse activation. On Qwen's own benchmarks, Qwen3.8-Flash-Next matches Qwen3-27B on agentic coding and approaches Claude Opus 4.6 on reasoning (91.7 on GPQA Diamond). v0.1.39 shipped Oct 4: byte-identical outputs, faster decoding, longer contexts, more concurrency. 40 releases and 845 commits in 12 days.
Why it matters
A 125-billion-parameter model runs on a gaming PC with 12GB of VRAM and 32GB of RAM; community tests show about 2x llama.cpp speed, versus the advertised 6x.
CompaniesAlibaba
Comments
Today
Oct 12 Monday- BriefAmazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
10 stories · Oct 11
- Quiz
- Call
Latest news
All →- Oct 11Amazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
- Oct 11GSK Expands Chai Deal After Wet-Lab Validation of AI Designs
- Oct 11PPT Master Hits GitHub Trending: Documents Become Native PowerPoint
- Oct 11context-mode Hits GitHub Trending: Tool Output, Sandboxed First
- Oct 11Cloudflare Acquires Deno: Deploy Shuts Down in Six Months
- Oct 11Anthropic Updates Claude Usage Policy: Armed Drones Named and Banned, Effective Nov 12
