Strata: run Qwen3.8-Flash-Next (125B) on consumer hardware
A 125B model on a 12 GB gaming GPU at 94 tokens/s. Strata runs Qwen3.8-Flash-Next by spreading weights across GPU, RAM and SSD, with a small draft model proposing tokens so the big one verifies them in batches. That number is an RTX 5070 with a Ryzen 5 7600 at Q2_0, so heavily quantised, but it’s MIT-licensed and speaks the OpenAI API. The “you need an H100” assumption keeps shrinking.