August 24, 2026

Adaptive Speculation

Draft length that follows measured acceptance, not a fixed guess, plus a Vulkan backend tuned for AMD's unified-memory APUs.

September 1, 2026

The benchmark was memorized

We needed a reliable way to evaluate our own quantization recipes. Wikitext-2 turned out to be memorized by Qwen3.8-Flash-Next's n-gram table, so we built our own held-out benchmark. A case study in what that changed.

August 28, 2026

30GB smaller, and higher quality

Our ROCmFPx build of Qwen3.8-Flash-Next is 30GB smaller than the leading community IQ4_XS quant, and it comes out ahead on perplexity too, not behind.