Adaptive Speculation
Draft length that follows measured acceptance, not a fixed guess, plus a Vulkan backend tuned for AMD's unified-memory APUs.
Systems work on inference speed and model compression, the work behind what we ship.
Draft length that follows measured acceptance, not a fixed guess, plus a Vulkan backend tuned for AMD's unified-memory APUs.
We needed a reliable way to evaluate our own quantization recipes. Wikitext-2 turned out to be memorized by Qwen3.8-Flash-Next's n-gram table, so we built our own held-out benchmark. A case study in what that changed.
Our ROCmFPx build of Qwen3.8-Flash-Next is 30GB smaller than the leading community IQ4_XS quant, and it comes out ahead on perplexity too, not behind.