Gyro 3.8 Flash Next
Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a single GPU. Gyro-S puts it on a 32 GB card in 27.6 GiB. Gyro-M matches ISTA-DASLab's IQ3_XXS (KLD 0.306 vs 0.311) with 20% less GPU memory, and beats unsloth's UD-Q2_K_XL with 25% less. Both finished all ten of our long-reasoning coding runs.
- 01Fits a 32 GB card. Gyro-S needs 27.64 GiB of GPU memory for its weights and 28.8 GiB with 64k of context. The model's n-gram table stays on disk.
- 023-bit quality at 2-bit size. Gyro-M (34.98 GiB) matches ISTA-DASLab IQ3_XXS (43.80 GiB) on KL divergence, 0.306 vs 0.311.
- 03Finishes what it starts. In ten long-reasoning coding runs, Gyro-M never looped and finished all ten; a 36.5 GiB 2-bit build looped in seven and finished six.
- 04Fast on one card. Up to 221 tok/s decode for Gyro-S on an RTX 5090 in Strata (140 on prose, ~4,900 tok/s prefill); 113 tok/s on agentionai/llama.cpp (CUDA), up to 218 with the MTP draft.
Pick your file
Both files keep the model's architecture: all 48 layers, all 512 experts per layer. The routed experts are stored in a rotated basis with our own rotor code, Agention Precision Rotor (APR); Gyro-M gives the experts and the shared layers more bits. The files need our llama.cpp build or Strata (see Install); stock llama.cpp cannot load them.
| Your hardware | File | Notes |
|---|---|---|
| 32 GB card (RTX 5090, Radeon AI PRO R9700) | Gyro-S | Up to 128k context, or 24k with the MTP draft on the same card. |
| 48 GB: one card or 2×24 GB | Gyro-M | About 37 GiB at 64k context; room for the MTP draft. |
| 64 GB unified memory | Gyro-M | Let the GPU use at least 48 GiB. |
| Strix Halo or other unified memory, 64 GB+ | Gyro-S or Gyro-M | Gyro-S is faster; Gyro-M is closer to the source. |
| 16–24 GB card | Gyro-S | With --n-cpu-moe N in llama.cpp, or Strata's expert cache. See Settings. |
| 64 GB card or machine, more quality | Gyro-L | Coming: about 50 GiB. |
| File | GPU memory (weights) | KLD ↓ | top-1 ↑ |
|---|---|---|---|
Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf | 27.64 GiB | 0.435 | 74.05% |
Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf | 34.98 GiB | 0.306 | 78.12% |
The TQ1_0 / TQ2_0 in the file names is a size class for the Hub's file browser; the files use our own rotor formats, not the standard ternary types. Abliterated versions of both files are in a separate repo.
Quality
How closely each file follows the source, token by token: KL divergence and top-1 agreement against the unsloth Q8_0 source, on a held-out corpus of 2026 technical writing and code (-c 2048, 60 chunks, one binary for every file). GPU memory is the GPU-resident weights; the n-gram table stays on disk for every file.
Gyro-M matches ISTA IQ3_XXS with 20% less GPU memory, beats unsloth UD-Q2_K_XL with 25% less, and beats ISTA IQ2_XS and Q2_0 at slightly less. Gyro-S is the same size as ISTA's 27.56 GiB Coder IQ1_M and well ahead of it (0.435 vs 0.504). Hover or tap a point for exact figures.
| File | GPU memory | KLD ↓ | top-1 ↑ |
|---|---|---|---|
| Gyro-S | 27.64 GiB | 0.435 | 74.05% |
| ISTA-DASLab GSQ-RCO Coder IQ1_M | 27.56 GiB | 0.504 | 72.96% |
| Gyro-M | 34.98 GiB | 0.306 | 78.12% |
| ISTA-DASLab GSQ-RCO Q2_0 | 35.03 GiB | 0.468 | 73.93% |
| ISTA-DASLab GSQ-RCO IQ2_XS | 36.52 GiB | 0.420 | 74.22% |
| ISTA-DASLab GSQ-RCO IQ3_XXS | 43.80 GiB | 0.311 | 77.68% |
| unsloth UD-Q2_K_XL | 46.63 GiB | 0.334 | 76.61% |
| unsloth UD-IQ3_XXS | 49.50 GiB | 0.264 | 79.39% |
| ISTA-DASLab GSQ-RCO IQ3_S | 51.04 GiB | 0.204 | 81.32% |
| unsloth UD-Q3_K_XL | 56.98 GiB | 0.185 | 81.93% |
| Agention AP-Q4_K_XL | 67.36 GiB | 0.112 | 85.08% |
We don't use wikitext for this model: its n-gram table has memorised it. See the research writeup for why.
Behaviour: long reasoning without loops
Low-bit quants of this model tend to loop in long reasoning: the same paragraph, again and again, until the context runs out. We test for it with one demanding one-shot coding prompt, a voxel pagoda scene in a single HTML file, at reasoning effort medium, over ten seeds (temperature 0.7, top-p 0.95, top-k 20, min-p 0, on an RTX A6000).
Loops also cost time. On the A6000 an answer took about 9 minutes on average with Gyro-M (median 8.4), and about 25 minutes with IQ2_XS (median 16.7).
Speed by hardware and engine
Decode at batch size 1, in tokens per second. Where we list prose, JSON and code separately, the speed depends on how predictable the output is.
| Hardware | Engine | File | Decode | Prefill | Notes |
|---|---|---|---|---|---|
| RTX 5090, 32 GB | Strata rc1 (release candidate) | Gyro-S | prose 140 JSON 221 code 209 | 4,883 | Greedy, MTP draft always on, all experts cached. Prefill on a 16k-token prompt. |
| RTX 5090, 32 GB | agentionai/llama.cpp main (from 2026-10-04), CUDA | Gyro-S | plain 113 JSON 210 code 180 copy 218 | ~3,000 | tg128; pp2048. Plain decode is without the draft; JSON, code and copy are with the MTP draft (greedy). Ryzen 9 9950X host. |
| RTX 5090, 32 GB | Strata rc1 (release candidate) | Gyro-M | prose 114 JSON 171 code 153 copy 176 | — | |
| RTX A6000, 48 GB | agentionai/llama.cpp, CUDA | Gyro-S | 57.4 | 740 | pp2048 |
| RTX A6000, 48 GB | agentionai/llama.cpp, CUDA | Gyro-M | 52.3 | 720 | pp2048 |
| Radeon AI PRO R9700, 32 GB | agentionai/llama.cpp, Vulkan | Gyro-S | 57.8 | 1,246 | pp512; measured by a tester. |
| AMD Strix Halo | agentionai/llama.cpp, Vulkan | Gyro-S | plain 32.7 prose 46.5 JSON 70.4 code 60.0 | 250 | pp512, balanced power. Plain decode is without the draft; prose, JSON and code are with the MTP draft. |
| AMD Strix Halo | agentionai/llama.cpp, Vulkan | Gyro-M | plain 27.5 prose 38.5 JSON 61.3 code 51.9 | 249 | pp512, balanced power. Same split as above. |
The MTP draft (cost-aware) pays most on predictable output: code, JSON and edits. On fast cards it can slow down prose and long reasoning, so try both.
Install
Gyro runs on our llama.cpp fork, agentionai/llama.cpp (Vulkan and CUDA), and on Strata. Tested on Linux with Vulkan (Strix Halo, R9700) and CUDA (RTX 5090, RTX A6000). ROCm, Windows and Apple (Metal) are untested or not supported yet.
1. The llama.cpp container
The quickest start. The container image ghcr.io/agentionai/agention-llama:server has our build ready; it runs through Vulkan on both AMD and NVIDIA.
hf download agentionai/Qwen3.8-Flash-Next-Gyro-GGUF Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf \
mtp-Qwen3.8-Flash-Next-draft.gguf mmproj-F16.gguf --local-dir ~/models
# AMD (and Intel) GPUs
docker run --rm -it --device /dev/dri --group-add "$(getent group render | cut -d: -f3)" \
-v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
-m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--mmproj /models/mmproj-F16.gguf
# NVIDIA GPUs: the same, with the NVIDIA container toolkit and its Vulkan driver
docker run --rm -it --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all \
-v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
-m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--mmproj /models/mmproj-F16.ggufThen open http://localhost:8080. For Gyro-M, download and pass Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf instead.
NVIDIA with Vulkan: many cloud GPU containers grant only compute,utility; then NVIDIA's Vulkan driver cannot start and llama.cpp silently falls back to the CPU (no usable GPU found). Keep -e NVIDIA_DRIVER_CAPABILITIES=all. Minimal images may also need libxext6 and libx11-6. For full speed on NVIDIA, build with CUDA (below).
2. Build from source
Needs Vulkan headers 1.4 or newer and glslc for the Vulkan build. On Ubuntu 22.04 the system packages are too old: install the LunarG Vulkan SDK and source its setup-env.sh first. NVIDIA: build with CUDA instead (CUDA toolkit 12.x+); the RTX 5090 and A6000 numbers above are from the CUDA build.
git clone https://github.com/agentionai/llama.cpp && cd llama.cpp
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j # NVIDIA: -DGGML_CUDA=ON instead (CUDA toolkit 12.x+)
./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:Gyro-S -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05With -hf the vision projector downloads and loads automatically, and so does the MTP draft when you add --spec-type draft-mtp. Use :Gyro-M for the larger file.
3. Strata (release candidate)
Strata has Gyro support on the rc1 branch. It is the fastest way to run Gyro on an RTX 5090, and its expert cache runs Gyro on 24 GB cards. One script sets it up; the full guide is docs/GYRO.md.
git clone -b rc1 https://github.com/agentionai/Strata && cd Strata
tools/gyro_setup.sh --model S # or --model MSettings and tuning
- →Sampling. Qwen's thinking-mode settings plus min-p:
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05. Min-p trims the low-probability tokens that a low-bit quant lifts slightly. - →Keep
--ngram-on-diskin every command, with the model on fast NVMe. It keeps the n-gram table off the GPU and reads its rows with a fast parallel reader. Plain memory-mapping (--lazy-mode on) roughly halves prompt-processing speed. - →MTP draft for speculative decoding: download
mtp-Qwen3.8-Flash-Next-draft.gguffrom the repo (about 3.75 GiB of GPU memory) and add--spec-type draft-mtp -md mtp-Qwen3.8-Flash-Next-draft.gguf --spec-draft-n-max 6 --spec-draft-n-min 2 --spec-draft-mtp-vocab 32768. With-hf, leave out-md. On a single 32 GB card the draft fits next to Gyro-S up to 24k context with-ub 256 -ctkd q8_0 -ctvd q8_0. - →Two GPUs? Put the whole model on one and the MTP draft on the other:
--device Vulkan0 --device-draft Vulkan1. Splitting the model's layers across both cards makes them take turns and is slower than one card. - →16–24 GB cards.
--n-cpu-moe Nkeeps the experts of N layers in system RAM, about 0.55 GiB each. Start with N = 26 on 16 GB or N = 12 on 24 GB and lower it until the model just fits; drop--mmprojto save about 1 GiB. Speed depends on your RAM bandwidth and CPU. 32 GB of system RAM minimum, 64 GB recommended. Or use Strata's expert cache. - →Ampere (RTX 30-series, A-series) with CUDA: the batched prompt path can crash (
MUL_MAT_ID failed,illegal memory access) when several requests with longer prompts run at once. Until the fix ships, start the server withGGML_CUDA_TQ_MMQ=0or--parallel 1. - →Vision. Both files read images through
mmproj-F16.gguf(about 1 GiB of GPU memory). To run text-only, drop--mmproj(with-m) or add--no-mmproj(with-hf).
Memory and context
What the GPU holds for Gyro-S (weights, KV cache, recurrent state and compute buffers), q8_0 KV cache, n-gram table on disk, one slot:
| Context | Gyro-S | + MTP draft |
|---|---|---|
| 32k | 28.1 GiB | ~31.9 GiB |
| 64k | 28.8 GiB | ~32.6 GiB |
| 128k | 30.2 GiB | ~34.0 GiB |
| 256k | ~33.3 GiB (extrapolated) | ~37.1 GiB |
On a 32 GB card: up to 128k context without the MTP draft, 24k with it, or put the draft on a second GPU. Gyro-M needs about 37 GiB at 64k context.
Status and roadmap
- ✓Gyro-S and Gyro-M: available.
- ✓Abliterated Gyro-S and Gyro-M: available in their own repo.
- …Gyro-L: coming. About 50 GiB, for 64 GB machines.
- …Gyro Signal: coming. Shorter answers, no extra loops.
- …Strata support: release candidate.
Gyro is experimental. The format and kernels may still change, and the files will be re-published when they do. We'd like to hear how it runs on your card.
Links
- ↗Gyro on HuggingFace (Gyro-S, Gyro-M, MTP draft, vision projector)
- ↗Gyro abliterated on HuggingFace
- ↗agentionai/llama.cpp and the container
ghcr.io/agentionai/agention-llama:server - ↗Strata,
rc1and its Gyro guide - →Our mainline-compatible Qwen3.8-Flash-Next builds (AP tiers, 76 GiB and up)
- ♥Sponsor AgentionAI on GitHub: a one-person team; GPU time funds the next quant.