AgentionAI — Gyro 3.8 Flash Next
  • 01Fits a 32 GB card. Gyro-S needs 27.64 GiB of GPU memory for its weights and 28.8 GiB with 64k of context. The model's n-gram table stays on disk.
  • 023-bit quality at 2-bit size. Gyro-M (34.98 GiB) matches ISTA-DASLab IQ3_XXS (43.80 GiB) on KL divergence, 0.306 vs 0.311.
  • 03Finishes what it starts. In ten long-reasoning coding runs, Gyro-M never looped and finished all ten; a 36.5 GiB 2-bit build looped in seven and finished six.
  • 04Fast on one card. Up to 221 tok/s decode for Gyro-S on an RTX 5090 in Strata (140 on prose, ~4,900 tok/s prefill); 113 tok/s on agentionai/llama.cpp (CUDA), up to 218 with the MTP draft.
Download on HuggingFace ↗ Install guides ↓

Pick your file

Both files keep the model's architecture: all 48 layers, all 512 experts per layer. The routed experts are stored in a rotated basis with our own rotor code, Agention Precision Rotor (APR); Gyro-M gives the experts and the shared layers more bits. The files need our llama.cpp build or Strata (see Install); stock llama.cpp cannot load them.

Your hardwareFileNotes
32 GB card (RTX 5090, Radeon AI PRO R9700)Gyro-SUp to 128k context, or 24k with the MTP draft on the same card.
48 GB: one card or 2×24 GBGyro-MAbout 37 GiB at 64k context; room for the MTP draft.
64 GB unified memoryGyro-MLet the GPU use at least 48 GiB.
Strix Halo or other unified memory, 64 GB+Gyro-S or Gyro-MGyro-S is faster; Gyro-M is closer to the source.
16–24 GB cardGyro-SWith --n-cpu-moe N in llama.cpp, or Strata's expert cache. See Settings.
64 GB card or machine, more qualityGyro-LComing: about 50 GiB.
FileGPU memory (weights)KLD ↓top-1 ↑
Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf27.64 GiB0.43574.05%
Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf34.98 GiB0.30678.12%

The TQ1_0 / TQ2_0 in the file names is a size class for the Hub's file browser; the files use our own rotor formats, not the standard ternary types. Abliterated versions of both files are in a separate repo.

Quality

How closely each file follows the source, token by token: KL divergence and top-1 agreement against the unsloth Q8_0 source, on a held-out corpus of 2026 technical writing and code (-c 2048, 60 chunks, one binary for every file). GPU memory is the GPU-resident weights; the n-gram table stays on disk for every file.

KL divergence vs. the Q8_0 source, by GPU memory. Lower is better and plotted higher.
KL divergence vs. GPU memory Eleven quantized builds of Qwen3.8-Flash-Next plotted by GPU-resident weight size against KL divergence from the unsloth Q8_0 source on a held-out corpus. Lower KL divergence is plotted higher. Gyro-S and Gyro-M sit above the ISTA-DASLab and unsloth builds of similar size; Gyro-M matches ISTA IQ3_XXS with 20% less memory. 0.1 0.2 0.3 0.4 0.5 KLD ↓ better ↑ 30 40 50 60 70 GiB 32 GB card Gyro-S Gyro-M AP-Q4_K_XL Gyro-S · 27.64 GiB · KLD 0.435 · 74.05% top-1 Gyro-S 27.64 GiB · KLD 0.435 · 74.05% Gyro-M · 34.98 GiB · KLD 0.306 · 78.12% top-1 Gyro-M 34.98 GiB · KLD 0.306 · 78.12% ISTA-DASLab GSQ-RCO Coder IQ1_M · 27.56 GiB · KLD 0.504 · 72.96% top-1 ISTA-DASLab Coder IQ1_M 27.56 GiB · KLD 0.504 · 72.96% ISTA-DASLab GSQ-RCO Q2_0 · 35.03 GiB · KLD 0.468 · 73.93% top-1 ISTA-DASLab Q2_0 35.03 GiB · KLD 0.468 · 73.93% ISTA-DASLab GSQ-RCO IQ2_XS · 36.52 GiB · KLD 0.420 · 74.22% top-1 ISTA-DASLab IQ2_XS 36.52 GiB · KLD 0.420 · 74.22% ISTA-DASLab GSQ-RCO IQ3_XXS · 43.80 GiB · KLD 0.311 · 77.68% top-1 ISTA-DASLab IQ3_XXS 43.80 GiB · KLD 0.311 · 77.68% unsloth UD-Q2_K_XL · 46.63 GiB · KLD 0.334 · 76.61% top-1 unsloth UD-Q2_K_XL 46.63 GiB · KLD 0.334 · 76.61% unsloth UD-IQ3_XXS · 49.50 GiB · KLD 0.264 · 79.39% top-1 unsloth UD-IQ3_XXS 49.50 GiB · KLD 0.264 · 79.39% ISTA-DASLab GSQ-RCO IQ3_S · 51.04 GiB · KLD 0.204 · 81.32% top-1 ISTA-DASLab IQ3_S 51.04 GiB · KLD 0.204 · 81.32% unsloth UD-Q3_K_XL · 56.98 GiB · KLD 0.185 · 81.93% top-1 unsloth UD-Q3_K_XL 56.98 GiB · KLD 0.185 · 81.93% Agention AP-Q4_K_XL · 67.36 GiB · KLD 0.112 · 85.08% top-1 Agention AP-Q4_K_XL 67.36 GiB · KLD 0.112 · 85.08%
Agention (Gyro, AP-Q4) ISTA-DASLab GSQ-RCO unsloth UD

Gyro-M matches ISTA IQ3_XXS with 20% less GPU memory, beats unsloth UD-Q2_K_XL with 25% less, and beats ISTA IQ2_XS and Q2_0 at slightly less. Gyro-S is the same size as ISTA's 27.56 GiB Coder IQ1_M and well ahead of it (0.435 vs 0.504). Hover or tap a point for exact figures.

FileGPU memoryKLD ↓top-1 ↑
Gyro-S27.64 GiB0.43574.05%
ISTA-DASLab GSQ-RCO Coder IQ1_M27.56 GiB0.50472.96%
Gyro-M34.98 GiB0.30678.12%
ISTA-DASLab GSQ-RCO Q2_035.03 GiB0.46873.93%
ISTA-DASLab GSQ-RCO IQ2_XS36.52 GiB0.42074.22%
ISTA-DASLab GSQ-RCO IQ3_XXS43.80 GiB0.31177.68%
unsloth UD-Q2_K_XL46.63 GiB0.33476.61%
unsloth UD-IQ3_XXS49.50 GiB0.26479.39%
ISTA-DASLab GSQ-RCO IQ3_S51.04 GiB0.20481.32%
unsloth UD-Q3_K_XL56.98 GiB0.18581.93%
Agention AP-Q4_K_XL67.36 GiB0.11285.08%

We don't use wikitext for this model: its n-gram table has memorised it. See the research writeup for why.

Behaviour: long reasoning without loops

Low-bit quants of this model tend to loop in long reasoning: the same paragraph, again and again, until the context runs out. We test for it with one demanding one-shot coding prompt, a voxel pagoda scene in a single HTML file, at reasoning effort medium, over ten seeds (temperature 0.7, top-p 0.95, top-k 20, min-p 0, on an RTX A6000).

Pagoda test, 10 seeds each
Runs that looped (fewer is better)
Gyro-M34.98 GiB
0 / 10
Gyro-S27.64 GiB
1 / 10, mild
ISTA-DASLab GSQ-RCO IQ2_XS36.52 GiB
7 / 10
Runs that finished with a complete page (more is better)
Gyro-M34.98 GiB
10 / 10
Gyro-S27.64 GiB
10 / 10
ISTA-DASLab GSQ-RCO IQ2_XS36.52 GiB
6 / 10

Loops also cost time. On the A6000 an answer took about 9 minutes on average with Gyro-M (median 8.4), and about 25 minutes with IQ2_XS (median 16.7).

Speed by hardware and engine

Decode at batch size 1, in tokens per second. Where we list prose, JSON and code separately, the speed depends on how predictable the output is.

HardwareEngineFileDecodePrefillNotes
RTX 5090, 32 GBStrata rc1 (release candidate)Gyro-Sprose 140
JSON 221
code 209
4,883Greedy, MTP draft always on, all experts cached. Prefill on a 16k-token prompt.
RTX 5090, 32 GBagentionai/llama.cpp main (from 2026-10-04), CUDAGyro-Splain 113
JSON 210
code 180
copy 218
~3,000tg128; pp2048. Plain decode is without the draft; JSON, code and copy are with the MTP draft (greedy). Ryzen 9 9950X host.
RTX 5090, 32 GBStrata rc1 (release candidate)Gyro-Mprose 114
JSON 171
code 153
copy 176
—
RTX A6000, 48 GBagentionai/llama.cpp, CUDAGyro-S57.4740pp2048
RTX A6000, 48 GBagentionai/llama.cpp, CUDAGyro-M52.3720pp2048
Radeon AI PRO R9700, 32 GBagentionai/llama.cpp, VulkanGyro-S57.81,246pp512; measured by a tester.
AMD Strix Haloagentionai/llama.cpp, VulkanGyro-Splain 32.7
prose 46.5
JSON 70.4
code 60.0
250pp512, balanced power. Plain decode is without the draft; prose, JSON and code are with the MTP draft.
AMD Strix Haloagentionai/llama.cpp, VulkanGyro-Mplain 27.5
prose 38.5
JSON 61.3
code 51.9
249pp512, balanced power. Same split as above.

The MTP draft (cost-aware) pays most on predictable output: code, JSON and edits. On fast cards it can slow down prose and long reasoning, so try both.

Install

Gyro runs on our llama.cpp fork, agentionai/llama.cpp (Vulkan and CUDA), and on Strata. Tested on Linux with Vulkan (Strix Halo, R9700) and CUDA (RTX 5090, RTX A6000). ROCm, Windows and Apple (Metal) are untested or not supported yet.

1. The llama.cpp container

The quickest start. The container image ghcr.io/agentionai/agention-llama:server has our build ready; it runs through Vulkan on both AMD and NVIDIA.

hf download agentionai/Qwen3.8-Flash-Next-Gyro-GGUF Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf \
  mtp-Qwen3.8-Flash-Next-draft.gguf mmproj-F16.gguf --local-dir ~/models

# AMD (and Intel) GPUs
docker run --rm -it --device /dev/dri --group-add "$(getent group render | cut -d: -f3)" \
  -v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
  -m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
  -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --mmproj /models/mmproj-F16.gguf

# NVIDIA GPUs: the same, with the NVIDIA container toolkit and its Vulkan driver
docker run --rm -it --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all \
  -v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
  -m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
  -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --mmproj /models/mmproj-F16.gguf

Then open http://localhost:8080. For Gyro-M, download and pass Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf instead.

NVIDIA with Vulkan: many cloud GPU containers grant only compute,utility; then NVIDIA's Vulkan driver cannot start and llama.cpp silently falls back to the CPU (no usable GPU found). Keep -e NVIDIA_DRIVER_CAPABILITIES=all. Minimal images may also need libxext6 and libx11-6. For full speed on NVIDIA, build with CUDA (below).

2. Build from source

Needs Vulkan headers 1.4 or newer and glslc for the Vulkan build. On Ubuntu 22.04 the system packages are too old: install the LunarG Vulkan SDK and source its setup-env.sh first. NVIDIA: build with CUDA instead (CUDA toolkit 12.x+); the RTX 5090 and A6000 numbers above are from the CUDA build.

git clone https://github.com/agentionai/llama.cpp && cd llama.cpp
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j      # NVIDIA: -DGGML_CUDA=ON instead (CUDA toolkit 12.x+)
./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:Gyro-S -ngl 999 -fa on --jinja --ngram-on-disk \
    -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05

With -hf the vision projector downloads and loads automatically, and so does the MTP draft when you add --spec-type draft-mtp. Use :Gyro-M for the larger file.

3. Strata (release candidate)

Strata has Gyro support on the rc1 branch. It is the fastest way to run Gyro on an RTX 5090, and its expert cache runs Gyro on 24 GB cards. One script sets it up; the full guide is docs/GYRO.md.

git clone -b rc1 https://github.com/agentionai/Strata && cd Strata
tools/gyro_setup.sh --model S      # or --model M

Settings and tuning

  • →Sampling. Qwen's thinking-mode settings plus min-p: --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05. Min-p trims the low-probability tokens that a low-bit quant lifts slightly.
  • →Keep --ngram-on-disk in every command, with the model on fast NVMe. It keeps the n-gram table off the GPU and reads its rows with a fast parallel reader. Plain memory-mapping (--lazy-mode on) roughly halves prompt-processing speed.
  • →MTP draft for speculative decoding: download mtp-Qwen3.8-Flash-Next-draft.gguf from the repo (about 3.75 GiB of GPU memory) and add --spec-type draft-mtp -md mtp-Qwen3.8-Flash-Next-draft.gguf --spec-draft-n-max 6 --spec-draft-n-min 2 --spec-draft-mtp-vocab 32768. With -hf, leave out -md. On a single 32 GB card the draft fits next to Gyro-S up to 24k context with -ub 256 -ctkd q8_0 -ctvd q8_0.
  • →Two GPUs? Put the whole model on one and the MTP draft on the other: --device Vulkan0 --device-draft Vulkan1. Splitting the model's layers across both cards makes them take turns and is slower than one card.
  • →16–24 GB cards. --n-cpu-moe N keeps the experts of N layers in system RAM, about 0.55 GiB each. Start with N = 26 on 16 GB or N = 12 on 24 GB and lower it until the model just fits; drop --mmproj to save about 1 GiB. Speed depends on your RAM bandwidth and CPU. 32 GB of system RAM minimum, 64 GB recommended. Or use Strata's expert cache.
  • →Ampere (RTX 30-series, A-series) with CUDA: the batched prompt path can crash (MUL_MAT_ID failed, illegal memory access) when several requests with longer prompts run at once. Until the fix ships, start the server with GGML_CUDA_TQ_MMQ=0 or --parallel 1.
  • →Vision. Both files read images through mmproj-F16.gguf (about 1 GiB of GPU memory). To run text-only, drop --mmproj (with -m) or add --no-mmproj (with -hf).

Memory and context

What the GPU holds for Gyro-S (weights, KV cache, recurrent state and compute buffers), q8_0 KV cache, n-gram table on disk, one slot:

ContextGyro-S+ MTP draft
32k28.1 GiB~31.9 GiB
64k28.8 GiB~32.6 GiB
128k30.2 GiB~34.0 GiB
256k~33.3 GiB (extrapolated)~37.1 GiB

On a 32 GB card: up to 128k context without the MTP draft, 24k with it, or put the draft on a second GPU. Gyro-M needs about 37 GiB at 64k context.

Status and roadmap

  • ✓Gyro-S and Gyro-M: available.
  • ✓Abliterated Gyro-S and Gyro-M: available in their own repo.
  • …Gyro-L: coming. About 50 GiB, for 64 GB machines.
  • …Gyro Signal: coming. Shorter answers, no extra loops.
  • …Strata support: release candidate.

Gyro is experimental. The format and kernels may still change, and the files will be re-published when they do. We'd like to hear how it runs on your card.