Hypothesis
At equal GPU memory, agentionai's Gyro files stay closer to the original model (lower KL divergence, higher top-1 agreement) than other published quants; the Pareto frontier of fidelity vs GPU-resident size is made of Gyro and AP files.
Shared results · Qwen3.8-Flash-Next quant comparison
Published 6 October 2026, updated 8 October 2026 as runs are recorded · 11 finished runs · 2 charts
Colour: vendor. The points span 2 comparison contexts; each point's is in the source data.
Frontier chosen by hand: 3 runs, joined in GPU-resident weights (GiB) order. The dashed line connects recorded runs; it does not imply results between them.
One point per successful run, no aggregation. Context series separate experiments, recorded comparison context, input references and environments; matching metadata does not establish experimental equivalence.
Displayed Y range: 0.1123 to 0.5043 · X range: 27.57 to 67.36
Schema b2b235ad-fe84-4c78-9970-147532f665f1 · 11 eligible runs · 1 excluded · Live as of 2026-10-08T09:29:39.894Z
| Run | Experiment | Revision | GPU-resident weights (GiB) | Mean KL divergence vs Q8_0 | Series | Context | Frontier |
|---|---|---|---|---|---|---|---|
| agentionai Gyro-S (TQ1_0): fidelity vs size | Shared experiment | 1 | 27.64 | 0.435165 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| agentionai Gyro-M (TQ2_0): fidelity vs size | Shared experiment | 1 | 34.98 | 0.305558 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| agentionai AP-Q4_K_XL: fidelity vs size | Shared experiment | 1 | 67.36 | 0.112272 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size | Shared experiment | 1 | 36.52 | 0.420224 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | |
| ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size | Shared experiment | 2 | 35.03 | 0.467805 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | |
| unsloth UD-Q2_K_XL: fidelity vs size | Shared experiment | 1 | 46.62 | 0.334233 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| unsloth UD-IQ3_XXS: fidelity vs size | Shared experiment | 1 | 49.5 | 0.264302 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| unsloth UD-Q3_K_XL: fidelity vs size | Shared experiment | 1 | 56.97 | 0.18548 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size | Shared experiment | 1 | 43.8 | 0.310866 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size | Shared experiment | 1 | 51.04 | 0.203961 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size | Shared experiment | 1 | 27.57 | 0.504256 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" |
Colour: vendor. The points span 2 comparison contexts; each point's is in the source data.
Frontier chosen by hand: 3 runs, joined in GPU-resident weights (GiB) order. The dashed line connects recorded runs; it does not imply results between them.
One point per successful run, no aggregation. Context series separate experiments, recorded comparison context, input references and environments; matching metadata does not establish experimental equivalence.
Displayed Y range: 72.96 to 85.08 · X range: 27.57 to 67.36
Schema b2b235ad-fe84-4c78-9970-147532f665f1 · 11 eligible runs · 1 excluded · Live as of 2026-10-08T09:29:39.902Z
| Run | Experiment | Revision | GPU-resident weights (GiB) | Top-1 agreement with Q8_0 (%) | Series | Context | Frontier |
|---|---|---|---|---|---|---|---|
| agentionai Gyro-S (TQ1_0): fidelity vs size | Shared experiment | 1 | 27.64 | 74.05 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| agentionai Gyro-M (TQ2_0): fidelity vs size | Shared experiment | 1 | 34.98 | 78.117 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| agentionai AP-Q4_K_XL: fidelity vs size | Shared experiment | 1 | 67.36 | 85.078 | vendor: agentionai | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | Yes |
| ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size | Shared experiment | 1 | 36.52 | 74.22 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | |
| ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size | Shared experiment | 2 | 35.03 | 73.928 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)" | |
| unsloth UD-Q2_K_XL: fidelity vs size | Shared experiment | 1 | 46.62 | 76.61 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| unsloth UD-IQ3_XXS: fidelity vs size | Shared experiment | 1 | 49.5 | 79.389 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| unsloth UD-Q3_K_XL: fidelity vs size | Shared experiment | 1 | 56.97 | 81.931 | vendor: unsloth | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size | Shared experiment | 1 | 43.8 | 77.682 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size | Shared experiment | 1 | 51.04 | 81.315 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" | |
| ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size | Shared experiment | 1 | 27.57 | 72.957 | vendor: ISTA-DASLab | hardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c" |
| Run | kld mean ↓ better | kld median ↓ better | top1 pct (%) ↑ better | gpu resident gib (GiB) ↓ better | file gib (GiB) | kld code ↓ better | kld reasoning ↓ better | mmlu pro (%) ↑ better | measured yyyymmdd |
|---|---|---|---|---|---|---|---|---|---|
| agentionai AP-Q4_K_XL: fidelity vs size | 0.112272 | 0.032691 | 85.078 | 67.36 | 94.18 | 0.042271 | 0.083649 | 81.43 | 20,260,900 |
| agentionai Gyro-M (TQ2_0): fidelity vs size | 0.305558 | 0.08974 | 78.117 | 34.98 | 85.65 | 0.121349 | 0.242431 | 80 | 20,261,000 |
| agentionai Gyro-S (TQ1_0): fidelity vs size | 0.435165 | 0.140503 | 74.05 | 27.64 | 54.47 | 0.188147 | 0.386166 | 78.1 | 20,261,000 |
| ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size | 0.504256 | 0.166775 | 72.957 | 27.57 | 54.4 | 0.207332 | 0.40311 | Missing | 20,261,000 |
| ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size | 0.420224 | 0.139201 | 74.22 | 36.52 | 63.34 | 0.186178 | 0.396463 | Missing | 20,260,900 |
| ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size | 0.203961 | 0.057718 | 81.315 | 51.04 | 77.88 | 0.094699 | 0.196154 | Missing | 20,261,000 |
| ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size | 0.310866 | 0.090238 | 77.682 | 43.8 | 70.63 | 0.130686 | 0.292607 | Missing | 20,261,000 |
| ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size | 0.467805 | 0.14989 | 73.928 | 35.03 | 61.86 | 0.209322 | 0.458675 | 80.95 | 20,261,000 |
| unsloth UD-IQ3_XXS: fidelity vs size | 0.264302 | 0.075732 | 79.389 | 49.5 | 76.33 | 0.10821 | 0.243885 | Missing | 20,261,000 |
| unsloth UD-Q2_K_XL: fidelity vs size | 0.334233 | 0.108277 | 76.61 | 46.62 | 73.45 | 0.139229 | 0.317638 | Missing | 20,261,000 |
| unsloth UD-Q3_K_XL: fidelity vs size | 0.18548 | 0.052315 | 81.931 | 56.97 | 83.81 | 0.076849 | 0.166222 | Missing | 20,261,000 |
Finished runs and what they recorded.
| Run | experts kept | file | ours | vendor | hardware (context) | protocol (context) | runtime (context) |
|---|---|---|---|---|---|---|---|
| agentionai Gyro-S (TQ1_0): fidelity vs size | 512 | Gyro-S (TQ1_0) | true | agentionai | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp (Vulkan) |
| agentionai Gyro-M (TQ2_0): fidelity vs size | 512 | Gyro-M (TQ2_0) | true | agentionai | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp (Vulkan) |
| agentionai AP-Q4_K_XL: fidelity vs size | 512 | AP-Q4_K_XL | true | agentionai | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp (Vulkan) |
| ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size | 512 | GSQ-RCO IQ2_XS | false | ISTA-DASLab | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp (Vulkan) |
| ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size | 512 | GSQ-RCO Q2_0 | false | ISTA-DASLab | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp (Vulkan) |
| unsloth UD-Q2_K_XL: fidelity vs size | 512 | UD-Q2_K_XL | false | unsloth | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
| unsloth UD-IQ3_XXS: fidelity vs size | 512 | UD-IQ3_XXS | false | unsloth | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
| unsloth UD-Q3_K_XL: fidelity vs size | 512 | UD-Q3_K_XL | false | unsloth | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
| ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size | 512 | GSQ-RCO IQ3_XXS | false | ISTA-DASLab | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
| ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size | 512 | GSQ-RCO IQ3_S | false | ISTA-DASLab | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
| ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size | 256 | GSQ-RCO Coder IQ1_M | false | ISTA-DASLab | AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory | KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions. | agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c |
Shareable comparison of published Qwen3.8-Flash-Next quantisations (agentionai and other vendors): output fidelity against the Q8_0 original versus GPU memory, behaviour, and speed by hardware and engine. Measurements and evaluation protocol only.