Shared researchMade with Fieldwork Ledger · Publish your first results page free — beta

Hypothesis

At equal GPU memory, agentionai's Gyro files stay closer to the original model (lower KL divergence, higher top-1 agreement) than other published quants; the Pareto frontier of fidelity vs GPU-resident size is made of Gyro and AP files.

Charts

KL divergence vs GPU memory (lower-left is better)

scatter · 11 runs plotted · 1 excluded

0.11230.30830.5043GPU-resident weights (GiB)Mean KL divergence vs Q8_027.5747.4767.36

Colour: vendor. The points span 2 comparison contexts; each point's is in the source data.

Frontier chosen by hand: 3 runs, joined in GPU-resident weights (GiB) order. The dashed line connects recorded runs; it does not imply results between them.

How this chart is drawn

One point per successful run, no aggregation. Context series separate experiments, recorded comparison context, input references and environments; matching metadata does not establish experimental equivalence.

Displayed Y range: 0.1123 to 0.5043 · X range: 27.57 to 67.36

Schema b2b235ad-fe84-4c78-9970-147532f665f1 · 11 eligible runs · 1 excluded · Live as of 2026-10-08T09:29:39.894Z

Chart display (this view only)

Unchecked numeric axes include zero. Fitted axes may omit zero; bars always start at zero. Hiding a series rescales the graph, not the source table. Labels may overlap; full titles remain in tooltips and the table.

Source data and excluded runs
Exact values, source revisions and context series
RunExperimentRevisionGPU-resident weights (GiB)Mean KL divergence vs Q8_0SeriesContextFrontier
agentionai Gyro-S (TQ1_0): fidelity vs sizeShared experiment127.640.435165vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
agentionai Gyro-M (TQ2_0): fidelity vs sizeShared experiment134.980.305558vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
agentionai AP-Q4_K_XL: fidelity vs sizeShared experiment167.360.112272vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs sizeShared experiment136.520.420224vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"
ISTA-DASLab GSQ-RCO Q2_0: fidelity vs sizeShared experiment235.030.467805vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"
unsloth UD-Q2_K_XL: fidelity vs sizeShared experiment146.620.334233vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
unsloth UD-IQ3_XXS: fidelity vs sizeShared experiment149.50.264302vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
unsloth UD-Q3_K_XL: fidelity vs sizeShared experiment156.970.18548vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs sizeShared experiment143.80.310866vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs sizeShared experiment151.040.203961vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs sizeShared experiment127.570.504256vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
  • d1bbd7c5-6b44-4c18-899b-a8d730231ac2: Superseded by excluded-from-comparison

Top-1 agreement vs GPU memory (upper-left is better)

scatter · 11 runs plotted · 1 excluded

72.9679.0285.08GPU-resident weights (GiB)Top-1 agreement with Q8_0 (%)27.5747.4767.36

Colour: vendor. The points span 2 comparison contexts; each point's is in the source data.

Frontier chosen by hand: 3 runs, joined in GPU-resident weights (GiB) order. The dashed line connects recorded runs; it does not imply results between them.

How this chart is drawn

One point per successful run, no aggregation. Context series separate experiments, recorded comparison context, input references and environments; matching metadata does not establish experimental equivalence.

Displayed Y range: 72.96 to 85.08 · X range: 27.57 to 67.36

Schema b2b235ad-fe84-4c78-9970-147532f665f1 · 11 eligible runs · 1 excluded · Live as of 2026-10-08T09:29:39.902Z

Chart display (this view only)

Unchecked numeric axes include zero. Fitted axes may omit zero; bars always start at zero. Hiding a series rescales the graph, not the source table. Labels may overlap; full titles remain in tooltips and the table.

Source data and excluded runs
Exact values, source revisions and context series
RunExperimentRevisionGPU-resident weights (GiB)Top-1 agreement with Q8_0 (%)SeriesContextFrontier
agentionai Gyro-S (TQ1_0): fidelity vs sizeShared experiment127.6474.05vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
agentionai Gyro-M (TQ2_0): fidelity vs sizeShared experiment134.9878.117vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
agentionai AP-Q4_K_XL: fidelity vs sizeShared experiment167.3685.078vendor: agentionaihardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"Yes
ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs sizeShared experiment136.5274.22vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"
ISTA-DASLab GSQ-RCO Q2_0: fidelity vs sizeShared experiment235.0373.928vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp (Vulkan)"
unsloth UD-Q2_K_XL: fidelity vs sizeShared experiment146.6276.61vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
unsloth UD-IQ3_XXS: fidelity vs sizeShared experiment149.579.389vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
unsloth UD-Q3_K_XL: fidelity vs sizeShared experiment156.9781.931vendor: unslothhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs sizeShared experiment143.877.682vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs sizeShared experiment151.0481.315vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs sizeShared experiment127.5772.957vendor: ISTA-DASLabhardware="AMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memory" · protocol="KL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions." · runtime="agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c"
  • d1bbd7c5-6b44-4c18-899b-a8d730231ac2: Superseded by excluded-from-comparison

Results

Finished runs and what they recorded.
Runkld mean ↓ betterkld median ↓ bettertop1 pct (%) ↑ bettergpu resident gib (GiB) ↓ betterfile gib (GiB)kld code ↓ betterkld reasoning ↓ bettermmlu pro (%) ↑ bettermeasured yyyymmdd
agentionai AP-Q4_K_XL: fidelity vs size0.1122720.03269185.07867.3694.180.0422710.08364981.4320,260,900
agentionai Gyro-M (TQ2_0): fidelity vs size0.3055580.0897478.11734.9885.650.1213490.2424318020,261,000
agentionai Gyro-S (TQ1_0): fidelity vs size0.4351650.14050374.0527.6454.470.1881470.38616678.120,261,000
ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size0.5042560.16677572.95727.5754.40.2073320.40311Missing20,261,000
ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size0.4202240.13920174.2236.5263.340.1861780.396463Missing20,260,900
ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size0.2039610.05771881.31551.0477.880.0946990.196154Missing20,261,000
ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size0.3108660.09023877.68243.870.630.1306860.292607Missing20,261,000
ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size0.4678050.1498973.92835.0361.860.2093220.45867580.9520,261,000
unsloth UD-IQ3_XXS: fidelity vs size0.2643020.07573279.38949.576.330.108210.243885Missing20,261,000
unsloth UD-Q2_K_XL: fidelity vs size0.3342330.10827776.6146.6273.450.1392290.317638Missing20,261,000
unsloth UD-Q3_K_XL: fidelity vs size0.185480.05231581.93156.9783.810.0768490.166222Missing20,261,000

Finished runs and what they recorded.

agentionai AP-Q4_K_XL: fidelity vs sizekld mean 0.112272
kld mean ↓
0.112272
kld median ↓
0.032691
top1 pct (%) ↑
85.078
gpu resident gib (GiB) ↓
67.36
file gib (GiB)
94.18
kld code ↓
0.042271
kld reasoning ↓
0.083649
mmlu pro (%) ↑
81.43
measured yyyymmdd
20,260,900
agentionai Gyro-M (TQ2_0): fidelity vs sizekld mean 0.305558
kld mean ↓
0.305558
kld median ↓
0.08974
top1 pct (%) ↑
78.117
gpu resident gib (GiB) ↓
34.98
file gib (GiB)
85.65
kld code ↓
0.121349
kld reasoning ↓
0.242431
mmlu pro (%) ↑
80
measured yyyymmdd
20,261,000
agentionai Gyro-S (TQ1_0): fidelity vs sizekld mean 0.435165
kld mean ↓
0.435165
kld median ↓
0.140503
top1 pct (%) ↑
74.05
gpu resident gib (GiB) ↓
27.64
file gib (GiB)
54.47
kld code ↓
0.188147
kld reasoning ↓
0.386166
mmlu pro (%) ↑
78.1
measured yyyymmdd
20,261,000
ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs sizekld mean 0.504256
kld mean ↓
0.504256
kld median ↓
0.166775
top1 pct (%) ↑
72.957
gpu resident gib (GiB) ↓
27.57
file gib (GiB)
54.4
kld code ↓
0.207332
kld reasoning ↓
0.40311
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs sizekld mean 0.420224
kld mean ↓
0.420224
kld median ↓
0.139201
top1 pct (%) ↑
74.22
gpu resident gib (GiB) ↓
36.52
file gib (GiB)
63.34
kld code ↓
0.186178
kld reasoning ↓
0.396463
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,260,900
ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs sizekld mean 0.203961
kld mean ↓
0.203961
kld median ↓
0.057718
top1 pct (%) ↑
81.315
gpu resident gib (GiB) ↓
51.04
file gib (GiB)
77.88
kld code ↓
0.094699
kld reasoning ↓
0.196154
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs sizekld mean 0.310866
kld mean ↓
0.310866
kld median ↓
0.090238
top1 pct (%) ↑
77.682
gpu resident gib (GiB) ↓
43.8
file gib (GiB)
70.63
kld code ↓
0.130686
kld reasoning ↓
0.292607
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
ISTA-DASLab GSQ-RCO Q2_0: fidelity vs sizekld mean 0.467805
kld mean ↓
0.467805
kld median ↓
0.14989
top1 pct (%) ↑
73.928
gpu resident gib (GiB) ↓
35.03
file gib (GiB)
61.86
kld code ↓
0.209322
kld reasoning ↓
0.458675
mmlu pro (%) ↑
80.95
measured yyyymmdd
20,261,000
unsloth UD-IQ3_XXS: fidelity vs sizekld mean 0.264302
kld mean ↓
0.264302
kld median ↓
0.075732
top1 pct (%) ↑
79.389
gpu resident gib (GiB) ↓
49.5
file gib (GiB)
76.33
kld code ↓
0.10821
kld reasoning ↓
0.243885
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
unsloth UD-Q2_K_XL: fidelity vs sizekld mean 0.334233
kld mean ↓
0.334233
kld median ↓
0.108277
top1 pct (%) ↑
76.61
gpu resident gib (GiB) ↓
46.62
file gib (GiB)
73.45
kld code ↓
0.139229
kld reasoning ↓
0.317638
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
unsloth UD-Q3_K_XL: fidelity vs sizekld mean 0.18548
kld mean ↓
0.18548
kld median ↓
0.052315
top1 pct (%) ↑
81.931
gpu resident gib (GiB) ↓
56.97
file gib (GiB)
83.81
kld code ↓
0.076849
kld reasoning ↓
0.166222
mmlu pro (%) ↑
Missing
measured yyyymmdd
20,261,000
The parameters of each run
Parameters each run was produced with, and the context it was measured in. Values marked "added later" were filled in after the run executed.
Runexperts keptfileoursvendorhardware (context)protocol (context)runtime (context)
agentionai Gyro-S (TQ1_0): fidelity vs size512Gyro-S (TQ1_0)trueagentionaiAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp (Vulkan)
agentionai Gyro-M (TQ2_0): fidelity vs size512Gyro-M (TQ2_0)trueagentionaiAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp (Vulkan)
agentionai AP-Q4_K_XL: fidelity vs size512AP-Q4_K_XLtrueagentionaiAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp (Vulkan)
ISTA-DASLab GSQ-RCO IQ2_XS: fidelity vs size512GSQ-RCO IQ2_XSfalseISTA-DASLabAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp (Vulkan)
ISTA-DASLab GSQ-RCO Q2_0: fidelity vs size512GSQ-RCO Q2_0falseISTA-DASLabAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp (Vulkan)
unsloth UD-Q2_K_XL: fidelity vs size512UD-Q2_K_XLfalseunslothAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
unsloth UD-IQ3_XXS: fidelity vs size512UD-IQ3_XXSfalseunslothAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
unsloth UD-Q3_K_XL: fidelity vs size512UD-Q3_K_XLfalseunslothAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
ISTA-DASLab GSQ-RCO IQ3_XXS: fidelity vs size512GSQ-RCO IQ3_XXSfalseISTA-DASLabAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
ISTA-DASLab GSQ-RCO IQ3_S: fidelity vs size512GSQ-RCO IQ3_SfalseISTA-DASLabAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
ISTA-DASLab GSQ-RCO Coder IQ1_M: fidelity vs size256GSQ-RCO Coder IQ1_MfalseISTA-DASLabAMD Strix Halo (Radeon 8060S, gfx1151), 128 GB unified memoryKL divergence and top-1 token agreement against the unsloth Q8_0 original on 60 x 2,048 held-out tokens (2026 technical writing and code), llama-perplexity --kl-divergence, n-gram table read from disk. Code/reasoning: same measure on pagoda-task code and reasoning traces written by other models. MMLU-Pro: 70 questions x 3 repetitions.agentionai/llama.cpp dev build (Vulkan), b11330 d8841825c
1 run is not shown: unfinished or replaced
  • DJLougen Mooney: fidelity vs size succeededsuperseded

How this was measured

Goal

Shareable comparison of published Qwen3.8-Flash-Next quantisations (agentionai and other vendors): output fidelity against the Q8_0 original versus GPU memory, behaviour, and speed by hardware and engine. Measurements and evaluation protocol only.