Qwen3.8-27B on 12, 16, 24 and 32 GB: what fits, what runs, and what each lever costs
Five consumer cards, one model, 91 measured points
Date
2026-09-09
Hardware
RTX 3060 · RTX 3090 · RTX 4070 Ti SUPER · RTX 4090 · RTX 5090
Everything below is published as the harness produced it: no post-processing,
no rounding, no re-ordering. The artifact table, the harness output and the
checksums are read from the files themselves at render time, so nothing on this
page can drift away from what was measured.
datasets/ggml-org/ci at 927b3642933080f1b0e811e2f916e14c292992f9, member wikitext-2-raw/wiki.test.raw
Per card
Read from each card's own bench.json. Nothing is filled in from a
neighbouring card: a field the harness did not record reads not reported.
Where a ladder failed before it could record the build, the commit comes from that
same card's sweeps, which did record it. The quant is not listed here — one card ran
several — it is named on each table of points below.
Prices come from the run notes, except the 4070 Ti SUPER, read off its published
instance panel frame.
Card
llama.cpp
Build
Series flags
Protocol
Taken on
Price
NVIDIA GeForce RTX 3060
df750f76b
b10871
-p 512 -n 128 -r 3 -t 8
bench-runner/1
2026-09-09
$0.221/hr
NVIDIA GeForce RTX 3090
df750f76b
b10871
-p 512 -n 128 -r 3 -t 8
bench-runner/1
2026-09-09
$0.172/hr
NVIDIA GeForce RTX 4070 Ti SUPER
df750f76b
b10871
-p 512 -n 128 -r 3 -t 8
bench-runner/1
2026-09-09
$0.131/hr
NVIDIA GeForce RTX 4090
df750f76b
b10871
-p 512 -n 128 -r 3 -t 8
bench-runner/1
2026-09-09
$0.390/hr
NVIDIA GeForce RTX 5090
df750f76b
b10871
-p 512 -n 128 -r 3 -t 8
bench-runner/1
2026-09-09
$0.074/hr
Weights and quants are not hosted here. The repository and revision above, plus the
llama-quantize command and sha256 of every quant below, are enough to
rebuild the same files byte for byte.
16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
256
256k
—
—
17.5 GiB
16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
128
256k
—
—
17.5 GiB
16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
NVIDIA GeForce RTX 5090 — KV cache type — Qwen3.8-27B-Q4_K_M.gguf
KV cache K
KV cache V
Prefill depth
Prompt processing
Token generation
Peak VRAM
Status
f16
f16
256k
—
—
17.5 GiB
16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
q8_0
f16
256k
—
—
17.5 GiB
13406.27 MiB short: failed to allocate CUDA0 buffer of size 14057490560
q4_0
f16
256k
—
—
17.5 GiB
13406.27 MiB short: failed to allocate CUDA0 buffer of size 14057490560
f16
q8_0
256k
—
—
30.5 GiB
timeout
Quality axis
Every quant against the Qwen3.8-27B-BF16.gguf baseline,
perplexity 6.782793 ± 0.073775, on a fixed text:
datasets/ggml-org/ci at revision 927b3642933080f1b0e811e2f916e14c292992f9,
member wikitext-2-raw/wiki.test.raw.
Quant
Perplexity
Ratio to baseline
Mean KLD
Median KLD
RMS Δp
Same top p
Peak VRAM
Status
Qwen3.8-27B-IQ4_XS.gguf
6.869776 ± 0.075071
1.012824
0.023808
0.010319
4.34 %
93.151 %
15.3 GiB
ok
Qwen3.8-27B-Q2_K.gguf
8.455446 ± 0.09406
1.246602
0.294271
0.148913
15.772 %
77.227 %
11.4 GiB
ok
Qwen3.8-27B-Q3_K_M.gguf
7.173957 ± 0.079915
1.05767
0.071103
0.032325
7.458 %
88.553 %
13.7 GiB
ok
Qwen3.8-27B-Q4_0.gguf
6.945883 ± 0.076069
1.024045
0.036346
0.016208
5.175 %
91.612 %
15.5 GiB
ok
Qwen3.8-27B-Q4_K_M.gguf
6.814875 ± 0.074022
1.00473
0.023021
0.009979
4.164 %
93.327 %
16.5 GiB
ok
Qwen3.8-27B-Q5_K_M.gguf
6.820789 ± 0.074328
1.005602
0.008207
0.003519
2.501 %
96.008 %
18.9 GiB
ok
Qwen3.8-27B-Q6_K.gguf
6.788637 ± 0.073886
1.000862
0.002654
0.0012
1.458 %
97.698 %
21.4 GiB
ok
Qwen3.8-27B-Q8_0.gguf
6.790667 ± 0.073943
1.001161
0.000797
0.000326
0.814 %
98.71 %
27.2 GiB
ok
Checksums
Source Qwen3.8-27B-BF16.gguf, sha256 5a3eedc837bcbd1365cdbf5b71e698df3122e76586ba872f07ce3ed4a9bfa97e,
quantised on 2026-09-09.