Measured points

Qwen3.8-27B on 12, 16, 24 and 32 GB: what fits, what runs, and what each lever costs

Five consumer cards, one model, 91 measured points

Date
2026-09-09
Hardware
RTX 3060 · RTX 3090 · RTX 4070 Ti SUPER · RTX 4090 · RTX 5090

Everything below is published as the harness produced it: no post-processing, no rounding, no re-ordering. The artifact table, the harness output and the checksums are read from the files themselves at render time, so nothing on this page can drift away from what was measured.

Conditions

ModelQwen3.8-27B
Repositoryggml-org/Qwen3.8-27B-GGUF
Revision0669b98607d47046c7c2b3f801011d54a08cfccf
Usable threshold10 TPS on token generation
Quality textdatasets/ggml-org/ci at 927b3642933080f1b0e811e2f916e14c292992f9, member wikitext-2-raw/wiki.test.raw

Per card

Read from each card's own bench.json. Nothing is filled in from a neighbouring card: a field the harness did not record reads not reported. Where a ladder failed before it could record the build, the commit comes from that same card's sweeps, which did record it. The quant is not listed here — one card ran several — it is named on each table of points below. Prices come from the run notes, except the 4070 Ti SUPER, read off its published instance panel frame.

Cardllama.cppBuildSeries flagsProtocolTaken onPrice
NVIDIA GeForce RTX 3060 df750f76b b10871 -p 512 -n 128 -r 3 -t 8 bench-runner/1 2026-09-09 $0.221/hr
NVIDIA GeForce RTX 3090 df750f76b b10871 -p 512 -n 128 -r 3 -t 8 bench-runner/1 2026-09-09 $0.172/hr
NVIDIA GeForce RTX 4070 Ti SUPER df750f76b b10871 -p 512 -n 128 -r 3 -t 8 bench-runner/1 2026-09-09 $0.131/hr
NVIDIA GeForce RTX 4090 df750f76b b10871 -p 512 -n 128 -r 3 -t 8 bench-runner/1 2026-09-09 $0.390/hr
NVIDIA GeForce RTX 5090 df750f76b b10871 -p 512 -n 128 -r 3 -t 8 bench-runner/1 2026-09-09 $0.074/hr

Weights and quants are not hosted here. The repository and revision above, plus the llama-quantize command and sha256 of every quant below, are enough to rebuild the same files byte for byte.

Downloads

Ladders

91 measured points in total across the tables on this page.

NVIDIA GeForce RTX 3060 — Qwen3.8-27B-Q4_K_M.gguf from the repository

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 — — 0.1 GiB 17402.38 MiB short: failed to allocate CUDA0 buffer of size 18247720960
32k 32768 — — — skipped
96k 98304 — — — skipped
256k 262144 — — — skipped

NVIDIA GeForce RTX 3090 — Qwen3.8-27B-Q4_K_M.gguf from the repository

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 1213.3 TPS 37.1 TPS 18.5 GiB ok
32k 32768 955.1 TPS 34.3 TPS 20 GiB ok
96k 98304 — — 17.3 GiB oom_prefill
256k 262144 — — — skipped

NVIDIA GeForce RTX 4070 Ti SUPER — Qwen3.8-27B-Q4_K_M.gguf from the repository

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 — — 0.2 GiB 17402.38 MiB short: failed to allocate CUDA0 buffer of size 18247720960
32k 32768 — — — skipped
96k 98304 — — — skipped
256k 262144 — — — skipped

NVIDIA GeForce RTX 4070 Ti SUPER — Qwen3.8-27B-Q4_K_M.gguf quantised by us, 15.41 GiB

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
32k 32768 — — — skipped
96k 98304 — — — skipped
256k 262144 — — — skipped

NVIDIA GeForce RTX 4090 — Qwen3.8-27B-Q4_K_M.gguf from the repository

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 2828.8 TPS 42.9 TPS 18.6 GiB ok
32k 32768 2343 TPS 40 TPS 20.1 GiB ok
96k 98304 — — 17.4 GiB oom_prefill
256k 262144 — — — skipped

NVIDIA GeForce RTX 5090 — Qwen3.8-27B-Q4_K_M.gguf from the repository

Prefill depth Tokens Prompt processing Token generation Peak VRAM Status
8k 8192 3696.2 TPS 67.7 TPS 18.8 GiB ok
32k 32768 2903.8 TPS 63.5 TPS 20.2 GiB ok
96k 98304 1536 TPS 54.5 TPS 24.2 GiB ok
256k 262144 — — 17.5 GiB 16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616

Sweeps

NVIDIA GeForce RTX 3060 — Layers on the GPU — Qwen3.8-27B-Q4_K_M.gguf

Layers on GPU Prefill depth Prompt processing Token generation Peak VRAM Status
-1 8k — — 0.1 GiB 17402.38 MiB short: failed to allocate CUDA0 buffer of size 18247720960
64 8k — — 0.1 GiB 17141.38 MiB short: failed to allocate CUDA0 buffer of size 17974035456
56 8k — — 0.1 GiB 15090.41 MiB short: failed to allocate CUDA0 buffer of size 15823440896
48 8k — — 0.1 GiB 13039.44 MiB short: failed to allocate CUDA0 buffer of size 13672846336
40 8k — — 10.8 GiB 513.51 MiB short: failed to allocate CUDA0 buffer of size 538454144
32 8k 273.1 TPS 2.6 TPS 9.8 GiB ok
24 8k 244.3 TPS 2.1 TPS 7.7 GiB ok
16 8k 223 TPS 1.8 TPS 5.6 GiB ok
8 8k 202.9 TPS 1.6 TPS 3.5 GiB ok
0 8k — — 1.8 GiB failed

NVIDIA GeForce RTX 3060 — Single point — Qwen3.8-27B-Q4_K_M.gguf

Prefill depth Prompt processing Token generation Peak VRAM Status
8k 185.5 TPS 1.4 TPS 1.8 GiB ok

NVIDIA GeForce RTX 3060 — Micro-batch — Qwen3.8-27B-Q4_K_M.gguf

Micro-batch Prefill depth Prompt processing Token generation Peak VRAM Status
512 8k — — 10.8 GiB 513.51 MiB short: failed to allocate CUDA0 buffer of size 538454144
256 8k 220.2 TPS 3.3 TPS 11.6 GiB ok
128 8k 142.7 TPS 3.3 TPS 11.5 GiB ok
64 8k 80.9 TPS 3.3 TPS 11.4 GiB ok

NVIDIA GeForce RTX 3060 — KV cache type — Qwen3.8-27B-Q4_K_M.gguf

KV cache K Prefill depth Prompt processing Token generation Peak VRAM Status
f16 8k 221.7 TPS 3.2 TPS 11.6 GiB ok
q8_0 8k — — 11.6 GiB oom_prefill
q4_0 8k 212.7 TPS 2.9 TPS 11.6 GiB ok

NVIDIA GeForce RTX 3060 — KV cache type — Qwen3.8-27B-Q4_K_M.gguf

KV cache K Prefill depth Prompt processing Token generation Peak VRAM Status
f16 32k — — 10.8 GiB 1300 MiB short: failed to allocate CUDA0 buffer of size 1363148800
q8_0 32k — — 10.8 GiB 995.31 MiB short: failed to allocate CUDA0 buffer of size 1043660800
q4_0 32k — — 10.8 GiB 832.81 MiB short: failed to allocate CUDA0 buffer of size 873267200

NVIDIA GeForce RTX 3060 — Layers on the GPU — Qwen3.8-27B-Q4_K_M.gguf

Layers on GPU Prefill depth Prompt processing Token generation Peak VRAM Status
-1 32k — — 0.1 GiB 17402.38 MiB short: failed to allocate CUDA0 buffer of size 18247720960
36 32k 185.2 TPS 2.6 TPS 10.5 GiB ok
32 32k 171.8 TPS 2.3 TPS 9.5 GiB ok
28 32k 161.1 TPS 2.1 TPS 8.5 GiB ok

NVIDIA GeForce RTX 4070 Ti SUPER — Micro-batch — Qwen3.8-27B-Q4_K_M.gguf

Micro-batch Prefill depth Prompt processing Token generation Peak VRAM Status
512 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
256 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
128 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
64 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184

NVIDIA GeForce RTX 4070 Ti SUPER — Layers on the GPU — Qwen3.8-27B-Q4_K_M.gguf

Layers on GPU Prefill depth Prompt processing Token generation Peak VRAM Status
-1 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
62 8k 1313 TPS 20.8 TPS 15.5 GiB ok
60 8k 1193.5 TPS 15.5 TPS 15 GiB ok
56 8k 1020.3 TPS 10.8 TPS 14.1 GiB ok

NVIDIA GeForce RTX 4070 Ti SUPER — Prefill depth — Qwen3.8-27B-IQ4_XS.gguf

Depth Prefill depth Prompt processing Token generation Peak VRAM Status
0 0 1772.5 TPS 37.8 TPS 14.5 GiB ok
8192 8k 1650.6 TPS 36.6 TPS 15 GiB ok
32768 32k — — 13.7 GiB 2080 MiB short: failed to allocate CUDA0 buffer of size 2181038080

NVIDIA GeForce RTX 4070 Ti SUPER — KV cache type — Qwen3.8-27B-IQ4_XS.gguf

KV cache K KV cache V Prefill depth Prompt processing Token generation Peak VRAM Status
f16 f16 32k — — 13.7 GiB 2080 MiB short: failed to allocate CUDA0 buffer of size 2181038080
q8_0 f16 32k — — 13.7 GiB 1758.27 MiB short: failed to allocate CUDA0 buffer of size 1843677312
q4_0 f16 32k — — 13.7 GiB 1758.27 MiB short: failed to allocate CUDA0 buffer of size 1843677312
f16 q8_0 32k — — 13.7 GiB 505 MiB short: failed to allocate CUDA0 buffer of size 529532928

NVIDIA GeForce RTX 4070 Ti SUPER — Single point — Qwen3.8-27B-IQ4_XS.gguf

Prefill depth Prompt processing Token generation Peak VRAM Status
32k 1325.2 TPS 33 TPS 15.2 GiB ok

NVIDIA GeForce RTX 4070 Ti SUPER — Flash attention — Qwen3.8-27B-IQ4_XS.gguf

Flash attention Prefill depth Prompt processing Token generation Peak VRAM Status
auto 8k 1655.7 TPS 36.6 TPS 15 GiB ok
on 8k 1652.3 TPS 36.6 TPS 15 GiB ok
off 8k 1413.5 TPS 35.6 TPS 15.2 GiB ok

NVIDIA GeForce RTX 4070 Ti SUPER — Threads — Qwen3.8-27B-Q4_K_M.gguf

Threads Prefill depth Prompt processing Token generation Peak VRAM Status
8 8k 1306 TPS 20.7 TPS 15.5 GiB ok
1 8k 1303.9 TPS 14.6 TPS 15.5 GiB ok
2 8k 1301.4 TPS 19.2 TPS 15.5 GiB ok
4 8k 1301.7 TPS 20.7 TPS 15.5 GiB ok
16 8k 1293.9 TPS 18.8 TPS 15.5 GiB ok

NVIDIA GeForce RTX 4070 Ti SUPER — Tensor overrides — Qwen3.8-27B-Q4_K_M.gguf

Tensor override Prefill depth Prompt processing Token generation Peak VRAM Status
default 8k — — 15 GiB 149.62 MiB short: failed to allocate CUDA0 buffer of size 156893184
blk\.(3|7|11|15|19|23|27|31|35|39|43|47|51|55|59|63)\..*=CPU 8k 894.4 TPS 9 TPS 12.9 GiB ok
\.ffn_down\.=CPU 8k 849.8 TPS 8.4 TPS 12.6 GiB ok

NVIDIA GeForce RTX 5090 — Micro-batch — Qwen3.8-27B-Q4_K_M.gguf

Micro-batch Prefill depth Prompt processing Token generation Peak VRAM Status
512 256k — — 17.5 GiB 16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
256 256k — — 17.5 GiB 16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
128 256k — — 17.5 GiB 16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616

NVIDIA GeForce RTX 5090 — KV cache type — Qwen3.8-27B-Q4_K_M.gguf

KV cache K KV cache V Prefill depth Prompt processing Token generation Peak VRAM Status
f16 f16 256k — — 17.5 GiB 16416 MiB short: failed to allocate CUDA0 buffer of size 17213423616
q8_0 f16 256k — — 17.5 GiB 13406.27 MiB short: failed to allocate CUDA0 buffer of size 14057490560
q4_0 f16 256k — — 17.5 GiB 13406.27 MiB short: failed to allocate CUDA0 buffer of size 14057490560
f16 q8_0 256k — — 30.5 GiB timeout

Quality axis

Every quant against the Qwen3.8-27B-BF16.gguf baseline, perplexity 6.782793 ± 0.073775, on a fixed text: datasets/ggml-org/ci at revision 927b3642933080f1b0e811e2f916e14c292992f9, member wikitext-2-raw/wiki.test.raw.

QuantPerplexityRatio to baselineMean KLD Median KLDRMS ΔpSame top pPeak VRAMStatus
Qwen3.8-27B-IQ4_XS.gguf 6.869776 ± 0.075071 1.012824 0.023808 0.010319 4.34 % 93.151 % 15.3 GiB ok
Qwen3.8-27B-Q2_K.gguf 8.455446 ± 0.09406 1.246602 0.294271 0.148913 15.772 % 77.227 % 11.4 GiB ok
Qwen3.8-27B-Q3_K_M.gguf 7.173957 ± 0.079915 1.05767 0.071103 0.032325 7.458 % 88.553 % 13.7 GiB ok
Qwen3.8-27B-Q4_0.gguf 6.945883 ± 0.076069 1.024045 0.036346 0.016208 5.175 % 91.612 % 15.5 GiB ok
Qwen3.8-27B-Q4_K_M.gguf 6.814875 ± 0.074022 1.00473 0.023021 0.009979 4.164 % 93.327 % 16.5 GiB ok
Qwen3.8-27B-Q5_K_M.gguf 6.820789 ± 0.074328 1.005602 0.008207 0.003519 2.501 % 96.008 % 18.9 GiB ok
Qwen3.8-27B-Q6_K.gguf 6.788637 ± 0.073886 1.000862 0.002654 0.0012 1.458 % 97.698 % 21.4 GiB ok
Qwen3.8-27B-Q8_0.gguf 6.790667 ± 0.073943 1.001161 0.000797 0.000326 0.814 % 98.71 % 27.2 GiB ok

Checksums

Source Qwen3.8-27B-BF16.gguf, sha256 5a3eedc837bcbd1365cdbf5b71e698df3122e76586ba872f07ce3ed4a9bfa97e, quantised on 2026-09-09.

QuantOn disksha256
Qwen3.8-27B-Q8_0.gguf 26.63 GiB f5c702d8820d36fb55985bb238fc83ee3a313e920f4b752a437c3a6a9e14e4c8
Qwen3.8-27B-Q6_K.gguf 20.57 GiB b40c8bc7e60c0481acb9e832b2515960cd268276745485e12ec88c816cbce4e6
Qwen3.8-27B-Q5_K_M.gguf 17.91 GiB 7737ab5692fc265a8b6b83b7635139d2cc8706673d0cf414f43141aec5b20842
Qwen3.8-27B-Q4_K_M.gguf 15.41 GiB 078fc466c6de18416f3515a108fd58a8c33fbe2215e730db661a2fe2fcd9c5f0
Qwen3.8-27B-Q4_0.gguf 14.41 GiB f7c1c020d7ecc58e7cc5d5f9a3f1948d38658c73ac89cd751d2f1fe6570f1e33
Qwen3.8-27B-Q3_K_M.gguf 12.39 GiB aa4d1df596360a425c9b5b9356ee8f43909376d831359916dea5e8cc5f18679a
Qwen3.8-27B-IQ4_XS.gguf 14.15 GiB e6e9b2c202a1673a4243452bb65c5c3c84b210ba57e582109fdea786aa69cb5a
Qwen3.8-27B-Q2_K.gguf 9.98 GiB ce220599529e4740552ed764999569c641476fde80233cd568ffdfd710c37dad