Five consumer cards, one model, 91 measured points, taken on 2026-09-09. The short version: memory decides the ceiling, but three settings that cost nothing decide whether a 16 GB card does 33 TPS at 32k or refuses to start.
Every figure below carries the configuration that produced it. Full tables of all 91 points, the reproduction script and checksums are on the raw page; the five machines are described on the stand page.
Setup#
Model ggml-org/Qwen3.8-27B-GGUF, revision 0669b98607d47046c7c2b3f801011d54a08cfccf. Build llama.cpp df750f76bb6126566621803b69ddaeb993be5b08 (b10871). Ladder flags -p 512 -n 128 -r 3 -t 8 -d <depth> -v.
A rung is prefill depth: 32k means 32768 tokens already sitting in the cache when the measurement starts, not "supports 32k context". Throughput is written as TPS, memory in GiB. The usable threshold is 10 TPS of token generation — below that the model technically runs and practically does not.
All five machines were rented, single card per host, consumer boards and desktop CPUs, PCIe x16 throughout. Nothing here was measured on a workstation with eight memory channels, because nobody watching this has one.
The four tiers#
| tier | card | best usable result | configuration |
|---|---|---|---|
| 12 GB | RTX 3060 | 3.3 TPS at 8k | stock Q4_K_M, 40 of 64 layers on GPU, -ub 256 |
| 16 GB | RTX 4070 Ti SUPER | 33.0 TPS at 32k | IQ4_XS, all layers, KV q8_0 both, -ub 256 |
| 24 GB | RTX 3090 / RTX 4090 | 34.3 / 40.0 TPS at 32k | stock Q4_K_M, all layers, KV f16, -ub 512 |
| 32 GB | RTX 5090 | 54.5 TPS at 96k | stock Q4_K_M, all layers, KV f16, -ub 512 |
Twelve gigabytes is out. Sixteen is the interesting tier. Twenty-four and thirty-two work out of the box.
The ladder, stock quant#
The file from the repository is 17.67 GiB, so the ladder only exists on cards that can hold it.
| depth | RTX 3090 | RTX 4090 | RTX 5090 |
|---|---|---|---|
| 8k | 1213.3 / 37.1 TPS · 18.5 GiB | 2828.8 / 42.9 TPS · 18.6 GiB | 3696.2 / 67.7 TPS · 18.8 GiB |
| 32k | 955.1 / 34.3 TPS · 20.0 GiB | 2343.0 / 40.0 TPS · 20.1 GiB | 2903.8 / 63.5 TPS · 20.2 GiB |
| 96k | fails, 505 MiB short (note 1) | fails, out of memory | 1536.0 / 54.5 TPS · 24.2 GiB |
| 256k | — | — | fails, 16416 MiB short |
Cells read prompt processing / token generation · peak VRAM.
Note 1. The 505 MiB figure comes from a separate manual run on the 3090 with -v, taken before the harness learned to record the missing buffer. On the 4090 the run reports out of memory without a magnitude, so none is quoted.
Three things fall out of this table.
The ceiling belongs to the memory, not the generation. The 3090 and the 4090 stop in exactly the same place, between 32k and 96k. Two generations apart, same wall.
Depth costs memory, not speed — until the memory runs out. From 8k to 32k the 3090 loses 8% of generation and gains 1.5 GiB. The 5090 goes all the way to 96k for 19% and 5.4 GiB. Right up to the wall the curve is nearly flat, and then there is no curve at all.
The gap between cards is uneven, and the unevenness is informative. Prompt processing on the 4090 is 2.3 to 2.45 times faster than on the 3090; generation is 16 to 17% faster. Prefill is compute-bound, where the newer card is far ahead. Generation is memory-bandwidth-bound, where the two are close. The numbers separate the two phases cleanly.
At 96k the 3090 was 505 MiB short — of the prompt compute buffer, not the KV cache. Half a gigabyte. The 4090 fails at the same rung with the same amount of memory, though its shortfall was never measured. Whether a smaller microbatch breaks that wall is untested here; the same lever did break the same kind of wall on 12 GB, so the hypothesis is strong, but it is a hypothesis.
Twelve gigabytes: the answer is no#
The RTX 3060 is the most common desktop card in the Steam survey. It has 11911 MiB available against a 17.67 GiB file. It does not load the model at all — oom_load on the first rung, before context is even allocated.
So the whole tier is about offloading, and offloading was measured properly, at -d 8192:
| layers on GPU | prompt processing | token generation | peak VRAM |
|---|---|---|---|
| 64 (all) | fails, 17.1 GiB short | — | — |
| 40 | fails, 513 MiB short | — | — |
| 32 | 273.1 TPS | 2.6 TPS | 9.8 GiB |
| 16 | 223.0 TPS | 1.8 TPS | 5.6 GiB |
| 0 | 185.5 TPS | 1.4 TPS | 1.8 GiB |
Abridged; the full sweep is on the raw page. The -ngl 0 row comes from a repeat run: in the first pass that rung failed for reasons unrelated to memory, and it was taken again on its own.
Note the 40-layer row: the weights fit, and the failure changes character to 513 MiB of compute buffer — practically the same half gigabyte that stopped the 3090 at 96k. One buffer, two very different cards.
Dropping -ub to 256 freed exactly that, put 40 layers on the card and raised generation from 2.6 to 3.3 TPS. Going further down to 128 and 64 changed generation not at all — it stays 3.3 — while prompt processing fell from 220.2 to 80.9 TPS. Lower the microbatch to the first value that fits and stop.
With cache quantisation, layer tuning and microbatch combined, the deepest the tier reached was 2.6 TPS at 32k with 36 layers and q4_0 cache. The threshold is 10. Every lever was pulled and the card stayed three times below it.
One more figure worth having: at -ngl 0, entirely on the CPU, the card still does 1.4 TPS. Half the model on the card buys you 2.6. Buying a 3060 for this model is worth less than a factor of two over not having a GPU at all.
Sixteen gigabytes: three decisions, none of them hardware#
The stock file does not load on 16376 MiB either. This is where the tier could have ended, and where most people stop.
Decision one: cut your own quant. Our Q4_K_M, quantised from the official BF16 with the same build, is 15.41 GiB against the repository's 17.67 GiB. Same quant name, different files, 2.26 GiB apart. Ours loads. "Q4_K_M fits in my card" is a claim about a specific file, not about a quant type.
Even then it does not fit whole: 149.62 MiB short. That number is the recurrent state of 48 of the model's 64 blocks. It is pinned to F32 and it does not shrink — sweeping -ub from 512 down to 64 moved the shortfall by nothing at all. Dropping two layers to the CPU gets you 20.8 TPS.
Decision two: IQ4_XS instead of Q4_K_M.
| quant | file size | placement | generation at 8k | same top token |
|---|---|---|---|---|
| stock Q4_K_M | 17.67 GiB | does not load | — | — |
| own Q4_K_M | 15.41 GiB | 62 of 64 layers | 20.8 TPS | 93.33% |
| IQ4_XS | 14.15 GiB | all 64 layers | 36.6 TPS | 93.15% |
1.26 GiB of file size buys 76% more throughput, because nothing has to cross the bus. The quality difference is 0.18 percentage points — noise. Most people pick Q4_K_M out of habit and pay for it with half their speed.
The layer curve explains why: 62 layers give 20.8 TPS, 60 give 15.5, 56 give 10.8. Every two layers cost roughly a quarter of generation. Offloading is not a gentle trade-off, it is a cliff.
Decision three: compress the cache — and it is the values that matter. At 32k with IQ4_XS and both halves at f16, the KV cache itself does not fit: 2080 MiB short, before any compute buffer is even allocated. Compress either half to q8_0 and the cache fits at 1592.50 MiB — the same size whether it is the keys or the values — and the failure moves to the prompt compute buffer. Keys in q8_0, values at f16: 1758 MiB short. Keys in q4_0: the same 1758 MiB. Values in q8_0, keys left at f16: 505 MiB short. Same cache size, three and a half times less missing, because the compute buffer is sized by what it has to read back. With both halves in q8_0 and -ub 256 it runs.
And one trap that looks entirely harmless:
| cache | prompt processing | token generation | peak VRAM |
|---|---|---|---|
-ctk q4_0 -ctv q8_0 |
4.2 TPS | 3.3 TPS | 15.3 GiB |
-ctk q8_0 -ctv q8_0 |
1325.2 TPS | 33.0 TPS | 15.2 GiB |
The mismatched-pair row was observed in a manual run that was stopped before the harness wrote its output files, so it has no artifact on the raw page and is not part of the recorded set. It is repeated here because the effect is unmissable; the exact figure deserves a re-run.
Memory is the same either way. There is no on-card kernel for a mismatched pair, so the work falls to the CPU and prompt processing collapses by a factor of 316. Mixing key and value cache types is the single most expensive setting in this whole article.
Put the three together — IQ4_XS, KV q8_0 for both, -ub 256 — and the tier lands at 33.0 TPS at 32k with all 64 layers on the card, 15.2 GiB peak. From "does not start" to comfortable, without touching the hardware.
Two walls that use the same words#
Every failure above reports out of memory. They are not the same failure.
On the 3060 and on the 3090 the missing piece was the prompt compute buffer: 513 MiB and 505 MiB. A smaller microbatch shrinks that buffer directly, and on 12 GB it broke the wall.
On the 5090 at 256k — the context the vendor advertises — the missing piece is 16416 MiB of KV cache. Dropping -ub to 256 and then to 128 changed the shortfall by exactly zero. Keys in q8_0 brought it to 13406 MiB; keys in q4_0 brought it to the same 13406. Values in q8_0 finally made it fit.
And then it did not work. The card sat at 0% while eight host cores ran at 810%, and an hour was not enough to finish. The vendor's maximum context is not practically reachable on 32 GB.
Which is the same answer the ladder gives at the other end. On the 3060 the model starts and produces 2.6 TPS. Fitting is not running.
The price of compression#
Measured once, on the 5090, because Q8_0 had to fit somewhere. Quality belongs to the quant, not to the card, so this table applies to every tier above.
Reference Qwen3.8-27B-BF16.gguf, perplexity 6.782793 ± 0.073775. Text datasets/ggml-org/ci revision 927b3642933080f1b0e811e2f916e14c292992f9, member wikitext-2-raw/wiki.test.raw, 200 windows at -c 512, KV f16 throughout. Metric is KL divergence; the column readers care about is how often the compressed model picks the same best token as the uncompressed one.
| quant | file size | KL divergence | same top token |
|---|---|---|---|
| Q8_0 | 26.63 GiB | 0.000797 | 98.71% |
| Q6_K | 20.57 GiB | 0.002654 | 97.70% |
| Q5_K_M | 17.91 GiB | 0.008207 | 96.01% |
| Q4_K_M | 15.41 GiB | 0.023021 | 93.33% |
| IQ4_XS | 14.15 GiB | 0.023808 | 93.15% |
| Q4_0 | 14.41 GiB | 0.036346 | 91.61% |
| Q3_K_M | 12.39 GiB | 0.071103 | 88.55% |
| Q2_K | 9.98 GiB | 0.294271 | 77.23% |
IQ4_XS and Q4_0 weigh almost the same and differ by 1.54 points: the name matters more than the size. Degradation is gradual down to Q3_K_M and then breaks — Q2_K changes roughly every fourth token, with four times the divergence of the step above it.
Which closes the 12 GB question from the other side. Q3_K_M needs 12.39 GiB against 11911 MiB available, so it misses by 776 MiB. The only quant that fits a 3060 whole is Q2_K, at 77%. Either a badly damaged model entirely on the card, or a decent one at 2.6 TPS.
An honest caveat: these quants were cut without an importance matrix, which none of the eight types requires but which some public repositories use. Ours are not byte-identical to those. It does not affect speed or memory, and it is worth saying out loud.
Negative results#
They are results, so they are here.
Selective offload by tensor mask lost to plain layer offload. Masks are the advanced tool: name any tensor, send it to the CPU. Sixteen blocks without their recurrent part gave 9.0 TPS; all FFN down projections gave 8.4. Two whole layers via -ngl 62 gave 20.8. The reason is visible in the memory column: the masks freed 2.5 to 3 GiB where 150 MiB was missing. Twenty times more than needed, paid for with half the speed. The tool works, but only if the mask is cut to the exact shortfall.
Flash attention is already at its optimum. auto and on are identical inside noise — 1655.7 and 1652.3 TPS of prefill. Turning it off costs 15% of prefill and 200 MiB. One of the settings people are told to tune has nothing to tune.
Threads saturate at four, and sixteen hurt. On a configuration with two layers on the CPU: 14.6 TPS at one thread, 19.2 at two, 20.7 at four, 20.7 at eight, 18.8 at sixteen. The host has eight physical cores; sixteen threads is hyperthreading, and it costs 9%. Prompt processing ignores the setting entirely — 1293 to 1306 TPS across the whole range, because it runs on the card.
Cache quantisation gives less than you expect here. Fourfold compression freed 467 MiB on the 3060, not four times the cache. This model is hybrid: 48 of its 64 blocks carry recurrent state pinned to F32, and -ctk / -ctv only reach the attention part.
One procedural lesson#
The BF16 source, 50.10 GiB, downloaded on a slow host, reached 100% and matched the repository byte count exactly. Its sha256 did not match. That surfaced an hour later, during quantisation, as an infinity value in the eighth tensor out of 851 — llama-quantize validates its output and refused. On a fast host the same file arrived correctly, so the fault was the link, not the source.
Size does not prove integrity. Check the sha256 against the x-linked-etag HuggingFace returns. The cost of skipping it here was an hour of download and a failed quantisation run; the cost in a worse case is a broken quant producing plausible numbers.
What was not measured#
Stated so that nobody reads a ceiling into a gap.
- A recorded run of the mismatched key/value cache pair. The effect was observed manually; the numbers above are not backed by an artifact.
- Whether
-ub 256breaks the 96k wall on 24 GB. - The ceiling on 16 GB. 32k is confirmed working; 96k was never attempted there.
-nkvo, tensor masks and thread sweeps on 24 and 32 GB — taken on 16 GB only.-fitt, llama.cpp's own solver, against our measured walls.- Depth
-d 0and the vendor maximum on 24 GB. - Memory channel count on any host. Two-channel operation is inferred from measured bandwidth: 34.6, 34.5, 43.3, 31.3 and 68.1 GB/s across the 3060, 3090, 4070 Ti SUPER, 4090 and 5090.
Reproducing this#
The raw page carries all 91 points, reproduce.sh with every command from cloning llama.cpp to the last sweep, results.json with the machine-readable set, and sha256 for the source BF16 and all eight quants. Weights are not hosted here — the repository, the revision and the quantise commands reproduce the same files byte for byte.