Field measurement

vLLM vs llama.cpp on a rented RTX 5090

Date
2026-09-27
Runs
4

Task

Serve Qwen3.8-27B at 4-bit on vLLM and on llama.cpp on the same rented RTX 5090. Per engine: billed seconds from launch to ready and to first token, stable tokens/s at one request; then raise concurrent requests step by step and record requests/hour, total output tokens/s and median time-to-first-token; mark the load where latency becomes unacceptable (our threshold: median TTFT > 2 s).

Stand

machineVast.ai rented instance 52918859 (machine 143669, Puerto Rico, verified host, reliability 99.56%): NVIDIA GeForce RTX 5090 32 GB (32607 MiB), driver 595.71.05, CUDA 13.2; AMD Ryzen 9 9900X 12-core/24-thread, 91 GB RAM, X870 Taichi Creator, PCIe 5.0 x8 (27.5 GB/s), Samsung SSD 9100 PRO, 2150/2024 Mbps; disk 60 GB; price $0.623/hr (rent_price.jpg); total bill $0.80 for ~1.3 h incl. ~35 min setup (bill.jpg)
reponone — no code task; benchmark client `vllm bench serve --backend openai-chat`, dataset random, input 1024 / output 256 tokens, seed 0, same prompt set and same client against both servers' OpenAI-compatible /v1/chat/completions
modelQwen3.8-27B (Alibaba, Apache 2.0, dense 27B, hybrid Gated DeltaNet). vLLM checkpoint cyankiwi/Qwen3.8-27B-AWQ-INT4 (community AWQ 4-bit, compressed-tensors, 21.0 GB; no official Qwen AWQ exists; 951k downloads). llama.cpp checkpoint ggml-org/Qwen3.8-27B-GGUF Qwen3.8-27B-Q4_K_M.gguf (19.0 GB)
versionsvLLM 0.30.0 (torch 2.13.0+cu132); llama.cpp build b11209, commit 187664b53, prebuilt ubuntu-cuda-12.8-x64 binary
conditions4-bit on both (formats differ: AWQ vs Q4_K_M); text-only on both (vLLM vision encoder disabled via --limit-mm-per-prompt, llama.cpp without mmproj); thinking on by default on both, capped identically by max_tokens 256; instance time 10:55–11:55 Europe/Paris
Vast search, row m:143669 $0.623/hr
rent_price.jpg — Vast search, row m:143669 $0.623/hr
instance card: host, GPU, IDs
stand.jpg — instance card: host, GPU, IDs
Billing, $0.80 total, spending rate $0.00 after destroy
bill.jpg — Billing, $0.80 total, spending rate $0.00 after destroy

r01 — baseline — llama.cpp, single stream

beat baseline · status done

Checks

same prompt setyes
same clientyes
same quant bitsyes
text onlyyes

setup

llama-server b11209, -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -c 32768 -np 1 --port 8001

launch

T0=$(date +%s); ./llama-server ... & until curl -s localhost:8001/health | grep -q '"ok"'; do sleep 1; done; echo $(( $(date +%s)-T0 ))

command

vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee llamacpp_c1.log

Results

download_s95 (hf download ggml-org/Qwen3.8-27B-GGUF, 19.0 GB at 218 MB/s)
load_s3 (weights still in page cache right after download; PCIe 5.0)
ttft_short_prompt_s0.109 (curl -w '%{time_starttransfer}', prompt "Explain what a mutex is.", max_tokens 64, stream)
benchvllm bench serve ... --num-prompts 20 --max-concurrency 1
tg_single_tps60.76 (output token throughput), peak 70.00; median TTFT on 1024-token prompt 514.92 ms; median TPOT 14.53 ms; duration 84.26 s

Files

llamacpp_c1.jpg
llamacpp_c1.jpg

Notes

vast base image had no vLLM; pip install failed building xformers, installed via `uv pip install vllm --torch-backend=auto`; llama.cpp taken as prebuilt release binary, no compile

r02 — baseline — vLLM, single stream

beat baseline · status done

Checks

same prompt setyes
same clientyes
same quant bitsyes
text onlyyes

setup

vllm serve /workspace/models/awq --served-model-name q --max-model-len 32768 --limit-mm-per-prompt '{"image":0,"video":0}' --max-num-seqs 64 --attention-backend FLASH_ATTN --port 8000; env VLLM_USE_FLASHINFER_SAMPLER=0

launch

T0=$(date +%s); vllm serve ... & for i in $(seq 1 900); do curl -s localhost:8000/health && break; sleep 1; done; echo $(( $(date +%s)-T0 ))

command

vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee vllm_c1.log

Results

download_s137 (hf download cyankiwi/Qwen3.8-27B-AWQ-INT4, 21.0 GB at 281 MB/s)
load_s61 (includes CUDA-graph capture)
ttft_short_prompt_s0.043 (same curl, same prompt)
benchvllm bench serve ... --num-prompts 20 --max-concurrency 1
tg_single_tps72.52 (output token throughput), peak 79.00; median TTFT on 1024-token prompt 278.09 ms; median TPOT 12.75 ms; duration 70.61 s

Files

vllm_c1.jpg
vllm_c1.jpg

Notes

three failed starts before this one, not counted in load_s: (1) OOM at profiling — multimodal model reserves memory for 255 images per prompt, fixed with --limit-mm-per-prompt; (2) FlashInfer wheel does not recognise Blackwell SM 12.0, fixed with --attention-backend FLASH_ATTN and sampler off; (3) torchaudio CUDA mismatch after reinstall, torchaudio uninstalled

r03 — subjects — llama.cpp, concurrency ladder

beat subjects · status done

Checks

same prompt setyes
same clientyes
same quant bitsyes
text onlyyes

setup

llama-server -ngl 99, -c 65536 -np 8 for steps 4 and 8; -c 131072 -np 16 for step 16 (context is split across slots)

command

for C in 4 8 16; do vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee llamacpp_c$C.log; done

ladder

Step tok/s req/s median TTFT, ms
c4 130.86 0.51 (1836/h) 2128.85
c8 195.86 0.77 (2772/h) 3443.10
c16 235.86 0.92 (3312/h) 4340.90
tok/s
0 62.5 125 187.5 250 c4 c8 c16 break_at c4: 130.86 130.86 c8: 195.86 c16: 235.86 235.86
median TTFT, ms
0 1250 2500 3750 5000 c4 c8 c16 break_at c4: 2128.85 2128.85 c8: 3443.10 c16: 4340.90 4340.90
break_at4 (median TTFT already 2.13 s at 4 concurrent)
ceiling235.86 tok/s at -np 16 (last tested point, not a hard limit; server started, no OOM)

Files

Notes

before step 16 the server was restarted with `-c 131072 -np 16`

r04 — subjects — vLLM, concurrency ladder

beat subjects · status done

Checks

same prompt setyes
same clientyes
same quant bitsyes
text onlyyes

setup

same vllm serve as r02

command

for C in 4 8 16 32 64; do vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee vllm_c$C.log; done

ladder

Step tok/s req/s median TTFT, ms
c4 221.35 0.86 (3096/h) 935.46
c8 348.60 1.36 (4896/h) 1204.36
c16 464.51 1.81 (6516/h) 1526.12
c32 580.02 2.27 (8172/h) 2189.60
c64 580.13 2.27 (8172/h) 14975.33
tok/s
0 150 300 450 600 c4 c8 c16 c32 c64 break_at c4: 221.35 c8: 348.60 c16: 464.51 c32: 580.02 580.02 c64: 580.13 580.13
median TTFT, ms
0 5000 10000 15000 20000 c4 c8 c16 c32 c64 break_at c4: 935.46 c8: 1204.36 c16: 1526.12 c32: 2189.60 2189.60 c64: 14975.33 14975.33
break_at32 (median TTFT 2.19 s)
ceiling580 tok/s — throughput flat between 32 and 64 concurrent, extra requests only queue

Files

Notes

none

Coverage

baseliner01 done, r02 done
subjectsr03 done, r04 done