Raw measurement log
vLLM vs llama.cpp on a rented RTX 5090
- Date
- 2026-09-27
- Runs
- 4
raw.md
# vLLM vs llama.cpp on a rented RTX 5090
date: 2026-09-27
machine: Vast.ai rented instance 52918859 (machine 143669, Puerto Rico, verified host, reliability 99.56%): NVIDIA GeForce RTX 5090 32 GB (32607 MiB), driver 595.71.05, CUDA 13.2; AMD Ryzen 9 9900X 12-core/24-thread, 91 GB RAM, X870 Taichi Creator, PCIe 5.0 x8 (27.5 GB/s), Samsung SSD 9100 PRO, 2150/2024 Mbps; disk 60 GB; price $0.623/hr (rent_price.jpg); total bill $0.80 for ~1.3 h incl. ~35 min setup (bill.jpg)
task: "Serve Qwen3.8-27B at 4-bit on vLLM and on llama.cpp on the same rented RTX 5090. Per engine: billed seconds from launch to ready and to first token, stable tokens/s at one request; then raise concurrent requests step by step and record requests/hour, total output tokens/s and median time-to-first-token; mark the load where latency becomes unacceptable (our threshold: median TTFT > 2 s)."
repo: none — no code task; benchmark client `vllm bench serve --backend openai-chat`, dataset random, input 1024 / output 256 tokens, seed 0, same prompt set and same client against both servers' OpenAI-compatible /v1/chat/completions
files_dir: manual/
stand_files: rent_price.jpg (Vast search, row m:143669 $0.623/hr), stand.jpg (instance card: host, GPU, IDs), bill.jpg (Billing, $0.80 total, spending rate $0.00 after destroy)
model: Qwen3.8-27B (Alibaba, Apache 2.0, dense 27B, hybrid Gated DeltaNet). vLLM checkpoint cyankiwi/Qwen3.8-27B-AWQ-INT4 (community AWQ 4-bit, compressed-tensors, 21.0 GB; no official Qwen AWQ exists; 951k downloads). llama.cpp checkpoint ggml-org/Qwen3.8-27B-GGUF Qwen3.8-27B-Q4_K_M.gguf (19.0 GB)
versions: vLLM 0.30.0 (torch 2.13.0+cu132); llama.cpp build b11209, commit 187664b53, prebuilt ubuntu-cuda-12.8-x64 binary
conditions: 4-bit on both (formats differ: AWQ vs Q4_K_M); text-only on both (vLLM vision encoder disabled via --limit-mm-per-prompt, llama.cpp without mmproj); thinking on by default on both, capped identically by max_tokens 256; instance time 10:55–11:55 Europe/Paris
=== r01 baseline — llama.cpp, single stream
beat_id: baseline
status: done
checks: same prompt set=yes; same client=yes; same quant bits=yes; text only=yes
setup: llama-server b11209, -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -c 32768 -np 1 --port 8001
launch: T0=$(date +%s); ./llama-server ... & until curl -s localhost:8001/health | grep -q '"ok"'; do sleep 1; done; echo $(( $(date +%s)-T0 ))
command: vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee llamacpp_c1.log
download_s: 95 (hf download ggml-org/Qwen3.8-27B-GGUF, 19.0 GB at 218 MB/s)
load_s: 3 (weights still in page cache right after download; PCIe 5.0)
ttft_short_prompt_s: 0.109 (curl -w '%{time_starttransfer}', prompt "Explain what a mutex is.", max_tokens 64, stream)
bench: vllm bench serve ... --num-prompts 20 --max-concurrency 1
tg_single_tps: 60.76 (output token throughput), peak 70.00; median TTFT on 1024-token prompt 514.92 ms; median TPOT 14.53 ms; duration 84.26 s
files: llamacpp_c1.log, llamacpp_c1.jpg, llamacpp_c1_server.log
notes: vast base image had no vLLM; pip install failed building xformers, installed via `uv pip install vllm --torch-backend=auto`; llama.cpp taken as prebuilt release binary, no compile
=== r02 baseline — vLLM, single stream
beat_id: baseline
status: done
checks: same prompt set=yes; same client=yes; same quant bits=yes; text only=yes
setup: vllm serve /workspace/models/awq --served-model-name q --max-model-len 32768 --limit-mm-per-prompt '{"image":0,"video":0}' --max-num-seqs 64 --attention-backend FLASH_ATTN --port 8000; env VLLM_USE_FLASHINFER_SAMPLER=0
launch: T0=$(date +%s); vllm serve ... & for i in $(seq 1 900); do curl -s localhost:8000/health && break; sleep 1; done; echo $(( $(date +%s)-T0 ))
command: vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee vllm_c1.log
download_s: 137 (hf download cyankiwi/Qwen3.8-27B-AWQ-INT4, 21.0 GB at 281 MB/s)
load_s: 61 (includes CUDA-graph capture)
ttft_short_prompt_s: 0.043 (same curl, same prompt)
bench: vllm bench serve ... --num-prompts 20 --max-concurrency 1
tg_single_tps: 72.52 (output token throughput), peak 79.00; median TTFT on 1024-token prompt 278.09 ms; median TPOT 12.75 ms; duration 70.61 s
files: vllm_c1.log, vllm_c1.jpg, vllm_server.log
notes: three failed starts before this one, not counted in load_s: (1) OOM at profiling — multimodal model reserves memory for 255 images per prompt, fixed with --limit-mm-per-prompt; (2) FlashInfer wheel does not recognise Blackwell SM 12.0, fixed with --attention-backend FLASH_ATTN and sampler off; (3) torchaudio CUDA mismatch after reinstall, torchaudio uninstalled
=== r03 subjects — llama.cpp, concurrency ladder
beat_id: subjects
status: done
checks: same prompt set=yes; same client=yes; same quant bits=yes; text only=yes
setup: llama-server -ngl 99, -c 65536 -np 8 for steps 4 and 8; -c 131072 -np 16 for step 16 (context is split across slots)
command: for C in 4 8 16; do vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee llamacpp_c$C.log; done
ladder: c4 → 130.86 tok/s, 0.51 req/s (1836/h), median TTFT 2128.85 ms; c8 → 195.86 tok/s, 0.77 req/s (2772/h), median TTFT 3443.10 ms; c16 → 235.86 tok/s, 0.92 req/s (3312/h), median TTFT 4340.90 ms
break_at: 4 (median TTFT already 2.13 s at 4 concurrent)
ceiling: 235.86 tok/s at -np 16 (last tested point, not a hard limit; server started, no OOM)
files: llamacpp_c4.log, llamacpp_c8.log, llamacpp_c16.log, llamacpp_np8_server.log, llamacpp_np16_server.log
notes: before step 16 the server was restarted with `-c 131072 -np 16`
=== r04 subjects — vLLM, concurrency ladder
beat_id: subjects
status: done
checks: same prompt set=yes; same client=yes; same quant bits=yes; text only=yes
setup: same vllm serve as r02
command: for C in 4 8 16 32 64; do vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee vllm_c$C.log; done
ladder: c4 → 221.35 tok/s, 0.86 req/s (3096/h), median TTFT 935.46 ms; c8 → 348.60 tok/s, 1.36 req/s (4896/h), median TTFT 1204.36 ms; c16 → 464.51 tok/s, 1.81 req/s (6516/h), median TTFT 1526.12 ms; c32 → 580.02 tok/s, 2.27 req/s (8172/h), median TTFT 2189.60 ms; c64 → 580.13 tok/s, 2.27 req/s (8172/h), median TTFT 14975.33 ms
break_at: 32 (median TTFT 2.19 s)
ceiling: 580 tok/s — throughput flat between 32 and 64 concurrent, extra requests only queue
files: vllm_c4.log, vllm_c8.log, vllm_c16.log, vllm_c32.log, vllm_c64.log, vllm_server.log
notes: none
=== coverage
baseline: r01 done, r02 done
subjects: r03 done, r04 done