Field measurement
vLLM vs llama.cpp on a rented RTX 5090
- Date
- 2026-09-27
- Runs
- 4
Task
Serve Qwen3.8-27B at 4-bit on vLLM and on llama.cpp on the same rented RTX 5090. Per engine: billed seconds from launch to ready and to first token, stable tokens/s at one request; then raise concurrent requests step by step and record requests/hour, total output tokens/s and median time-to-first-token; mark the load where latency becomes unacceptable (our threshold: median TTFT > 2 s).
Stand
| machine | Vast.ai rented instance 52918859 (machine 143669, Puerto Rico, verified host, reliability 99.56%): NVIDIA GeForce RTX 5090 32 GB (32607 MiB), driver 595.71.05, CUDA 13.2; AMD Ryzen 9 9900X 12-core/24-thread, 91 GB RAM, X870 Taichi Creator, PCIe 5.0 x8 (27.5 GB/s), Samsung SSD 9100 PRO, 2150/2024 Mbps; disk 60 GB; price $0.623/hr (rent_price.jpg); total bill $0.80 for ~1.3 h incl. ~35 min setup (bill.jpg) |
|---|---|
| repo | none — no code task; benchmark client `vllm bench serve --backend openai-chat`, dataset random, input 1024 / output 256 tokens, seed 0, same prompt set and same client against both servers' OpenAI-compatible /v1/chat/completions |
| model | Qwen3.8-27B (Alibaba, Apache 2.0, dense 27B, hybrid Gated DeltaNet). vLLM checkpoint cyankiwi/Qwen3.8-27B-AWQ-INT4 (community AWQ 4-bit, compressed-tensors, 21.0 GB; no official Qwen AWQ exists; 951k downloads). llama.cpp checkpoint ggml-org/Qwen3.8-27B-GGUF Qwen3.8-27B-Q4_K_M.gguf (19.0 GB) |
| versions | vLLM 0.30.0 (torch 2.13.0+cu132); llama.cpp build b11209, commit 187664b53, prebuilt ubuntu-cuda-12.8-x64 binary |
| conditions | 4-bit on both (formats differ: AWQ vs Q4_K_M); text-only on both (vLLM vision encoder disabled via --limit-mm-per-prompt, llama.cpp without mmproj); thinking on by default on both, capped identically by max_tokens 256; instance time 10:55–11:55 Europe/Paris |
r01 — baseline — llama.cpp, single stream
beat baseline · status done
Checks
| same prompt set | yes |
|---|---|
| same client | yes |
| same quant bits | yes |
| text only | yes |
setup
llama-server b11209, -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -c 32768 -np 1 --port 8001
launch
T0=$(date +%s); ./llama-server ... & until curl -s localhost:8001/health | grep -q '"ok"'; do sleep 1; done; echo $(( $(date +%s)-T0 ))
command
vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee llamacpp_c1.log
Results
| download_s | 95 (hf download ggml-org/Qwen3.8-27B-GGUF, 19.0 GB at 218 MB/s) |
|---|---|
| load_s | 3 (weights still in page cache right after download; PCIe 5.0) |
| ttft_short_prompt_s | 0.109 (curl -w '%{time_starttransfer}', prompt "Explain what a mutex is.", max_tokens 64, stream) |
| bench | vllm bench serve ... --num-prompts 20 --max-concurrency 1 |
| tg_single_tps | 60.76 (output token throughput), peak 70.00; median TTFT on 1024-token prompt 514.92 ms; median TPOT 14.53 ms; duration 84.26 s |
Files
Notes
vast base image had no vLLM; pip install failed building xformers, installed via `uv pip install vllm --torch-backend=auto`; llama.cpp taken as prebuilt release binary, no compile
r02 — baseline — vLLM, single stream
beat baseline · status done
Checks
| same prompt set | yes |
|---|---|
| same client | yes |
| same quant bits | yes |
| text only | yes |
setup
vllm serve /workspace/models/awq --served-model-name q --max-model-len 32768 --limit-mm-per-prompt '{"image":0,"video":0}' --max-num-seqs 64 --attention-backend FLASH_ATTN --port 8000; env VLLM_USE_FLASHINFER_SAMPLER=0
launch
T0=$(date +%s); vllm serve ... & for i in $(seq 1 900); do curl -s localhost:8000/health && break; sleep 1; done; echo $(( $(date +%s)-T0 ))
command
vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 20 --max-concurrency 1 2>&1 | tee vllm_c1.log
Results
| download_s | 137 (hf download cyankiwi/Qwen3.8-27B-AWQ-INT4, 21.0 GB at 281 MB/s) |
|---|---|
| load_s | 61 (includes CUDA-graph capture) |
| ttft_short_prompt_s | 0.043 (same curl, same prompt) |
| bench | vllm bench serve ... --num-prompts 20 --max-concurrency 1 |
| tg_single_tps | 72.52 (output token throughput), peak 79.00; median TTFT on 1024-token prompt 278.09 ms; median TPOT 12.75 ms; duration 70.61 s |
Files
Notes
three failed starts before this one, not counted in load_s: (1) OOM at profiling — multimodal model reserves memory for 255 images per prompt, fixed with --limit-mm-per-prompt; (2) FlashInfer wheel does not recognise Blackwell SM 12.0, fixed with --attention-backend FLASH_ATTN and sampler off; (3) torchaudio CUDA mismatch after reinstall, torchaudio uninstalled
r03 — subjects — llama.cpp, concurrency ladder
beat subjects · status done
Checks
| same prompt set | yes |
|---|---|
| same client | yes |
| same quant bits | yes |
| text only | yes |
setup
llama-server -ngl 99, -c 65536 -np 8 for steps 4 and 8; -c 131072 -np 16 for step 16 (context is split across slots)
command
for C in 4 8 16; do vllm bench serve --backend openai-chat --base-url http://localhost:8001 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee llamacpp_c$C.log; done
ladder
| Step | tok/s | req/s | median TTFT, ms |
|---|---|---|---|
| c4 | 130.86 | 0.51 (1836/h) | 2128.85 |
| c8 | 195.86 | 0.77 (2772/h) | 3443.10 |
| c16 | 235.86 | 0.92 (3312/h) | 4340.90 |
| break_at | 4 (median TTFT already 2.13 s at 4 concurrent) |
|---|---|
| ceiling | 235.86 tok/s at -np 16 (last tested point, not a hard limit; server started, no OOM) |
Files
- llamacpp_c4.log 6.1 KiB
- llamacpp_c8.log 6.8 KiB
- llamacpp_c16.log 8.2 KiB
- llamacpp_np8_server.log 198.5 KiB
- llamacpp_np16_server.log 114.2 KiB
Notes
before step 16 the server was restarted with `-c 131072 -np 16`
r04 — subjects — vLLM, concurrency ladder
beat subjects · status done
Checks
| same prompt set | yes |
|---|---|
| same client | yes |
| same quant bits | yes |
| text only | yes |
setup
same vllm serve as r02
command
for C in 4 8 16 32 64; do vllm bench serve --backend openai-chat --base-url http://localhost:8000 --endpoint /v1/chat/completions --model q --tokenizer /workspace/models/awq --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 100 --max-concurrency $C 2>&1 | tee vllm_c$C.log; done
ladder
| Step | tok/s | req/s | median TTFT, ms |
|---|---|---|---|
| c4 | 221.35 | 0.86 (3096/h) | 935.46 |
| c8 | 348.60 | 1.36 (4896/h) | 1204.36 |
| c16 | 464.51 | 1.81 (6516/h) | 1526.12 |
| c32 | 580.02 | 2.27 (8172/h) | 2189.60 |
| c64 | 580.13 | 2.27 (8172/h) | 14975.33 |
| break_at | 32 (median TTFT 2.19 s) |
|---|---|
| ceiling | 580 tok/s — throughput flat between 32 and 64 concurrent, extra requests only queue |
Files
- vllm_c4.log 6.6 KiB
- vllm_c8.log 8.1 KiB
- vllm_c16.log 7.7 KiB
- vllm_c32.log 7.6 KiB
- vllm_c64.log 7.6 KiB
- vllm_server.log 82.0 KiB
Notes
none
Coverage
| baseline | r01 done, r02 done |
|---|---|
| subjects | r03 done, r04 done |