Field measurements
Bench
Numbers taken on real hardware, not quoted from model cards. Every run publishes the article, the raw artifacts it was built from, and the passport of the stand it was taken on.
vLLM vs llama.cpp on a rented RTX 5090
Serve Qwen3.8-27B at 4-bit on vLLM and on llama.cpp on the same rented RTX 5090. Per engine: billed seconds from launch to ready and to first token, stable tokens/s at one request; then raise concurrent requests step by step and record requests/hour, total output tokens/s and median time-to-first-token; mark the load where latency becomes unacceptable (our threshold: median TTFT > 2 s).
Qwen3.8-27B
ggml-org/Qwen3.8-27B-GGUF
Prefill ladder and generation throughput of Qwen3.8-27B under llama.cpp, measured per GPU and per quant on the GREZA stand.