vLLM vs SGLang Benchmark Dashboard
Unified A/B over all merged benchmark records; fair win = output tok/s better by >0.5%
Across 149 fair head-to-head runs, vLLM came out ahead in 119 of them — winning 82% of the decided match-ups. SGLang took 27, with 3 ties.
A “fair win” means one engine’s output throughput (tokens/second) beat the other’s by more than 0.5% on the same host, model and workload. A further 4 runs were set aside as no-parity (one engine hit a bug or config skew, so the match-up wasn’t apples-to-apples).
Who wins, by model
Each bar is one model's paired runs across every host. Longer cyan = vLLM won more of that model's workloads.
Throughput, run by run
Every dot is one paired run — vertical axis is vLLM tokens/sec, horizontal is SGLang. Above the dashed line, vLLM was faster; below it, SGLang was. Hover any dot for the exact run.