vLLM vs SGLang · Benchmarks
Unified A/B benchmarkSM90 H20-3e / SM10.3 B300 / SM120 RTX PRO 6000

vLLM vs SGLang Benchmark Dashboard

Unified A/B over all merged benchmark records; fair win = output tok/s better by >0.5%

Headline verdict

Across 149 fair head-to-head runs, vLLM came out ahead in 119 of them — winning 82% of the decided match-ups. SGLang took 27, with 3 ties.

A “fair win” means one engine’s output throughput (tokens/second) beat the other’s by more than 0.5% on the same host, model and workload. A further 4 runs were set aside as no-parity (one engine hit a bug or config skew, so the match-up wasn’t apples-to-apples).

vLLM · 11927 · SGLang
80% of all runs3 ties18% of all runs
Fair wins · vLLM
119
Fair wins · SGLang
27
Ties
3
Models
9
Host × model pairs
17
Measured runs
369
Fair pairings
149
No-parity runs
4
Raw benchmark rows
369
Root-caused bugs
20

Who wins, by model

Each bar is one model's paired runs across every host. Longer cyan = vLLM won more of that model's workloads.

By host
rtx6000-159 · 1 · 3
h20-035 · 18 · 1
h20-115 · 1 · 2
b3002 · 7
h20-0+h20-18 · 0 · 1
By workload type
sharegpt43 · 7 · 1
short35 · 13 · 1 · 2
long41 · 7 · 1 · 2

Throughput, run by run

Every dot is one paired run — vertical axis is vLLM tokens/sec, horizontal is SGLang. Above the dashed line, vLLM was faster; below it, SGLang was. Hover any dot for the exact run.