Tag: roofline model
-

Benchmark Your Own LLM Serving Stack: Two Measurements, One Afternoon
Run an LLM serving benchmark on your own GPU in a single afternoon: measure the hardware roofline, sweep concurrency for TTFT and…
-

The Roofline Model: Why LLM Decode Is Memory-Bound
The roofline model in one chart: why LLM decode at batch 1 uses under 1 percent of your GPU, where batching stops…