Tag: inference optimization
-

The Roofline Model: Why LLM Decode Is Memory-Bound
The roofline model in one chart: why LLM decode at batch 1 uses under 1 percent of your GPU, where batching stops…

The roofline model in one chart: why LLM decode at batch 1 uses under 1 percent of your GPU, where batching stops…