Systems for Machine Learning · Final Project · UT Austin · Spring 2026
Energy-Efficient Multi-Modal LLM Serving
A Sub-phase Characterization Study
Aashrith Attelli  ·  Amaury Jayr  ·  Nathan Lemma  ·  Mahin Naveen
University of Texas at Austin
OVERVIEW

When you ask a multimodal AI to describe an image, the request doesn't hit the GPU as one uniform job. It moves through distinct phases — looking at the image, reading your prompt, then writing the answer one word at a time — and each phase stresses the hardware completely differently. Yet today's serving systems hand the entire GPU to every phase equally.

We profiled InternVL3-8B phase by phase on an NVIDIA A100, then swept how much of the GPU each phase actually gets — from 2 SMs up to all 108 — across 1,080 configurations. The headline: the phase that takes the most time uses the GPU the least, and roughly two-thirds of the chip often sits powered but idle.

KEY FINDINGS
89% time, 14% GPU
Decode — generating the answer token by token — eats 88.7% of every request's wall-clock time, yet sustains only 14% SM and 23% memory throughput. The slowest phase is also the one that wastes the most hardware.
Plateaus at SM = 36
Output-heavy throughput flattens once a third of the GPU is active. The remaining 67% of streaming multiprocessors add under 9% more throughput. Input-heavy work, dominated by the vision encoder, keeps scaling all the way to SM = 108.
230 W < 250 W TDP
Power saturates at ~230 W at that same SM = 36 elbow — below the A100's 250 W limit. Because the GPU never reaches TDP, masking a single job's stranded SMs can't recover the energy. Batching can: 13× more efficient for output-heavy work.
THE FOUR PHASES

One multimodal request, four phases. Pick any to see what it does and how hard it pushes the GPU.

Input pipeline · 3 phases
Generation · 1 phase
WHERE THE TIME GOES

Decode dominates the clock while barely touching the silicon — the gap this study is about.

Wall-time per request InternVL3-8B · BS 1 · 1344² image · 64 tokens
Per-phase GPU utilization % of peak sustained

Vision Encoder, MLP Connector and Prefill all keep the compute units busy. Decode is the outlier — low on both axes, for the longest time. Toggle the metric to compare.

How much headroom is in decode? Batching reveals it.

At batch 1, decode spends 23.2 ms per output token. Pack 16 requests together and that drops to 1.6 ms/token — a 14× speedup from the same hardware. The compute was always there; a single serial request just couldn't use it.

BS 1: 23.2 ms/tok BS 4: 6.5 ms/tok BS 8: 3.2 ms/tok BS 16: 1.6 ms/tok
Decode time per output token vs batch size, showing 14x headroom
14× headroom at batch 1. Decode time per output token follows a clean y = 22.9 / batch curve — textbook memory-bound behavior. The serial token dependency, not the GPU, is the limit.
RESULTS — SWEEPING THE GPU

Two workloads, every batch size, SM count swept from 2 to 108. The two classes behave nothing alike.

Input-Heavy vs. Output-Heavy
Input-Heavy vision-bound
1.9×batch efficiency
throughput keeps climbing through SM = 108
SM scalingno plateau (2 → 108)
bottleneckvision encoder (compute)
energy/req @ BS 1139 J
energy/req @ BS 161,177 J
Output-Heavy decode-bound
12.9×batch efficiency
throughput plateaus at SM = 36
SM scalingelbow @ 36 · 67% idle
bottleneckmemory bandwidth
energy/req @ BS 11,114 J
energy/req @ BS 161,378 J
Throughput vs active SMs for input-heavy and output-heavy workloads
Finding 1 — Throughput. Output-heavy (right) saturates at SM ≈ 36 across every batch size; the elbow is unmistakable. Input-heavy (left) keeps climbing to the full chip because it spends real time in compute-bound vision encoding and prefill.
Average GPU power vs active SMs, plateauing near 230W
Finding 2 — Power. Power rises linearly with SM count then flattens at ~230 W (below the 250 W TDP) at the same SM ≈ 36 threshold. Above it, output-heavy work draws no more power and does no more work — the extra 72 SMs are clocked but stranded.
Energy per request vs active SMs across batch sizes
Finding 3 — Energy & batching. Batching is a strong energy lever for output-heavy work (per-request energy barely moves while serving 16×) but a weak one for input-heavy: each request re-runs the ~1 s vision encoder, a fixed cost that never amortizes across a batch.
PHASE FINGERPRINTS

Per-kernel SM and DRAM throughput over each phase's lifetime. Three phases run hot; decode flatlines.

Vision encoder kernel trace
Vision Encoder — high SM, periodic DRAM spikes.
MLP connector kernel trace
MLP Connector — brief and compute-dense.
Prefill kernel trace
Prefill — sustained SM, healthy memory use.
Decode kernel trace
Decode — ~14% SM, ~25% DRAM for 1.5 s straight.
WHAT THIS ENABLES

The stranded capacity is real and measurable: ~67% of the GPU sits powered but unused during small-batch output-heavy decode. Single-job SM masking can't reclaim it — the GPU is already below its power limit — but two unmasked knobs plausibly can. Frequency throttling during decode exploits the memory-bound regime directly, and cross-workload spatial multiplexing could place compute-hungry input-heavy work onto the idle SMs of a decoding job. Both motivate a phase-aware serving system that provisions hardware to each phase's actual demand rather than treating inference as one monolithic block.

HOW WE MEASURED
Model
InternVL3-8B
Server
patched vLLM
GPU
NVIDIA A100 SXM4 · 108 SM · 250 W
Phase attribution
NVTX ranges
Kernel profiling
Nsight Compute
Partitioning
SM masking (libsmctrl)
SM sweep
2 → 108
Batch sizes
1 · 2 · 4 · 8 · 16
Configurations
1,080 · 5 runs each
CITE & REFERENCES
BibTeX
@misc{attelli2026mllmenergy, title = {Energy-Efficient Multi-Modal LLM Serving: A Sub-phase Characterization Study}, author = {Attelli, Aashrith and Jayr, Amaury and Lemma, Nathan and Naveen, Mahin}, year = {2026}, note = {University of Texas at Austin} }
  1. ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving. Qiu et al., ACM SoCC 2026. arXiv:2502.00937.
  2. CornServe: Efficiently Serving Any-to-Any Multimodal Models. Ma et al., 2025. arXiv:2512.14098.
  3. ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism. Liu et al., NeurIPS 2025. arXiv:2507.10069.
  4. Hardware Compute Partitioning on NVIDIA GPUs (libsmctrl). Bakita & Anderson, IEEE RTAS 2023.