Systems for Machine Learning · Final Project · UT Austin · Spring 2026
Energy-Efficient Multi-Modal LLM Serving
A Sub-phase Characterization Study
Aashrith Attelli · Amaury Jayr · Nathan Lemma · Mahin Naveen
University of Texas at Austin
OVERVIEW
When you ask a multimodal AI to describe an image, the request doesn't hit the GPU as one uniform job. It moves
through distinct phases — looking at the image, reading your prompt, then
writing the answer one word at a time — and each phase stresses the hardware completely
differently. Yet today's serving systems hand the entire GPU to every phase equally.
We profiled InternVL3-8B phase by phase on an NVIDIA A100, then swept how much of the GPU each
phase actually gets — from 2 SMs up to all 108 — across 1,080 configurations. The headline:
the phase that takes the most time uses the GPU the least, and roughly two-thirds of the chip often
sits powered but idle.
KEY FINDINGS
Decode — generating the answer token by token — eats 88.7% of every request's wall-clock time, yet sustains only 14% SM and 23% memory throughput. The slowest phase is also the one that wastes the most hardware.
Output-heavy throughput flattens once a third of the GPU is active. The remaining 67% of streaming multiprocessors add under 9% more throughput. Input-heavy work, dominated by the vision encoder, keeps scaling all the way to SM = 108.
Power saturates at ~230 W at that same SM = 36 elbow — below the A100's 250 W limit. Because the GPU never reaches TDP, masking a single job's stranded SMs can't recover the energy. Batching can: 13× more efficient for output-heavy work.
THE FOUR PHASES
One multimodal request, four phases. Pick any to see what it does and how hard it pushes the GPU.
Input pipeline · 3 phases
WHERE THE TIME GOES
Decode dominates the clock while barely touching the silicon — the gap this study is about.
Wall-time per request
InternVL3-8B · BS 1 · 1344² image · 64 tokens
Per-phase GPU utilization
% of peak sustained
Vision Encoder, MLP Connector and Prefill all keep the compute units busy. Decode is the outlier — low on both axes, for the longest time. Toggle the metric to compare.
14× headroom at batch 1. Decode time per output token follows a clean y = 22.9 / batch curve — textbook memory-bound behavior. The serial token dependency, not the GPU, is the limit.
RESULTS — SWEEPING THE GPU
Two workloads, every batch size, SM count swept from 2 to 108. The two classes behave nothing alike.
Input-Heavy vs. Output-Heavy
Input-Heavy vision-bound
1.9×batch efficiency
throughput keeps climbing through SM = 108
SM scalingno plateau (2 → 108)
bottleneckvision encoder (compute)
energy/req @ BS 1139 J
energy/req @ BS 161,177 J
Output-Heavy decode-bound
12.9×batch efficiency
throughput plateaus at SM = 36
SM scalingelbow @ 36 · 67% idle
bottleneckmemory bandwidth
energy/req @ BS 11,114 J
energy/req @ BS 161,378 J
Finding 1 — Throughput. Output-heavy (right) saturates at SM ≈ 36 across every batch size; the elbow is unmistakable. Input-heavy (left) keeps climbing to the full chip because it spends real time in compute-bound vision encoding and prefill.
Finding 2 — Power. Power rises linearly with SM count then flattens at ~230 W (below the 250 W TDP) at the same SM ≈ 36 threshold. Above it, output-heavy work draws no more power and does no more work — the extra 72 SMs are clocked but stranded.
Finding 3 — Energy & batching. Batching is a strong energy lever for output-heavy work (per-request energy barely moves while serving 16×) but a weak one for input-heavy: each request re-runs the ~1 s vision encoder, a fixed cost that never amortizes across a batch.
PHASE FINGERPRINTS
Per-kernel SM and DRAM throughput over each phase's lifetime. Three phases run hot; decode flatlines.
Vision Encoder — high SM, periodic DRAM spikes.
MLP Connector — brief and compute-dense.
Prefill — sustained SM, healthy memory use.
Decode — ~14% SM, ~25% DRAM for 1.5 s straight.
WHAT THIS ENABLES
The stranded capacity is real and measurable: ~67% of the GPU sits powered but unused during small-batch
output-heavy decode. Single-job SM masking can't reclaim it — the GPU is already below its power limit — but two
unmasked knobs plausibly can. Frequency throttling during decode exploits the memory-bound regime
directly, and cross-workload spatial multiplexing could place compute-hungry input-heavy work onto
the idle SMs of a decoding job. Both motivate a phase-aware serving system that provisions hardware to each phase's
actual demand rather than treating inference as one monolithic block.
HOW WE MEASURED
GPU
NVIDIA A100 SXM4 · 108 SM · 250 W
Phase attribution
NVTX ranges
Kernel profiling
Nsight Compute
Partitioning
SM masking (libsmctrl)
Batch sizes
1 · 2 · 4 · 8 · 16
Configurations
1,080 · 5 runs each
CITE & REFERENCES
BibTeX
@misc{attelli2026mllmenergy,
title = {Energy-Efficient Multi-Modal LLM Serving:
A Sub-phase Characterization Study},
author = {Attelli, Aashrith and Jayr, Amaury and
Lemma, Nathan and Naveen, Mahin},
year = {2026},
note = {University of Texas at Austin}
}
- ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving. Qiu et al., ACM SoCC 2026. arXiv:2502.00937.
- CornServe: Efficiently Serving Any-to-Any Multimodal Models. Ma et al., 2025. arXiv:2512.14098.
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism. Liu et al., NeurIPS 2025. arXiv:2507.10069.
- Hardware Compute Partitioning on NVIDIA GPUs (libsmctrl). Bakita & Anderson, IEEE RTAS 2023.