Helion Developers Claim Over 10% Higher AI Serving Throughput on Some H100 Workloads
The vLLM integration automates workload-specific GPU optimization, while keeping existing libraries for larger inputs. Tuning time and configuration upkeep remain adoption hurdles.
Helion’s vLLM integration pairs autotuned kernels for small-token quantized matrix multiplications with existing GPU libraries for larger inputs, rather than replacing the serving stack. In the developers’ H100 tests across Qwen3 models and three eight-bit formats, the fork reports workload-specific serving gains, while kernel-only comparisons reached 11.0%–17.8% depending on format and baseline. Users still face hours of tuning and cold-start compilation; upstream adoption remains unsettled, with workload-specific configuration upkeep a key obstacle to easier use.
01
Helion handles inputs of up to 32 tokens under CUDA Graph replay; larger inputs fall back to the existing backend.
02
Kernel speedups over CUTLASS were 11.0% for FP8_Dynamic and 17.8% for W8A8_INT8; Block_FP8 gains were measured against other baselines.
03
The tests used an 80GB NVIDIA H100 and Qwen3 models from 1.7B to 32B parameters, plus Qwen3.8-27B.
The Helion backend keeps vLLM’s established GPU libraries for larger inputs and uses tuned kernels for smaller ones. In an October 2 PyTorch write-up, its developers report more than 10% higher end-to-end serving throughput on some H100 workloads. The implementation and pre-tuned configurations are available in their vLLM fork.
vLLM is software for running and serving language models. Helion is a PyTorch-native language for writing GPU kernels—the routines that perform individual computations. This integration targets NVIDIA Hopper GPUs and three eight-bit formats: FP8_Dynamic, W8A8_INT8 and Block_FP8. Its focus is the matrix multiplication used in quantized linear layers.
One implementation, several ways to multiply
Rather than write separate kernels and rules for choosing them, the developers expose different execution methods within one implementation. Helion’s autotuner benchmarks candidate settings for each input shape—the dimensions of the matrices—and selects the best-performing combination. Its search covers memory layout, scheduling and algorithm choice.
The choices include ordinary matrix multiplication, Split-K, which divides work across GPU thread blocks, and Swap-AB, which rearranges the multiplication to improve GPU use for small inputs. Keeping these choices tunable replaces hand-written selection rules with measured, workload-specific decisions.
The published setup limits Helion to inputs of up to 32 tokens under CUDA Graph replay. That avoids CPU dispatch and launch overhead that could otherwise erase the kernel gains. Restricting tuning to this small-token range also reduces the configurations developers must generate and maintain.
What the H100 numbers measure
The developers benchmarked an NVIDIA H100 with 80GB of HBM3 memory. Tests covered Qwen3 models from 1.7 billion to 32 billion parameters, plus Qwen3.8-27B. The results describe those models, that GPU and the three supported formats—not a universal serving-speed improvement.
For isolated kernels, the reported geometric-mean speedups over CUTLASS were 11.0% for FP8_Dynamic and 17.8% for W8A8_INT8. Block_FP8 showed 14.9% over FlashInfer and 17.7% over DeepGEMM. These summarize kernel tests, not whole-server throughput; individual input shapes also produced different gains.
The work moves ahead of serving
Fine-grained tuning can still take hours. Startup can also trigger just-in-time compilation during CUDA Graph capture, increasing cold-start latency. Caching compiled artifacts can largely remove that compilation penalty on warm starts, but it does not eliminate the initial preparation.
Availability in the fork should not be confused with default upstream adoption. The June 23 vLLM RFC describes explicit opt-in and says startup fails if required tuned configurations are missing. It directs users to generate configurations for their own workloads.
The October write-up identifies configuration upkeep as a remaining upstream hurdle: large pre-tuned files are difficult to validate exhaustively. The developers currently ship them in the vLLM fork, but are exploring keeping kernels and integration upstream while leaving workload-specific tuning to users. That approach favors performance and maintainability over out-of-the-box usability.
Sources
github.com[RFC]: Add Helion linear backend for vLLM · Issue #46526 · vllm-project/vllm
pytorch.orgBuilding a High-Performance and Portable vLLM Linear Backend with Helion
Reader comments
Newest comments first. Replies stay oldest first.