TileRT Support in vLLM Could Improve GPU Interactivity
By Larbi Belkhit |
17 Sep 2026 |
IN-8253
Log In to unlock this content.
You have x unlocks remaining.
This content falls outside of your subscription, but you may view up to five pieces of premium content outside of your subscription each month
You have x unlocks remaining.
By Larbi Belkhit |
17 Sep 2026 |
IN-8253
NEWSvLLM and TileRT Combine for Low-Latency GPU Inference |
Disaggregated inference has become a widely adopted pattern in 2026. Many inference providers optimize different compute clusters to serve prefill and decode separately to improve throughput and latency. More advanced disaggregated serving architectures involve leveraging compute architectures other than a Graphics Processing Unit (GPU), with companies like Cerebras and SambaNova seeing growing demand for their compute. The main value proposition for this silicon has always been its benefits within the decode phase of inference, providing much stronger Tokens per Second (TPS) as agentic workloads continue to grow. The added benefit of disaggregating inference—beyond leveraging the best silicon architecture for each phase—is that the underlying inference engine can be modularized.
In July 2026, vLLM introduced a second decode option for its inference engine from TileRT, integrated via its public connector interface. The vLLM prefill pool, scheduler, and serving Application Programming Interface (API) remain the same as before, but the decode pool uses a router-gated dispatch depending on whether the traffic is latency critical, enabling the engine to route it to the TileRT decode compute pool. TileRT is a new inference runtime built to optimize interactivity (TPS/user), with its initial beta release back in November 2025, and hitting production support for Zhipu AI’s GLM 5.2 model in May 2026. TileRT’s persistent engine statically compiles the entire inference decode graph onto a single persistent kernel on NVIDIA GPUs, rather than generating many kernels continually through the decode process.

IMPACTGPU Utilization, Premium Tiers, and Token Pricing |
As discussed in a previous ABI Insight, “Inference Is Disaggregating to Balance Performance, Latency, and Cost,” one of the principal benefits of disaggregated inference is its ability to reshape cloud tokenomics. The current market trend is toward throughput isolation—segmenting token traffic across premium and base model access tiers so providers can dedicate necessary compute to meet Service-Level Agreements (SLAs). By improving decode performance on top of existing GPU infrastructure, providers can offer higher-performance tiers and further justify higher premiums, which typically begin at around 50% above base pricing or above. At the other end of the spectrum, offline batch inference—where latency requirements are minimal—has a 50% haircut compared to base mode levels. Table 1 highlights several examples across popular models and compares the distinction between base and fast model service tiers.

All of this helps inference providers maximize the Return on Investment (ROI) of their compute fleets as agentic workloads scale. Improving GPU decode performance enhances the value of already-deployed GPU capacity by increasing overall utilization and enabling more differentiated service tiers. Given the strong demand from buyers and the still capacity-constrained nature of the market, this represents a path of least resistance that should be well received by both inference specialists and cloud providers. Support from vLLM—one of the largest open-source inference engines—for TileRT exposes its software technology to an enormous amount of developers within the cloud ecosystem, which can fast-track adoption.
A software-layer optimization such as TileRT does not eliminate the underlying bandwidth constraints that limit GPU performance for inference decode. As a result, while these optimizations extend the economic life and competitiveness of GPU infrastructure, they are unlikely to displace the longer-term shift toward custom inference silicon. Custom silicon will continue to see increasing deployments across the cloud ecosystem over time.
RECOMMENDATIONSThe Industry Needs an Answer—Real Breakthrough or Cool Proof of Concept? |
Short term, TileRT’s adoption will be constrained by practical factors, particularly with model support. Unless this meaningfully expands through community contributions or directly through the TileRT team, it will remain a narrow technological innovation that does not get leveraged in production workflows. Furthermore, the commercial impact of TileRT may come less from its direct adoption and more from imitation. Inference providers with proprietary stacks, such as Fireworks AI, may choose to replicate the architectural principles of TileRT into their own engines, if they are not already using a similar approach.
For GPU vendors other than NVIDIA, the response should be measured rather than reactive. Qualcomm should explore whether a similar architectural approach is feasible and can be implemented within Modular’s MAX engine, particularly given its compute-agnostic strategy. For Intel & AMD, as TileRT was designed primarily with NVIDIA deployments in mind, both should engage TileRT and explore integration with their GPU stacks to publish early performance data for Prefill-Decode (PD) disaggregation. For developers, clear benchmark evidence that comparable decode-side gains can be achieved on non-NVIDIA accelerators would be valuable in reinforcing TileRT’s credibility & portability.
NVIDIA should be slightly more proactive, particularly as it should look to integrate similar support for TileRT into TensorRT-LLM, rather than leaving this path solely on vLLM. Doing so would provide developers with more native alternatives inside NVIDIA’s stack, as SGLang has not yet formally announced support either. Combining this with the performance optimizations of Dynamo, NVIDIA could concretely provide evidence for the industry whether TileRT’s decode architecture is a step-change for interactivity for GPU compute.
Written by Larbi Belkhit
Related Service
- Competitive & Market Intelligence
- Executive & C-Suite
- Marketing
- Product Strategy
- Startup Leader & Founder
- Users & Implementers
Job Role
- Telco & Communications
- Hyperscalers
- Industrial & Manufacturing
- Semiconductor
- Supply Chain
- Industry & Trade Organizations
Industry
Services
Spotlights
5G, Cloud & Networks
- 5G Devices, Smartphones & Wearables
- 5G, 6G & Open RAN
- Data Centers
- Enterprise Connectivity
- Space Technologies & Innovation
- Telco AI
AI & Robotics
Automotive
Bluetooth, Wi-Fi & Short Range Wireless
Cyber & Digital Security
- Citizen Digital Identity
- Digital Payment Technologies
- eSIM & SIM Solutions
- Quantum Safe Technologies
- Trusted Device Solutions