TokenRouter: Efficient Serving System for Token-Level LLM Routing
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.
Key points
- While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains.
- To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference.
- TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously.
- Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing.
Sources (2)
- [1]TokenRouter: Efficient Serving System for Token-Level LLM RoutingHugging Face Daily Papers · Oct 8, 12:00 AM
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.
While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains.
- [2]TokenRouter: Efficient Serving System for Token-Level LLM RoutingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:21 PM · same content
Extractive summary: sentences quoted from the sources.