LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
AuthorsTao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You
Resources
LLMRouter is an open platform and benchmark for choosing the best language model for each query while balancing answer quality, personalization, and inference cost.
Key results
Number of heterogeneous LLMs evaluated, including Llama, Mistral, Qwen, GPT-OSS, and DeepSeek-V3.1.
Total test instances spanning generic, memory, vision, video, time-series, and personalized routing.
Representative single-turn, multi-turn, and personalized router implementations in the library.
Improvement of learned routers over the strongest fixed-model baseline.
Accuracy on personalized routing evaluated with a persona-conditioned DeepSeek-V3.1 judge.
Held-out Slack user-preference accuracy.
What the paper found
LLMRouter presents a unified infrastructure for choosing among heterogeneous language models by casting routing as a sequential decision process with five components: context encoders, model encoders, scoring functions, decision rules, and learning signals. This abstraction covers single-turn, multi-turn, agentic, and personalized routing, while its automated data engine runs every query across an 18-model pool—including Meta’s Llama, Mistral, Qwen, GPT-OSS, and DeepSeek-V3.1—then records task quality, token usage, and cost in a dense query–model matrix. The resulting xRouteBench contains 4,767 test instances spanning generic reasoning, memory, vision, video, time series, and personalized dialogue. The open-source library provides 17 built-in routers, supports pointwise, pairwise, and reinforcement-learning objectives, and exposes deployments through an OpenAI-compatible server, OpenClaw, and ComfyUI. Across the benchmark, learned routers deliver a 14.6% relative improvement over the strongest fixed-model baseline, while router rankings change sharply as cost constraints tighten, demonstrating that always selecting the largest model is inefficient. Personalization is effective but depends on how user context is represented: GMTRouter reaches 68.78 accuracy with a persona-conditioned DeepSeek-V3.1 judge, whereas PersonalizedRouter achieves 83.05 accuracy on held-out real-user preferences collected through Slack. The study shows that no router dominates every task or budget, and multi-turn decomposition does not consistently justify its added inference cost.
Original abstract
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.