ResearchPod Summary
As organizations scale their use of Large Language Models (LLMs), they face a critical trade-off: using the most capable models for every query is prohibitively expensive, while smaller models may lack the necessary performance for complex tasks. Existing routing methods often fail to enforce global, workload-level budget constraints (e.g., a monthly API quota) or require expensive, dense supervision to train routing policies. This paper asks: how can we perform efficient LLM routing that respects a global budget while learning from sparse, real-time interactions?
The authors formulate LLM routing as a Constrained Contextual Multi-Armed Bandit (CCB) problem. They introduce WISERouter (WR), which uses Adaptive Linear Programming (ALP) to translate a global budget into dynamic per-query constraints.
To bridge the gap between continuous query embeddings and the discrete requirements of ALP, the authors use clustering to group queries into a finite set of contexts, allowing the router to make decisions based on the expected performance and cost of each model within that cluster.
Empirical evaluations on RouterBench and SWE-Bench show that WISERouter consistently outperforms existing baselines. In data-rich scenarios, WR-Offline improves average response quality by 14% over the best baseline on SWE-Bench under tight budget constraints. In data-scarce scenarios, WR-Online achieves performance comparable to offline methods while requiring 90% less exploration data. The framework is also significantly faster to train than neural-network-based routing approaches, as it avoids complex model retraining when budget constraints change.
This work provides a theoretically grounded and practically efficient solution for enterprises managing high-throughput LLM workloads. By shifting the focus from per-query constraints to workload-level budget management, WISERouter allows organizations to maximize the utility of their API spend, ensuring that high-complexity queries receive the necessary compute power while simpler queries are handled by more cost-effective models.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.