Online Learning for Cost-Efficient LLM Routing
This article describes how an LLM gateway optimizes model routing using an online learning approach. It specifically employs Thompson sampling over lognormal latency posteriors to dynamically select the fastest and most reliable LLM route for each incoming request. This strategy is designed to achieve cost-efficient operation by ensuring optimal performance and reliability in real-time.
Llm RoutingThompson Sampling+2Jul 2026