What actually is Model Routing? A deep dive into cost, efficiency & risk

SEP 22, 202624 MIN

Description

Model routing sounds like a small architectural detail, but it’s quickly becoming the control point that determines whether AI systems are fast, affordable, and trustworthy. In this episode of Pop Goes the Stack, F5's Lori MacVittie and Joel Moses are joined by Patrick Roughan to unpack why routing decisions for LLM workloads can’t be treated like ordinary traffic distribution. The real challenge isn’t simply getting requests to an available endpoint, it’s choosing the right model and the right execution path based on intent, context, and policy. Patrick explains how context-aware routing changes everything, starting with KV cache locality. When similar prompts land on infrastructure that already holds relevant cached state, you avoid expensive recomputation and improve response times. But “smart routing” goes beyond cache. Different models have different strengths, costs, and risk profiles, and enterprises increasingly need to steer requests based on what’s being asked, who’s asking, and what data is included. The conversation also touches on how some systems are evolving toward specialization, including architectures that effectively route within a model family, and why governance has to be part of the routing layer. When data sovereignty, privacy, and regulations come into play, routing becomes a policy decision, not just a performance decision. Sometimes the right answer is to send a request to a local model, and sometimes it’s to block a request entirely after inspecting it for sensitive content. The takeaway: GPUs are constrained and costs are real, so the winning strategy isn’t throwing more hardware at the problem. It’s building routing intelligence that optimizes for efficiency, correctness, and compliance at the same time.