Six LLM Routers Tested, And None Beat a Weighted Coin Flip

0:00Paige: Six AI routing systems were tested on a simple promise: save money by sending each question to the right language model. None of them beat randomly choosing between two well-selected models at the same average model cost. Some effectively tied. One configuration finished ... more than ten accuracy points behind. So how can reading the question be worth less than ignoring it?

0:24Eric: Maybe reading it was never where the value sat, because the phrase "well-selected" makes me suspicious, in a useful way. Somebody made a smart decision before the coin ever started flipping. So is the valuable decision really which models to keep, rather than which one to call next?

0:41Paige: That's the central tension in "Dynamic LLM Routers are Often Misguided," a preprint from Sam Wang and colleagues at Fastino Labs. This is AI Papers: A Deep Dive. The paper audits routing products, and then asks whether the way they're scored rewards the behavior users expect.

0:57Eric: A large language model router is a dispatcher. It reads your question and picks a model before any answer gets generated. The pitch is sensible: cheap questions go to a cheap model, and difficult questions justify a stronger one. This paper studies that upfront choice. It's not about a system that tries one answer and retries if it fails.

1:17Paige: And the comparison is a weighted coin, not necessarily a fifty-fifty split. The researchers use Gemini 3.7 Flash as the cheap option, and Claude Opus 5, at high reasoning effort, as the expensive one. Then they adjust how often the coin picks Opus, until the expected cost matches the router they're evaluating.

1:35Eric: And the coin's expected accuracy is just the matching blend of those two models' accuracies. A smart dispatcher should beat that, by spending its expensive calls where they help most. If it only matches the blend, reading the question hasn't shown any extra value. This baseline also comes from earlier routing research, so it isn't a stunt invented for this paper.

1:57Paige: The audit covers six systems in fourteen configurations, accessed in late August 2026. They include OpenRouter and Microsoft's Azure router. The researchers record which model each system picks. Then they generate and grade that model's answer themselves, the same way every time. So any difference comes from which model got chosen, not from how each service produces its answers.

2:19Eric: What kind of questions are these routers choosing models for? A router serving mostly short factual questions, could have a very different job from one serving hard programming tasks.

2:31Paige: They use eight hundred questions drawn from seventeen benchmarks. Those questions fall into eight task categories, weighted equally, including coding, math, knowledge, and tool use. Some systems are scored on slightly fewer questions, because requests were rejected or the chosen models weren't available. So this is a controlled benchmark mix, not a sample of real customer traffic.

2:54Eric: That makes the comparison easy to interpret, but it also puts a boundary around the opening result. These systems failed to beat randomness on this particular set of questions.

3:05Paige: Yes, and the strongest fairness check is about which models were available. Maybe the baseline got a newer, better bargain than a commercial router could pick. So the authors rerun the comparison using only models each router had access to. None clearly beats the baseline there either, but several now land within statistical noise of it. So we should say "no demonstrated win," not "every product lost badly."

3:29Eric: And these aren't identical products. OpenRouter's system goes by what's been popular lately with users making similar requests, rather than directly predicting which model will be correct. So this is an audit of routing behavior you can actually buy, not six versions of one algorithm.

3:46Paige: The really interesting part is the explanation. Across the commercial systems, sending harder questions to stronger models is a weak tendency at best. And five configurations lean toward upgrading questions that have shorter answers. The authors argue that one standard goal can reward both behaviors: maximizing accuracy under a cost budget.

4:07Eric: The difficulty part sounds backward at first. But suppose a question is so easy that either model gets it right. Upgrading buys you almost nothing. Now suppose it's so hard that both models fail. Again, upgrading buys almost nothing. The useful territory is in between, where the stronger model succeeds and the cheaper one struggles.

4:27Paige: The authors formalize that with a model borrowed from educational testing: each question has a difficulty, and each model has an overall ability. Under that assumption, the payoff from upgrading rises and then falls as questions get harder. Their measured pass rates fit this pattern reasonably well. But it isn't a universal law of how models behave.

4:49Eric: Then I'm not ready to call that behavior a mistake. If the goal is correct answers per dollar, spending money on a question neither model can solve sounds wasteful. "Hardest first" and "biggest accuracy improvement first" are different policies.

5:03Paige: They are, and that disagreement comes back when we get to the authors' replacement score. But first, there's another incentive that surprised me more: answer length. Language model services charge by the token, meaning the chunks of text they read and write. Longer answers generally cost more, yet an accuracy score usually counts each question once.

5:24Eric: So an extra correct short answer earns the same point as an extra correct long answer, but the short one can be cheaper to buy. The dispatcher doesn't have to be confused about difficulty at all. It could just be responding sensibly to the accounting.

5:39Paige: The paper tests that directly. For this demonstration, they pair Opus with a cheap model that's much weaker than Flash. Using measured difficulty and average answer length, they compare upgrading different fifths of the questions. Upgrading the fifth with the shortest answers costs ... about a dollar per thousand questions. That's roughly one sixty-seventh of what it costs to upgrade the hardest fifth. And it finishes only 1.6 accuracy points lower.

6:05Eric: Wow, that's a startling price difference for such a small accuracy gap. But this is a diagnostic that uses measured answer lengths. It isn't evidence that a deployed router can predict those lengths perfectly.

6:19Paige: Right. It reveals the incentive. The authors call it "accuracy hacking": call the stronger model when it's cheapest to do so. In their sample, longer answers tend to go with harder questions. So the score can push the premium-model budget away from exactly those harder questions.

6:35Eric: And that could explain another shortcut: recognizing what kind of question this is, rather than judging how hard this particular one is. If a benchmark usually produces short answers, spotting that benchmark could pay off financially.

6:49Paige: And most of the routers they tested, send nearly every question from a given benchmark to one model, even though difficulty varies within that benchmark. The authors test that kind of benchmark-based routing on its own. When they score it with each answer's real cost, it beats random routing at the two cheapest budgets. But when they charge every answer its chosen model's average cost instead, that advantage disappears.

7:14Eric: Oh, that's a revealing control. They take away the discount for picking short jobs, without changing any of the answers being graded. That suggests the benchmark-based policy was profiting from length-related costs, not necessarily from matching each question to some model's special expertise.

7:32Paige: There's a distinction we shouldn't lose, though. Charging the average cost is just a scoring convention. It doesn't change anyone's bill. The authors use it to separate model choice from answer length. They also don't look inside any vendor's proprietary training. Their analysis shows what the common score rewards, not how every vendor trained its system.

7:54Eric: And real bills still matter. If I asked for the most correct answers within a fixed budget, those short-answer savings could be completely legitimate. A score is only misguided relative to some goal, not just because the thing optimizing it makes a choice we find unappealing.

8:10Paige: And the authors spell out their alternative goal. They call a question hard if at least half the tested models fail it. On those hard questions, the score doesn't check accuracy; it rewards sending the request to the stronger model. On the easier questions, they compare accuracy against average model cost. They also check how well a router predicts difficulty within each benchmark.

8:33Eric: So the hard-question score rewards giving the best available model a shot, even without a demonstrated accuracy improvement. That's a preference about service quality. It isn't a finding that the extra spending produces better answers.

8:49Paige: The authors acknowledge that, and they make this correction optional. Their position is that users would rather have the stronger model attempt an exceptionally difficult question. You can disagree and stick with a strict accuracy-per-dollar goal. What you shouldn't do is confuse those two promises.

9:07Eric: That settles the scoring dispute, but not my suspicion from the opening. If the shortlist of models is poor, even a sensible policy can struggle. How much of this result comes down to which models the router can choose between?

9:21Paige: A great deal of it comes down to which models the router can choose between, as far as the evidence goes. On this evaluation, Flash trails Opus by ... just 1.6 accuracy points. And it costs ninety-six percent less. So the cheap model already captures most of the measured performance. Meanwhile, several commercial lineups include models that give you a worse deal on cost and accuracy.

9:44Eric: A small average gap doesn't settle it, though. Two models could have similar scores while succeeding on different questions. A dispatcher might exploit those complementary strengths.

9:55Paige: The authors test for that. Letting each model have its own strengths on different tasks improves their statistical fit only modestly. They also give difficulty-based routing perfect knowledge of each question's difficulty, and let it tune directly on the evaluation questions. Even with that help, a two-model lineup stays within ... one accuracy point of the best larger lineup they tested, and those had up to five models.

10:20Eric: That supports a small shortlist here, but it doesn't prove specialization never matters. They also replicate on SPROUT, an older dataset with a more varied set of models, and there, routing by benchmark beats routing by difficulty. Different models can create a different routing problem.

10:38Paige: Which makes their final experiment especially satisfying. They build a router designed to avoid the four patterns they found. A predictor reads internal signals from another language model as that model processes the question. It estimates difficulty and sends harder-looking questions to Opus. Everything else goes to Flash. They also test it on task categories the predictor never saw in training.

11:03Eric: I can feel the expected ending coming: the audit finds the flaws, the authors fix them, and their router wins. Does it?

11:10Paige: It doesn't. Their router passes the behavioral checks, including telling hard questions from easy ones within the same benchmark. But under standard scoring, its accuracy is statistically indistinguishable from random routing at every budget they tried. And under the revised scoring, it shows no clear accuracy gain on the easier questions for this pair of models either.

11:32Eric: That's an admirably inconvenient result. It still leaves two possible explanations. Either there's not much useful difference between the two models, or their difficulty predictor isn't good enough to exploit the difference that's there.

11:47Paige: Both are still on the table. The predictor is imperfect, especially on unfamiliar categories, so this isn't proof that every future router has to fail. But the small gap between the models, the weak evidence for specialization, and the perfect-difficulty comparisons all point to limited room for improvement here. Routing that looks better doesn't automatically produce more correct answers.

12:10Eric: And "here" still means an English benchmark mix, with one sampled answer per model per question. That makes the difficulty labels noisy. It doesn't establish the result for real customer traffic, or for routing again and again over a long conversation.

12:25Paige: My main takeaway is that model selection deserves testing before routing sophistication. For this set of questions, picking a strong cheap model and a strong premium model, captures most of the value anyone demonstrated. I'd compare any proposed dispatcher against a weighted random baseline, on the questions it'll actually serve.

12:44Eric: Mine is that cost and accuracy don't fully describe the experience you're promising. Maximizing correct answers per dollar can favor short questions and skip the hardest ones. If you're promising the strongest available attempt on difficult requests instead, you have to measure that separately.

13:02Paige: So why didn't reading the question beat ignoring it? On this test, the well-chosen models left little measurable advantage to exploit, and several routing behaviors that look strange made sense under the usual score. The lesson isn't that dispatching is useless. It's that a clever dispatcher needs both a choice worth making and a clearly stated job. You can find the annotated episode at paperdive dot AI: the full transcript, with every technical term tap-to-define and related papers linked by theme. If you want every major AI paper taken apart like this, daily, that's what this channel does, and you can subscribe to get them.

13:41Eric: Here's our disclosure. The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Paige and I are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is "Dynamic LLM Routers are Often Misguided," by Sam Wang and colleagues, posted October 2nd, 2026.