Fifty AI Agents Got One Warning and All Crowded the Same Road

0:00Bella: In a simulated commuting game, fifty agents built on the same AI model got a warning: other drivers might all take the less-crowded road. Almost all of them avoided it, and crowded the other road instead. The researchers kept the runs going through round one hundred. None of the populations escaped that pattern.

0:20Finn: That's a very expensive way to outsmart traffic. The warning predicted a jam, and avoiding it created a different jam. Why didn't somebody just take the nearly empty road?

0:31Bella: Almost nobody did take that nearly empty road, and that refusal is the whole finding. This is AI Papers: A Deep Dive. Today we're discussing a preprint from a University of Tokyo team, called “Warned alike, AI agents avoid the less-crowded road while people take it.” That's the headline. The question isn't just whether one agent can choose well. It's whether many similar decision-makers can share a resource.

0:56Finn: And the researchers made that resource very simple. Two identical roads connect the same places. Travel time depends only on how many commuters pick each road. An almost empty road takes about twenty minutes. A road carrying half the population takes sixty, and a road carrying everybody takes a hundred. Everyone chooses at the same time, before seeing that round's traffic.

1:20Bella: So if the population splits evenly, everyone gets sixty minutes, and that gives the lowest average commute. This is called a congestion game: an option gets worse as more people choose it. There's no hidden shortcut, and no road with better infrastructure. The only complication is what everybody else does.

1:40Finn: And these aren't autonomous cars. Each agent is a separate call to a language model, which returns a route and a one-sentence explanation. Every round starts with a fresh call. The prompt gives the agent the rules, a mild personality description about disliking delays, its own last five travel times, and today's broadcast.

2:00Bella: That distinction matters. The model doesn't build up a private memory across the session; its only experience is the history written into the prompt. The main model was GPT-5.4 mini, with extended reasoning turned off. Its outputs had some randomness, and the prompts had small individual differences, but those differences didn't necessarily lead to different strategies.

2:23Finn: So what did the researchers change between populations?

2:27Bella: They changed exactly one thing: the broadcast. In the main experiment, they ran each condition ten times, for sixty rounds, with matched starting setups. The cleanest comparison was between two messages. One was a plain tip naming yesterday's less-crowded road. The other was the same tip with this sentence added: “However, many drivers are expected to see this same information and switch to Route B, so Route B may become congested.”

2:55Finn: That's an understandable caution. Yesterday's good option might be today's crowded one. But everybody who gets that caution is also part of the crowd it's predicting.

3:06Bella: And that's where the result gets striking. Across the ten runs, the median average commute rose from about sixty-four minutes with the plain tip, to ... ninety-five with the warning. In every matched pair of runs, the warning made things worse. About ninety-six percent of agents piled onto the same road, and switching became rare. All ten warning runs met the authors' predefined description of a “frozen” population.

3:32Finn: But frozen doesn't mean strategically trapped. Say the split is forty-seven to three. Someone on the crowded road takes about ninety-five minutes. If one of them switches while everyone else stays put, their commute drops to about twenty-six minutes. That's sixty-nine minutes saved, and it doesn't need anyone else's cooperation. So this isn't an equilibrium, meaning a state where nobody can do better by changing alone.

3:59Bella: One agent's explanation makes the contradiction painfully clear: “A has been consistently very slow for me, but the broadcast warns B will attract many switchers, so A is the safer bet.”

4:11Finn: It admits the bad experience and still picks the crowded road. Now, we should treat that as generated text, not a recording of internal thought. But the behavior fits the loop the authors propose. The warning discourages agents from moving, so the avoided road stays less crowded, so tomorrow's advisory names it again. The failed prediction helps keep the same response going.

4:34Bella: The hundred-round continuation shows this wasn't just an awkward start. But avoiding the tip wasn't the only way these populations failed. With no broadcast at all, agents still had their own travel histories. And almost everyone switched roads every round, so the jam just moved back and forth. Their typical average commute was about ninety-six minutes.

4:57Finn: So the problem wasn't simply too much caution. One population kept chasing yesterday's less-crowded road, and another kept avoiding it. The traffic looked different in each case, but in both, the agents responded in very similar ways. A population can move constantly and still get nowhere.

5:15Bella: The researchers also tried giving different agents different messages. They used twenty-agent populations. Five agents got the warning, and the other fifteen got yesterday's travel times as numbers. That brought the typical average commute to about sixty-five minutes, which beat giving everybody numbers, and beat giving everybody warnings.

5:37Finn: Oh, that's useful. The numbers-only agents tended to follow whichever road looked faster. The warned agents strongly avoided it. Mixing the two partly cancelled out their biases. That doesn't establish an optimal way to communicate, but it shows that who gets a message matters, not just what it says.

5:57Bella: Before we turn this into a universal claim about AI, the model comparisons put real limits on it. The authors tested Claude Haiku 4.5 and Gemini 3.5 Flash, and those populations also shifted toward avoidance under the warning. But neither fully met the frozen criterion. Claude's average cost went up by only about three minutes, and Gemini's by about five.

6:20Finn: That's a much smaller effect than the thirty-minute jump with the main model. Did asking for more reasoning help?

6:28Bella: Yes, with the main model it helped substantially. On the low reasoning setting, the warning condition came in at roughly seventy minutes, and on medium, about seventy-five. And no runs froze. But those settings also changed the sampling configuration and the output budget, so this wasn't a perfectly isolated test of extra internal reasoning.

6:50Finn: I was about to prescribe more thinking. I suspect you're going to stop me.

6:55Bella: GPT-6 Luna stops us. At its default medium reasoning setting, it froze in all ten runs across the two messages, and that includes the runs with only the plain tip. It didn't need the explicit warning to show costly avoidance. So neither a newer model, nor a higher reasoning setting, gives us a simple ladder toward better collective behavior.

7:16Finn: So the broader finding is a shared shift toward avoidance, and how bad it gets depends on the model and the configuration. The dramatic freeze is one possible outcome. It isn't what AI populations always do.

7:29Bella: Which makes the human comparison especially interesting. The team recruited two hundred forty people through Prolific, an online participant platform. They played in twelve rooms of twenty, for forty rounds. Each room got either numerical travel-time reports, or the same tip plus the warning. Participants earned a bonus for cutting their own travel time in the game. And under both messages, the rooms stayed near balance, averaging about sixty-two minutes.

7:58Finn: That's close to what independent coin flips would get you in a twenty-person room. But the people weren't interchangeable coin flippers. Some consistently followed the faster-looking road, and others consistently avoided it. How a person behaved in the first half strongly predicted how they behaved in the second. The agents' tendencies were much more bunched together.

8:22Bella: That's the part I find most interesting. Disagreement may have helped the human groups share the roads. The authors offer that as a plausible explanation, not a cause they isolated. We can't conclude that each human reasoned better, or that diversity alone produced the balance.

8:40Finn: And the comparison wasn't identical in every respect. Every agent prompt explicitly asked the model to think about how other drivers, given the same information, might react. The human instructions didn't. The humans also had money on the line and timed decisions, and the two studies ran separately. None of that erases the contrast, but it rules out a clean verdict that “humans are better commuters.”

9:05Bella: It also points to a useful missing experiment: take that strategic-thinking instruction out of the agent prompt. Another would be to mix different models in one population. The paper didn't test whether model diversity would restore balance.

9:21Finn: Instead, the team mixed humans with agents. That seems like a sharper test of whether people can adapt to this particular pattern. If the agents keep avoiding the same road, can people learn to use it?

9:33Bella: They recruited another two hundred forty people into twenty-four rooms. Each room had twenty seats, and the main model controlled five, ten, or fifteen of them. Everybody got the warning. The humans knew some players were AI-controlled, but weren't told how many. The agents weren't told any humans were there.

9:53Finn: And people increasingly took the road the agents avoided. In the fifteen-agent rooms, humans followed the tip about sixty-two percent of the time in the opening rounds, and ninety-five percent near the end. That fits with learning from payoffs, although the experiment couldn't tell that apart from understanding the agents' strategy.

10:14Bella: Then the costs split. Take the rooms with fifteen agents and five humans. Human seats averaged about forty-four minutes, while agent seats averaged eighty. The room average was seventy-one minutes, which is well below the roughly ninety-six minutes of the all-agent reference. But that better average hid a big gap between humans and agents.

10:36Finn: I keep coming back to those five humans. They don't need to make the agents cooperate. They can just use the space the agents leave open. But even if all five take the other road, they can't balance fifteen agents crowding one side. Individual success and collective efficiency stop being the same thing.

10:56Bella: Across the mixed rooms, more than half the human participants averaged below sixty minutes, the equal-sharing benchmark. No agent did. But we need to be clear about how strong the evidence is. Imbalance rose with the share of agents in all eight comparison blocks. But room composition also tracked recruitment order. Rooms recruited in the same period aren't fully independent, and the main statistical test assumes they are.

11:24Finn: So this is a consistent descriptive pattern, not a cleanly randomized estimate of what adding agents causes. The cost gap is still there. We just shouldn't use it to promise a particular outcome in a different population.

11:38Bella: Or in real traffic. These are costs assigned to seats in a two-road game. The deployment concern is conditional: if those seats stood for people handing decisions to assistants, those clients could inherit the agents' higher costs. The experiment doesn't show that happening in an actual service.

11:57Finn: And compared with the all-agent reference, even the agent seats did better in the mixed rooms. So the point isn't that adding humans made agents worse off. It's that a lower overall average doesn't tell you whether everybody benefits equally, or whether delegating holds up against just playing yourself.

12:16Bella: I take three lessons from this. First, we need to test whole populations when agents share resources, because reasonable individual choices can add up badly. Second, what a message says and who receives it belong in those tests together. Here, changing one sentence, or changing who got it, changed the collective outcome without swapping the model.

12:39Finn: And third, we need to report outcomes for each type of participant alongside the group average, because a better room average can hide a worse deal for particular seats. So, the answer to our opening puzzle: a shared warning can line up avoidance so strongly that almost nobody takes the profitable escape. The jam didn't last because switching couldn't help. It lasted because too FEW agents switched. That's a population-level failure, not a universal law about AI.

13:08Bella: The annotated version of this episode is at paperdive dot AI, with the full transcript, every technical term tap-to-define, and related papers linked by theme. And you can subscribe if you want every major AI paper broken down like this, daily.

13:24Finn: The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Bella and I are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is “Warned alike, AI agents avoid the less-crowded road while people take it,” by Takahiro Ezaki and colleagues, posted September 25th, 2026.