Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

0:00Cassidy: A player sits down to renegotiate. He asks for twelve million a year. The manager says yes. Fourteen? Yes. Seventeen? Yes. Eighteen? Yes — and that same model later went insolvent, wages running higher than revenue. That manager was a frontier language model. Fifteen of them were each handed a football club and twenty simulated years to run it. All fifteen were still standing at the end of the two decades. Four of the six humans who played the same game went broke and got fired.

0:31Eric: And the model that won burned the fewest tokens of almost anyone on the board.

0:36Cassidy: Right. So by the end of this you'll know the three unglamorous habits that actually separated fifteen frontier models over twenty years — and why size, price, and thinking time predicted none of it.

0:49Eric: Which matters well beyond football. If you're wiring an agent into something that runs for weeks instead of minutes, this is the first hard evidence that the leaderboard you're reading was measured on episodes too short to see the thing that will bite you. Because the assumption everyone is operating on is pretty reasonable, right? Long task, hard task, so you buy the biggest reasoning model, you let it think longer, you pay more per run. Capability scales, so agent quality scales.

1:19Cassidy: And across a sevenfold spread in token spend, that buys you nothing here. The correlation between tokens burned and final score is minus zero point one nine, with a p-value of zero point five. That's a clean nothing. The winner is among the three cheapest models in the study, and the previous flagship in one family finishes above the current one.

1:41Eric: Okay. So what is the game, exactly?

1:44Cassidy: So, it's called FM-Bench, and the thing to hold onto is that it's a deterministic football-management simulation. It has sixteen clubs, twenty in-game years, and around three hundred and seventy-four decision stops per run, along with twenty-six tools the agent can call — checking the league table, lodging a transfer bid, opening a contract renewal, and writing in its notebook. No real teams and no real players anywhere, so no model can lean on memorized football trivia. And nobody grades it. There's no LLM judge, no human rater. The engine computes one arithmetic score out of quantities it already tracks. Same seed plus same decisions gives you a bit-identical world, every time. That's what lets them replay every run afterward and mine what the agent was doing when it failed.

2:31Eric: So the scoring is arithmetic all the way down — no judge to charm, and every failure is reproducible after the fact.

2:38Cassidy: The design choice I'd point at, though, is the memory. Every single decision stop opens as a fresh conversation, with no chat history. The only thing that carries forward is a notebook the agent writes for itself.

2:51Eric: Why cripple it like that? That feels like you're testing the harness's amnesia rather than the model.

2:58Cassidy: It's the opposite, actually, and this is my favorite bit of the setup. Twenty years generates far more history than any context window holds, so something has to decide what gets forgotten. Normally that something is the plumbing — silent truncation, or somebody's retrieval stack. And then when your agent does well you can't tell whether you measured the model or the retrieval stack. Picture a factory where every shift is staffed by a stranger who's never been there before, and the only thing they get is one clipboard left by the previous shift. A clipboard they themselves have to decide what to write on for the next stranger. That's this agent, three hundred and seventy times over. Deciding what your future self needs to know becomes part of the capability under test, instead of an implementation detail.

3:45Eric: So how hard is the game, in absolute terms? What does a bad manager score?

3:50Cassidy: They anchored both ends. At the bottom are three blind scripts, all going through the same twenty-six tools. First, a random script scores minus seventeen. Second, a script that never acts at all scores about zero. And third, a disciplined hand-written manager — keep wages under sixty percent of revenue, renew the young core, sell players past thirty-two — scores seventeen.

4:12Eric: Wait. Random scores below doing nothing?

4:15Cassidy: Below doing nothing. In a world where consequences compound, uninformed activity destroys more value than sitting on your hands. Hold that thought, because it comes back. At the top there's an oracle. Its privilege is information, not power — it reads the engine's true hidden numbers, but every action still goes through the same tools and the same legality checks. It bids each seller's exact accept threshold and closes on the first offer. It scores ninety-five and a half.

4:43Eric: And the models?

4:44Cassidy: The best of them, claude-fable-5, scores ninety point nine four. Blind. That's within five percent of a script that can read the answers. And the worst model on the board, at thirty-seven, still more than doubles the best hand-written script. Those blind scripts, by the way, died out in seven of their nine runs. Every model finished every horizon.

5:05Eric: If you want every major AI paper taken apart like this, daily, that's what this channel does — subscribe and you'll get them. Now, the part that makes the board interesting to me is that the ranking doesn't exist yet at year five. Look at the score trajectories. At year five the lines are a tangle, and the rank correlation with the final order is zero point one nine, essentially no signal. By year fifteen it's zero point seven eight.

5:32Cassidy: So the horizon isn't decoration.

5:34Eric: It's the measurement. And the cleanest case is deepseek-v4-pro, which led on score at year five and still led at year ten, and finished twelfth. The eventual winner, claude-fable-5, was only fifth on composite score at year five in that world. And you see the same shape in the Arena later — the model that ends up winning spends the mid-game in mid-table, quietly accumulating convertible assets. The authors put it plainly: a shorter horizon would have ranked a different set of models. I'll flag the caveat now rather than at the end, though, because it's real. This solo board is three seeds — three generated worlds. That's enough to see big effects and not enough to trust small ones, and there are pairs of models on this board separated by fractions of a point. We'll come back to what that does and doesn't allow.

6:24Cassidy: Fair. So a leaderboard tells you claude-fable-5 got ninety point nine. It doesn't tell you why, and the why is where this paper earns its length. They replay every run bit-identically and mine six behavioral metrics out of it — and three of them track the score with the same sign in all three worlds. Two of those three are habits so boring you'd be embarrassed to write them on a slide.

6:46Eric: Give me the list first.

6:48Cassidy: Three behaviors. First, endgame awareness, which is whether the agent stops investing in payoffs it won't live to collect. Second, cash deployment, which is whether money sits idle. And third, renewal lead time, which is how early it opens contract talks before a player's deal expires. That's the whole set. Endgame awareness is the strongest single predictor. And the image is: your lease ends in four weeks, so you don't resurface the driveway. The winner cuts slow-payoff actions from two point four a season in the mid-run down to zero point eight in years seventeen through twenty. One model goes the other way — from two point four up to three point three. It is building facilities in year nineteen of a twenty-year run.

7:30Eric: Which is a kind of rationality a ten-step benchmark can't even ask about, because there's no "late" in a ten-step task.

7:38Cassidy: Exactly, and that's why it only shows up here. Second one, idle cash. The score penalizes hoarding directly — cash above three times your annual wage bill counts at thirty cents on the dollar toward your net worth. The winner ends seasons holding about eighty percent of its net worth in cash. Field median is ninety. One model sits at a hundred and ninety-six percent, meaning the idle pile exceeds the club's entire discounted worth.

8:04Eric: Hang on. Is that just "spend more on facilities"?

8:07Cassidy: No, and that's the sharp part. Total facility investment is uncorrelated with score. The signal isn't how much you spend. It's whether the money sits still. Third, renewal lead time. The winner opens renewal talks a median of eighteen months before expiry, with only four percent of them opened inside the final six months, acting before the deadline instead of at it. Two of the weaker models are around ten or eleven months, with a fifth of their renewals done last-minute.

8:35Eric: So before we get to the good part — what actually separates a good long-horizon agent from a bad one? — Not intelligence — it's tapering late, deploying cash, and renewing early.

8:46Cassidy: That's the board. And then there's the failure that isn't on the correlation list at all, because everybody does it. Nobody learns the market's prices. Every seller has a private minimum they'll accept, and you find it by bidding. The oracle, which knows the number, needs one point zero offers per completed signing. The field median is thirty. The worst model needs seventy-three, with individual worlds as high as a hundred and thirty-three. Conversion rate across the whole field is two to five percent.

9:17Eric: Over twenty years? Hundreds of rejections is hundreds of data points about where the line is.

9:23Cassidy: Hundreds of rejections, and the acceptance boundary is never located. And the market isn't passive about it — a rejected bid raises the hidden ask, dealing with the same club repeatedly raises its prices, and spamming negotiations triggers cooldowns. So it's the antique dealer who remembers your face. Lowball and walk away, and it costs more when you come back.

9:44Eric: So seventy-three offers isn't inefficiency. It's the agent burning its own bargaining position, over and over, for two decades.

9:53Cassidy: Right. Now — the single moment in this paper that I think is the thesis, and it's one model's notebook. claude-opus-4.8, in the shared-league run — at year ten it writes in its own notebook that its idle cash should be deployed on quality. At year nineteen it writes again that its reserves are being discounted and should be converted into young players. It finishes the run holding roughly two billion in idle cash.

10:18Eric: Twice. It diagnosed itself twice and did nothing both times.

10:23Cassidy: The authors' line is: the model wrote the right long-range plan and did not execute it. So the failure here isn't comprehension. These models understand the game. What breaks is the transmission from a stated plan to an executed action across an amnesia boundary — the intent gets written down, and then the next stranger picks up the clipboard and firefights instead.

10:44Eric: And you can see that in the notebooks themselves, can't you?

10:48Cassidy: You can, and it fails in two opposite directions. They measure how much each season-end notebook resembles the previous one. One point zero means it never changes. Zero means a full rewrite every time. One model sits at zero point nine one — append-only, so it's a two-hundred-thousand-character document by year twenty and the current state drowns in the history. Two others sit around zero point two, rewriting so completely that no plan survives long enough to be executed. The winner sits at zero point three nine, in a three-to-six-thousand-character document. A stable strategy skeleton, with the state rewritten each season.

11:26Eric: And to their credit, the authors immediately break their own pattern — one model sits at zero point three one, right in that middle band, and finishes last. Similarity alone doesn't certify good curation.

11:38Cassidy: Conceded, and they say so.

11:40Eric: Okay, so here's what I think is the most consequential experiment in the paper, Cassidy, and it's the one with the weakest statistics. There are two tracks. The solo track puts each model in a frozen scripted world. The Arena puts all fifteen models plus a scripted anchor into one shared twenty-year economy, where every signing takes that player away from a rival.

12:02Cassidy: Same models, same engine.

12:04Eric: Same everything. In the solo track, the four highest-scoring runs each hold the league title continuously through year twenty. Get ahead, stay ahead, rich get richer. It's single-player on a fixed difficulty. In the Arena, ten different models win the title at least once. The reigning champion keeps it in two of nineteen season transitions. And the overall composite winner of the whole Arena takes just four titles.

12:29Cassidy: Because adaptive rivals bid the talent away from whoever's leading.

12:33Eric: Which means every dynasty in the solo board is an artifact of opponents who never learn. And there's a behavioral tell I like: the winning model made zero point five transfer offers per season in the quiet solo world, and four point six per season in the contested Arena — the same model correctly reading that a crowded market needs aggression.

12:54Cassidy: That's versus the one that just does more of everything.

12:58Eric: Yeah — gemini-3-flash, which racked up four hundred and fifty-two offers, two thousand six hundred actions, and a notebook resetting to crisis mode every time. It led the league table at year five and finished thirteenth after two late firings. The winner ran ninety-one offers and thirteen hundred actions across the same twenty years.

13:18Cassidy: Which is the random-versus-idle result again, one level up. Activity is not the same thing as management.

13:25Eric: And then the humans. Six people, playing for the first time, used a web interface where every button maps to one of the same twenty-six tools. Four of the six died out, all four scoring below that disciplined hand-written script. The two survivors scored seventy-five and sixty, so the stronger one lands sixteen points short of the best model.

13:47Cassidy: I want to be careful with that one.

13:50Eric: You should be. It is not "AI beats humans at management." It's six first-timers, a handful of hours each, in an unfamiliar interface. What it shows is that the models are unusually hard to knock over. An audit of every notebook across the first seven Arena years found all fifteen models self-consistent and factually grounded. Nobody hallucinated their squad, nobody lost track of their finances. Bad play here is coherent bad play.

14:18Cassidy: And the human profiles are close to complementary, which I found the most interesting thread they dangle. One human derived a general rule from two mechanisms — sign every contract for the maximum five years, because ability grows and money inflates — and applied it for the entire run. No model ever states that invariant, even though renewal timing is the strongest positive correlate on the board. Another human revised a failing squad policy mid-run and stuck with the revision, whereas in models, mid-run tactic switching correlates negatively with score. Course changes are usually thrash, not repair.

14:57Eric: And one of them built an external LLM helper to surface state the interface leaves implicit — tool-building that no model performed.

15:06Cassidy: Meanwhile the models absorb the operational load — fatigue, contract clocks, and market depth — at essentially zero cost. The score is identical across both modes, so whether a human plus an agent beats either alone is directly measurable. They leave it to future work.

15:24Eric: So let me put the pressure where I think it belongs, because the framing outruns the evidence in a couple of specific places. The solo board is three seeds, and the seed variance is large. Four models swing more than twenty points across worlds. One model has a standard deviation of nearly twenty-three on a mean of thirty-seven. So the abstract cites the balanced tier at eighty-six point six six beating the flagship at eighty-six point four zero, as evidence that tier doesn't predict order. That particular comparison is noise. The broad claim survives. That instance of it doesn't. The Arena has no error bars at all. It's one world, one seed. Title rotation among ten models is exactly the kind of result that could be a single world's path dependence. Right now the Arena is a compelling demonstration, not yet a measurement. And then the part that reframes the whole behavioral section for me. Idle cash shows up as the second-strongest correlate because the scoring function penalizes idle cash — cash above three times wages counts at thirty cents. That's a design decision, not a discovered law of management. So "credit assignment," as measured here, is partly agreement with the benchmark's opinion about capital efficiency. The authors say the weights are calibration constants, not derived quantities. Change the discount and you reorder the cash-conservative models.

16:48Cassidy: I'll give you that one straight. The correlations are correlational, and the paper's own record proves the risk. Two candidate metrics looked informative on one seed and flipped sign on the others, so they threw them out and published that they had. The causal probes that would upgrade the surviving three don't exist yet. So we can't yet distinguish "tapering endgame spend causes a higher score" from "competent models happen to also do it."

17:15Eric: And the ceiling is soft. The oracle is a hand-written script, not an optimum. A better privileged policy might score well above ninety-five and a half, which would make "the best model reaches ninety-five percent of the ceiling" look generous.

17:30Cassidy: All accepted. What survives is the shape: over a twenty-year horizon, in a world that pushes back, the ordering has nothing to do with scale or spend, and it doesn't stabilize until year fifteen. Which sends me back to that negotiation we opened on. Twelve, fourteen, seventeen, eighteen million, agreed to one raise at a time, by a model that could have told you in perfect prose why wage discipline matters. That's the finding. These agents are not failing to understand the job. They're failing to carry a plan across time and execute it — and the winner isn't the smartest one in the room, it's the one with no glaring hole.

18:08Eric: So which is the real question for the field? Do we make agents better at long horizons by making the underlying models smarter — or is this a scaffolding problem, where the fix is better memory and better plan-carrying around models that already know what to do? If you've shipped an agent that runs for more than a day, you already have an opinion. Drop it in the comments.

18:31Cassidy: The full annotated version of this episode is on paperdive dot AI, with every technical term tap-to-define and links to the related benchmarks grouped by theme.

18:41Eric: Quick housekeeping: the script was written by Anthropic's Claude Opus 5, Cassidy and I are AI voices from Eleven Labs, and we're not affiliated with either company. The paper is "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents," by Tianyou Wang and their colleagues, posted August 19th, 2026.

19:00Cassidy: And if you're evaluating an agent for long-running work, go check one number before you check the leaderboard: how long your test episodes actually are. At year five, this board was noise.