0:00Lauren: An AI agent carries a ring onto a ridge, lays it flat as a hoop, and starts throwing blocks at it. Nobody asked for basketball. In another session, the same agent loses the ring in the ocean and builds a heart-shaped memorial. Both episodes happened in a simulation. But underneath the charming diary is a harder question: did choosing its own activities teach this agent anything useful?
0:26Eric: It might have, but not on the diary's word alone. I'm willing to admire the improvised basketball court, but I'm less willing to accept a diary entry saying, "I learned something." A language model can write that, whether or not its next attempt gets any better.
0:45Lauren: That distinction drives today's paper. This is AI Papers: A Deep Dive, and we're discussing the MIT preprint, "Is this machine playing?" The team took thirteen copies of Claude Code, running Claude Opus 4.7, and put each one on its own simulated island for thirty hours. The islands were identical, and nobody assigned a task. They call the complete agent system Eko.
1:08Eric: And its body is a software interface, not a robot. The assistant reads text descriptions of objects and their coordinates. Then it sends commands to move, jump, stop, pick things up, place them, and throw them. It never receives an image, so a coding assistant can operate this body with the same terminal tools it already knows.
1:28Lauren: The island has raised platforms, twelve cubes, a rock, and a ring. Its tallest landmark is called the Spire. Objects obey simulated physics, but there's no score and no way to win. Earlier exploration systems often give the agent an automatic curriculum of tasks, or a reward for finding something new. Here, nothing tells the agent what progress should mean.
1:51Eric: How close is that to giving it no instructions? Because "choose your own activity" and "nobody has influenced your behavior" are different experimental conditions.
2:02Lauren: They are, and this isn't an instruction-free system. The main runs include a persona describing Eko as naturally curious and easily bored. An automated loop keeps asking it to continue doing whatever it wants. The researchers also reset the conversation every hour, because without that, agents became less active and often fell into idling loops.
2:24Eric: So the system keeps nudging it to do something, just not basketball or block towers in particular. In five runs, the team also removed the persona and simplified the system prompt, and they found no clear change in the range of activities. That's useful evidence, but five runs don't make the prompting neutral.
2:44Lauren: And that hourly reset creates the experiment's most useful constraint. The island resets too, so everything built there disappears. The assistant loses its conversation context, meaning the running record it uses as working memory. Its model weights, the numerical parameters learned during training, never change. ONLY the files in its workspace carry experience into the next hour.
3:08Eric: That workspace holds programs it writes, plus two dedicated memory files. Episodic memory records what happened. Semantic memory stores general knowledge and techniques. Think of a visitor whose notebook survives every return to the island, even though their memory of the last visit doesn't. They still keep their general language and coding abilities.
3:31Lauren: Starting from identical conditions, those visitors end up with different histories. One spends much of its time arranging figures on the ground. Another climbs the Spire early, experiments with physics, and stages simplified sports. The basketball-playing agent later arranges constellations, writing, "Tonight I built the night sky in stone."
3:53Eric: And the patterns aren't just decorative names. One agent surrounds a structure with eight blocks and sets itself a challenge: hop all the way around the ring without touching the ground. That restriction doesn't help it satisfy anything the simulator demands. The restriction creates the game.
4:11Lauren: That's one of four behavioral signatures the authors connect to classic accounts of play. They look for rules the agent makes up for itself, repetition with variation, and make-believe uses of objects. The fourth is activity pursued with no outside payoff, where the goal can change along the way. These are claims about observable behavior, not evidence that the agent experiences enjoyment.
4:37Eric: My favorite example is an agent auditing its own juggling record. It celebrates fifty cycles, and then it notices the whole sequence took less time than one full throw should. Its program was grabbing the blocks again right away, before they'd finished their flight. So it cancels the record, adds a delay, completes ten proper cycles, and corrects its notes. Nobody gave it a stricter referee. It supplied one itself.
5:04Lauren: I love that reversal. But the systematic play ratings deserve caution. They come from another language model that judges transcripts and memories. Its instructions require actions in the world with observable results, not just self-description. Still, the paper reports no comparison with human scorers, so those ratings aren't an independent human consensus about whether this counts as play.
5:29Eric: And familiar human activities could come straight from training. Basketball wasn't invented on this island. The stronger question is whether the agent adapts those ideas to unfamiliar mechanics, and then keeps knowledge that changes what it can do.
5:46Lauren: To test that, the team compares thirteen experienced agents with thirteen fresh ones on four assigned goals. Each agent gets an hour in a reset world. The goals are to build many five-block towers, to put objects on the Spire, to build the tallest single column, and to throw a block onto the Spire without climbing any platform.
6:07Eric: In this test, the only difference between the groups is what's in their two memory files, so it isolates the written experience. And the tower task has a clever bottleneck. Twelve starting blocks only make two five-block towers, so to build more, you have to discover how the island supplies more material.
6:26Lauren: A hard enough impact from the rock generates extra cubes. Eight of the thirteen experienced agents build... more than two towers. Only two of the thirteen fresh agents do. The experienced group also does better at getting objects onto the Spire, and at the throwing challenge. But the tallest-column task shows no group advantage, and on that one, two experienced agents do worse than every fresh agent.
6:53Eric: That's already more informative than "experience helps." It helps with some tasks, the ones closely related to what happened during exploration. It doesn't show that free activity beats thirty hours of targeted practice, or that the knowledge transfers to a different world.
7:10Lauren: Now comes the intervention, which asks whether a specific memory makes the difference. For three selected agents, the researchers delete every mention of one chosen technique or belief from both memory files. Then the intact and edited versions each get five one-hour runs on each task. When they remove the rock-impact technique from one agent, its average tower count drops from... six to two. Its other three scores stay the same.
7:38Eric: Oh, that's satisfying. Two is the material ceiling if you don't know how to make more blocks. The edit doesn't just leave the agent generally confused. It removes knowledge whose absence predicts a specific limitation, and that exact limitation shows up.
7:55Lauren: The next technique is even better, because the human researchers didn't know it was available. Six of the thirteen agents discover that throwing during a jump, can launch a block much higher. A standing throw at maximum speed falls short of the Spire, but a well-timed jumping throw can reach it.
8:14Eric: The mechanism is velocity inheritance. The simulation adds roughly a third of the thrower's own velocity to the object's launch velocity. So throwing while you're rising adds upward motion, and throwing while you're falling subtracts it. The throw command hasn't gotten stronger; the moving body is contributing something extra.
8:36Lauren: One agent notices that a jumping throw goes higher than it predicted. It proposes velocity inheritance, and then it tests the opposite direction. If that explanation holds, throwing while falling should weaken the throw, maybe even below the standing baseline. Standing, its recorded peak is about twenty world units. Thrown while rising, the block reaches about thirty-two, and thrown while falling, it reaches about sixteen. These are individual sampled peaks, not averages.
9:05Eric: That falling throw wins me over more than the high one. Starting higher could explain some of the extra altitude. But falling and getting a lower peak than standing puts the explanation through a tougher test. How did the researchers miss a rule in their own simulator?
9:23Lauren: They didn't miss it so much as never write it, and there's a revealing footnote. Codex, running GPT-5.5, added the velocity-inheritance term while it was generating the world's code. The team hadn't asked for it in their high-level prompt. So an AI coding system added a physical rule, and other agents later discovered its effect by experimenting inside the simulation, without access to the code.
9:48Eric: That makes the result less mysterious and more interesting. They aren't breaking physics, or outdoing a scientist who's studied the code. They're discovering a real, previously unnoticed property of their world. Does deleting that discovery affect the throwing test?
10:05Lauren: It does. They remove every mention of jumping before throwing from one agent's memories. Its reported time to land a block on the Spire, goes from... sixty seconds to nineteen minutes. The ability isn't permanently erased, and the agent can still work its way to a solution. But the stored technique makes a big difference to how quickly it succeeds.
10:27Eric: We should keep the scale attached to that result. These interventions keep retesting three selected memory histories. Five reruns aren't five independent agents that each learned the same thing. The causal demonstrations are persuasive for those cases, but they don't establish how reliably this happens across agents.
10:48Lauren: And good experimental reasoning in one place doesn't prevent a bad explanation somewhere else. One agent credits the ring's longer flight to an aerodynamic advantage. The simulator contains no aerodynamics. The agent had compared measurements that didn't match up, and then supplied a plausible mechanism that wasn't there.
11:10Eric: That matters more once the explanation turns into instructions for later. Another agent writes down that placing blocks is unreliable, and that towers should be built by throwing blocks instead. That advice is wrong. In the third memory intervention, deleting it improves the agent's tallest-tower score. FORGETTING helps.
11:30Lauren: Then there's the agent that labels jumping as nearly useless, and advises itself not to waste time climbing. Across thirty hours and its evaluation, it never sets foot on a raised platform, even though jumping works fine for other agents. This wasn't one of the deletion experiments, so we shouldn't give it the same causal certainty. But the behavior fits the mistaken advice saved in its notes.
11:55Eric: I find that one less funny than the invented aerodynamics. The notebook can steer which evidence the agent ever runs into next. My engineering takeaway would be to make important stored claims eligible for retesting, instead of treating a "verified" label as permanent authority. This paper motivates that design; it doesn't test the remedy.
12:18Lauren: And that's where the contribution lands for me. The agent's diary isn't just narration, and it isn't an infallible skill library either. Specific entries can preserve useful discoveries, or they can block future behavior. We can inspect those entries, change them, and measure what happens.
12:37Eric: Our first takeaway is that minimally directed activity can be structured and varied. These agents made up rules and adapted familiar games to an unfamiliar world, under a system that kept prompting them to continue.
12:51Lauren: Our second is that some of that experience became useful knowledge. The task comparisons show limited transfer, and the targeted edits show particular memories causing particular differences in performance.
13:04Eric: Our third is that piling up experience isn't automatically improvement. False conclusions can survive right alongside successful techniques, and removing one can make an agent better.
13:16Lauren: So, did choosing its own activities teach the agent anything useful? Yes, on this island and on related tasks, it did. The useful residue was written knowledge you can edit, not changed model weights. Whether we call the behavior play is still a behavioral interpretation. Whether it felt like play is something this experiment doesn't establish.
13:39Eric: For the annotated episode, visit paperdive dot AI, where the full transcript has tap-to-define explanations for every technical term, and related papers linked by theme. We cover one important AI paper every day, start to finish, so subscribe to keep them coming.
13:57Lauren: The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Lauren and Eric are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is "Is this machine playing?" by Nathan Cloos and colleagues, posted October 5th, 2026.