Two Random Networks Teach Each Other To Predict Real Data

0:00Paige: Imagine a hypothetical tutor who's rewarded whenever a student fails to predict the next number. The tutor could print fresh random numbers forever. The exercises would stay difficult, but there'd be no hidden rule to discover. Now make the tutor and student neural networks. Could they invent useful lessons without any examples from the world?

0:19Eric: Yes, they can, provided the tutor is rewarded by something better than difficulty. A learner trained on the outputs of self-generated programs gets better at predicting real text, image pixels, and speech samples. The interesting question is what transfers between those very different settings.

0:35Paige: This is AI Papers: A Deep Dive. We're discussing “Self-Play Pretraining with Zero Data,” a preprint by Aditya Cowsik and colleagues, from a collaboration including researchers at Stanford and Tel Aviv University.

0:47Eric: We should pin down that title. “Zero data” doesn't mean the networks learn without examples. It means they generate their own training examples. And neither network starts as a pretrained model that already carries knowledge from human text.

1:00Paige: Both start with random weights, and no natural data enters their weight updates. There's an important qualification, though. The researchers check how well models predict web text and DNA, and they use those scores to choose training settings. So natural data influences which models get picked, even though it doesn't supply the training examples.

1:19Eric: That makes this a controlled experiment about what pretraining buys you. Some of it must be information about the world. But perhaps some of it is practice at recognizing structure, and that part doesn't have to come from the world.

1:33Paige: Mechanically, pretraining here means predicting the next byte. The model assigns probabilities to the possible continuations, then adjusts its weights to give the actual continuation more probability. A byte has 256 possible values. The researchers encode every evaluation domain as byte sequences, so one predictor can handle all of them.

1:52Eric: And they score prediction error in bits per byte. Lower is better, because it means the model is less surprised by what comes next. You can also think of it as a compression score. But predicting image bytes isn't recognizing objects, and predicting speech samples isn't transcribing speech. These are tests of sequence prediction.

2:11Paige: To manufacture training sequences, the team uses two transformers with the same architecture but separate weights. The first one, the generator, writes short programs. A small virtual machine runs those programs, and the programs can read random input bytes. Whatever they print becomes training data for the second network, the learner. The learner sees the outputs, not the programs.

2:33Eric: Why programs, rather than having the generator write sequences directly? The programming language can, in principle, express any computable process. That gives the search a broad space, without specifying that the output should resemble English or pictures. In practice, though, execution has strict memory and time limits, so theoretical universality doesn't mean unlimited computation.

2:55Paige: The machine is also designed so that every generated instruction string can run. Unmatched loop brackets don't cause syntax errors, and execution stops once it uses up its budget. The machine also has shortcut instructions, designed by humans, that make common operations easier to express. The comparison baselines, which sample programs without learning, get those same shortcuts, so the shortcuts can't explain the advantage of self-play.

3:21Eric: But expressiveness doesn't tell the generator where to search. Most possible programs won't make useful lessons. And our hypothetical tutor's problem still applies: if you reward prediction difficulty, you can end up favoring random output. What replaces that reward?

3:35Paige: What replaces it is a question about the learner rather than about difficulty: does a candidate exercise engage the directions the learner has already been learning in? To measure that, the authors compare the exercise's loss gradient with how the learner's weights have moved, over a growing stretch of training history. A gradient describes how changing each weight would change prediction error. So, the comparison measures how strongly that exercise connects to the path the learner is already on.

4:04Eric: There's a technical distinction here. The score uses the training algorithm's per-weight scaling, and it takes the absolute value of the alignment, so either sign counts. That means it isn't literally asking whether the exercise pushes the learner forward. It's closer to asking whether the exercise is sensitive, along the directions that recent learning has established.

4:24Paige: The authors' intuition is that exercises the learner has mastered produce little signal. Unlearnable noise produces changes that don't line up with sustained progress. Useful new structure should fall between those two. That's a heuristic, not a guarantee. They chose this reward because it worked in experiments, and in small-scale tests, shuffling the rewards between programs makes transfer worse.

4:47Eric: The generator also mixes fresh programs with mutations of promising ones, and it replays older programs. So this isn't just two networks exchanging guesses. There's machinery for exploration and for memory. How do we know that adapting the curriculum adds value?

5:02Paige: We know the adaptation adds value from two controls that keep everything except the adaptation. The first samples programs from the same language without learning which ones to favor. It weights shorter programs more heavily, but it never adapts to the learner. Earlier research had already explored training on random computations like this. Here, as compute grows, the adaptive system improves substantially faster than that fixed program source. So access to an expressive language isn't enough.

5:29Eric: The second control is more specialized. It generates sequences from probabilistic context-free grammars, which produce hierarchical, language-shaped structure. Training on those grammars is stronger than self-play on text and code. But self-play is substantially better on images, music, audio, and speech. So the result is broader transfer, not a win on every domain.

5:50Paige: Across the evaluation suite, giving self-play more training compute produces predictable drops in next-byte error. The reported curves take the best tested combinations of model size, training length, and ensembles, which average several models' predictions. So they aren't simply following one model as it trains longer. All the models have fewer than 25 million parameters, and they see about four thousand bytes of context at a time.

6:14Eric: The researchers fit power laws with a floor. That means each time you multiply compute, you get a roughly consistent fractional cut in the error that remains above that fitted floor. The regularity matters, because it shows useful transfer doesn't appear only in one lucky run or one domain.

6:31Paige: The authors compare these improvement rates with published rates for training directly on natural data, and they're broadly comparable. But this isn't a head-to-head contest. The studies differ in scale and setup, and some of the comparison values are mathematically derived. A steep improvement curve also doesn't guarantee a good destination. Self-play can't supply missing world-specific information.

6:52Eric: The programs themselves give us a concrete view of what the search finds. By training round 512, the saved generator outputs included Fibonacci-like sequences, where each new value is the sum of the previous two, wrapping around so it fits in a byte. The generator also found geometric, quadratic, and cubic sequences.

7:09Paige: Meanwhile, sampling 164 million programs from the fixed random source, produced no matches for any of those four families. That's evidence that the search finds these structures efficiently, not evidence that random search could never find them. And it doesn't establish that Fibonacci-like lessons caused the improvements on natural data. The authors name that causal question as future work.

7:30Eric: We can test something more directly than spotting attractive programs: can the learner pick up a new rule from examples in its input? That's in-context learning. Its weights stay fixed, and the examples themselves provide the information it needs to answer the next query.

7:45Paige: With enough examples, self-play learners approach perfect accuracy on three tasks: reversing a string, stack operations, and retrieving values from a dictionary printed in the input. That's measured on the predictions that get scored. Pretraining on the fixed program source produces little effective in-context learning. Grammar pretraining is strong on dictionary retrieval, but it transfers weakly to the other tasks.

8:08Eric: The addition experiment gives a more detailed picture. The task adds two byte values, wrapping around at 256. Across sampled trials, as demonstrations pile up, the model first favors common output bytes. Then it tends to copy bytes from the context, which is usually wrong, and after that its predictions become less confident.

8:26Paige: After roughly four examples, it starts getting the lower four bits of the answer right. Around eight examples, it starts getting the upper four bits right too, and then its confidence rises. Those are patterns in its answers and probability distributions, not a report of its thoughts. But they show useful rule inference without any extra weight updates.

8:45Eric: That seems stronger than saying it learned a few recurring byte frequencies. It's applying relationships that were supplied in the input. Still, these are deliberately mechanical tasks. We shouldn't turn success at byte addition and stack operations, into a claim that it can learn any unfamiliar task from a prompt.

9:02Paige: And some of the natural-data improvements may come from quite simple regularities. The DNA benchmark uses only eight symbols out of those 256 possible byte values. Just recognizing that restricted alphabet can cut prediction error substantially, without learning any richer biological structure. That's a plausible contributor, not something the experiments isolate. Better DNA prediction doesn't automatically mean biological understanding.

9:27Eric: The authors organize these findings around two resources. One is contingent information: facts specific to a particular world or dataset. The other is transferable predictive structure, like copying and recognizing repetition. Their mathematical model separates those two, but the experiments don't measure how much of ordinary pretraining belongs to each.

9:47Paige: They also ask whether self-play helps once real data becomes available. They take a self-play model of roughly 24 million parameters, and use it as the starting point for normal training. On the ESC-50 audio benchmark, that warm start reaches their convergence criterion after about 320 million training tokens. Starting from random weights takes about 496 million. The advantage narrows toward the end of training.

10:09Eric: Those are tokens from the later training, not a saving in total compute. The comparison leaves out the cost of producing the self-play starting point. The authors argue that one starting point can be reused across domains, which is reasonable, but the experiment doesn't settle the overall economics. What it shows is faster learning afterward.

10:28Paige: Nor does it settle whether the approach keeps working at much larger scales. Absolute prediction quality is still far below a practical model. The authors frame starting from scratch as a scientific control, not necessarily the best engineering recipe. What they've shown is narrower: generated computational structure can improve prediction beyond what it was trained on.

10:47Eric: My first takeaway is about the teacher. Difficulty alone is a poor goal for a curriculum. This particular learning-progress signal makes adaptive program generation more useful, than continuing to sample from a fixed program distribution.

11:00Paige: My second is about transfer. The learner gets better at predicting diverse real datasets, and it gains in-context learning on tasks it never trained on. That supports the idea that some benefits of pretraining are reusable sequence skills, rather than knowledge tied to one domain.

11:16Eric: The third is the boundary: reusable skills aren't world knowledge. So can two models invent useful lessons without examples from the world? Within this controlled setup, yes. They manufacture practice that transfers, but they don't manufacture the facts that only experience with the world can supply.

11:36Paige: You can find the annotated version of this episode at paperdive dot AI. It has the full transcript, with every technical term tap-to-define and related papers linked by theme. If you want every major AI paper taken apart like this, daily, that's what this channel does, so subscribe and you'll get them.

11:55Eric: The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Paige and I are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is “Self-Play Pretraining with Zero Data,” by Aditya Cowsik and colleagues, posted September 24th, 2026.