Can You Measure Research Taste If The AI Isn't Allowed To Code?

0:00Bella: In this study, an AI researcher isn't allowed to write experiment code. It can decide what to try, but a separate coding agent has to carry out the plan. Even so, the best model reaches expert-level results using roughly half the experimental compute. That raises a surprisingly slippery question: can we measure good research judgment separately from the ability to build things?

0:25Finn: We can measure research judgment separately from the ability to build things, but only partly, and it depends on what that measurement captures. Choosing a useful experiment matters. So does recognizing that a promising result is noise. Neither one necessarily means you've invented anything. This paper gives us reasons to keep those abilities separate.

0:47Bella: This is AI Papers: A Deep Dive. Today we're discussing TasteVal, a preprint by Oliver Jaffe and Dane Sherburn at P-Zero Research. Their target is experimental research taste: deciding which experiments to run, and interpreting what happens. Existing evaluations often mix that judgment with coding skill, or they score proposals by whether experts think they sound promising. TasteVal asks something different: do the chosen experiments deliver results?

1:16Finn: Delivering results is the whole bar, so the model doesn't get credit for writing a persuasive research proposal. Someone has to run it. The team splits the work into two roles: the Researcher chooses the experiments, and the Coder implements them. Every contestant, human or AI, directs the same Coder model, Opus 4.8. And the Coder isn't told which kind of researcher it's working for.

1:40Bella: The separation goes further than taking away the keyboard. The Coder reports measurements, but it can't suggest next steps or explain what the results mean. Two automated monitors enforce that boundary. They reject experiment descriptions that leave research decisions to the Coder, and they reject reports that sneak research advice back to the Researcher.

2:03Finn: I like that restriction. An assistant saying, “You should probably try this next,” would contaminate the very ability they're measuring. Everyone gets one H100 graphics processor. They get forty hours of active GPU time, plus a separate limit of a hundred and twenty hours of elapsed time. The model's own deliberation doesn't count against that experimental-compute budget, although it does cost money.

2:28Bella: Researchers can inspect the training data, which is what the experimental model learns from. They can also use validation data, which gives them feedback while they improve the model. Then a separate test set supplies the final grade, and the Researcher never sees those test scores. That split lets the team check whether an apparent improvement holds up, beyond the feedback that was used to choose it.

2:53Finn: So what does one of these problems look like? The appendix gives an illustrative task that wasn't used in the benchmark. You're building a training collection from raw web text in nine languages. You can keep five hundred million tokens, the small text units models process, out of a pool of about five billion. That pool mixes useful writing with spam, cookie notices, and boilerplate. You decide which documents make the cut, and in what order.

3:22Bella: The example Researcher's first move is deliberately plain: give each language an equal share, don't filter the text, and use default training settings. That sets a starting point before trying any clever selection methods. I love that example, because good taste can begin with resisting the urge to be clever.

3:41Finn: The human comparison wasn't casual, either. The team recruited twenty-four experts with relevant research experience, including people who'd recently worked at OpenAI and Google DeepMind. They paid them to prepare by doing literature reviews. At least two experts attempted each task, and the strongest human attempt on each one became the baseline. Models generally got six runs per task, and they weren't scored only on their best attempt.

4:08Bella: There was still some friction with the interface. Humans had more of their experiment descriptions rejected than recent models did, usually for leaving details unspecified. Rejections didn't use up GPU time, and rejection rates fell as the humans learned the interface. But they could still use up attention. Identical rules don't necessarily mean everyone's equally at home in the working environment.

4:34Finn: Then there's the central measurement, the compute multiplier. Suppose a human reaches a particular result after forty GPU hours, and a model reaches the same result after twenty. The model gets a multiplier of ... two. But two runs usually finish at different scores, so the comparison takes the weaker run's final score, and asks when the stronger run first reached it. It measures how efficiently you reach comparable quality, not simply who finishes with the highest score.

5:05Bella: And these are experiments run one after another on a single GPU. Halving the GPU time, isn't the same as halving the number of machines in a research cluster. The authors also compare models spanning several years of capability, and to do that, they chain comparisons through multiple reference points. So the reported multiplier is an aggregate estimate, not one stopwatch reading.

5:30Finn: With that machinery in place, Opus 5.5 gets a compute multiplier of about ... 2.3 relative to the expert baseline. That's the basis for saying it uses roughly half the experimental compute. But the ninety-five-percent confidence interval is wide: it runs from about 1.2 to 4.4. So the advantage is supported here, but its size is uncertain, especially because performance varies substantially across a small set of tasks.

5:56Bella: There's one more condition attached to that result. Opus 5.5 refused one task, so it was scored on seven of the eight. Recalculating everyone's comparison on those same seven barely changes its multiplier. The cost difference is striking, too. Its average run cost under three hundred dollars, versus about nine thousand for the human runs. Those totals include the supporting agents and the GPU time, and on the human side they include human pay.

6:24Finn: Under their accounting, that's roughly ... one-thirtieth of the cost. I didn't expect the gap to be that large. But the final answers aren't twice as good. They use a normalized performance scale, where a weak starting solution scores zero and the expert result scores one. Opus 5.5 scores 1.14. That's extra progress beyond the expert baseline, not a fourteen-percent improvement in accuracy across the board.

6:51Bella: The historical trends back up that distinction. Across the models tested, compute efficiency improved much faster after December 2025. The fitted doubling time drops from about fourteen months to ... three. Final performance on the normalized scale shows no statistically significant break in its trend. So the sharp recent change is in reaching comparable results with less experimental compute. Final answer quality didn't accelerate to match.

7:18Finn: That three-month estimate describes the models they observed, not a promise about future releases. Its confidence interval spans roughly two to five months, and only relatively few releases fall in that recent stretch. The authors' sensitivity checks support the bend, but extending that curve into the future is a separate assumption.

7:39Bella: Which raises the question of what those efficient experiments actually contained. The authors sort experiments into four kinds: adjusting an existing recipe, combining published components, structurally modifying a component, and inventing a component that isn't in the literature at all. Their audit covered five hundred and forty selected model submissions. It also covered about sixteen hundred more experiments from top-performing runs, including failures and rejected proposals. None were classified as INVENTED.

8:12Finn: That sounds devastating until you add the human result: the humans registered no inventions either. Most human submissions combined existing components. There's also a caveat about the labeling. Another AI model did the novelty labeling, using web search, and the paper doesn't report how well it agrees with human novelty judges. So this is evidence about what the audit found in this setting, not proof that models can't invent.

8:40Bella: Nor does it make the efficiency result disappear. Picking useful combinations can be productive research. The tension is that TasteVal shows strong optimization without demonstrating new scientific components. And the authors identify a more specific weakness in how one strong model handled evidence.

8:58Finn: Training involves randomness, controlled partly by a random seed, so running the same setup twice can produce different scores. In the transcripts they reviewed, the strongest humans measured that run-to-run variation in their first experiment. Then they discounted any single-run gain smaller than the variation they'd measured. Fable 5.1 rarely repeated an experiment with a different SEED. Take a hypothetical case where the score naturally wobbles by half a point. A third-of-a-point gain isn't persuasive evidence that your change helped. The humans checked the wobble before celebrating.

9:36Bella: That feels more revealing than just saying the model has less judgment. It points to a practice you could improve. Did the team test whether changing how the model works could encourage more of that discipline?

9:49Finn: Yes, they tested exactly that with Opus 5.0, and the result is suggestive. In a variant called Super-Think, the software wrapper wouldn't let the Researcher sit idle while an experiment ran. It had to keep going, until it had generated two hundred thousand tokens of deliberation and output, or until the experiment finished. In the transcripts they reviewed, the model searched the literature more widely, wrote down numerical predictions before results came in, and set explicit criteria for stopping unpromising experiments. The same model, given a different working loop, showed more deliberate research habits.

10:29Bella: Across six tasks, its estimated compute efficiency improved by about ... forty percent over the ordinary maximum-reasoning setup. That came at about fifty percent extra cost. But the uncertainty interval includes no gain at all. So I find the behavioral changes promising, but the efficiency improvement isn't statistically settled.

10:50Finn: And careful habits still don't solve the problem of deciding what success means. The ordinary tasks hand you a clean validation signal. To test how much the models depend on it, the authors made messier versions of four tasks. They took away the validation set, the validation scores, and the scoring description, but they told the Researcher that a hidden score existed. For both models tested, the estimated compute multipliers fell to roughly half their normal values.

11:19Bella: That drop deserves attention, but with only four tasks, neither model's drop was statistically significant. More importantly, the humans weren't rerun under those messy conditions. So you can't conclude that removing feedback makes models fall behind humans facing the same handicap. What you can say is that the model results suggest clean feedback contributes substantially to performance.

11:44Finn: How they responded to losing that feedback was interesting, too. Both models built their own evaluation sets in almost every run. The newer model often drew on public datasets. The older one mostly used portions of its training data, and those predicted final test scores poorly. They tried to rebuild the measuring instrument they'd lost.

12:05Bella: That's where the boundary becomes clear. TasteVal hands contestants a problem, and usually a useful progress signal. It doesn't measure the step before that: choosing which problems deserve a research program. Its eight tasks are also kept private to prevent training contamination, and that limits how much outsiders can inspect them. The contribution is a repeatable instrument for one part of research judgment, with real experimental outcomes behind it. That's useful, even though it isn't a complete test of scientific ability.

12:38Finn: My first takeaway is that the model advantage here is mainly compute efficiency and cost, not dramatically better final answers. My second is that efficient use of existing ideas counts as useful work, but it shouldn't be mistaken for demonstrated invention.

12:54Bella: And my takeaway is that experimental judgment depends on the feedback, and the working process around the researcher. So can we measure research taste separately from coding? We can measure a bounded experimental slice, and the strongest model tested performs well on it. Choosing the next useful experiment is becoming measurable, but choosing the next great research problem remains outside this test.

13:19Finn: You can find the annotated episode at paperdive dot AI: the full transcript with every technical term tap-to-define, and related papers linked by theme. It adds no new analysis. We break down a major AI paper every day, so if you subscribe, tomorrow's is in your feed.

13:36Bella: Here's our disclosure. The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Finn and I are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is "TasteVal," by Oliver Jaffe and Dane Sherburn, posted October 5th, 2026.