Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%

0:00Cassidy: Here's a question people have been arguing about for three years without a real answer: what fraction of new scientific papers are written with help from a chatbot? A team just answered it for biomedicine, across about 1.2 million full-text papers. By December of last year, roughly nine in ten showed the linguistic signs of LLM-assisted writing.

0:20Finn: Nine in ten. And the number that should stop you is the one it replaced. Studies on comparable data have reported everything from two percent to fifty-seven percent, with the most-quoted ones landing around fifteen. This paper's claim is that those earlier studies weren't wrong measurements. They were measurements of a different quantity, and structurally they could only come out low.

0:42Cassidy: So the thing to hold onto here is the counting move — how you measure something you can never detect in any individual case, and why the honest version of that measurement had been sitting in plain sight as a floor that everyone was quoting as an answer.

0:56Finn: Well, and it lands in the middle of a policy vacuum. Journals, funders and universities are writing AI disclosure rules right now. A rule designed for a fringe behavior and a rule designed for the default condition of scientific prose are not the same object. And you can't write either one if your best estimate spans two percent to fifty-seven.

1:16Cassidy: So, start with the obvious approach, which is a detector. You paste the text in, you get a verdict. That dream has mostly collapsed. Commercial detectors have been tested repeatedly and come out unreliable, and they fail in the worst direction — flagging real human writing, disproportionately from people whose first language isn't English. The authors mention dropping one detector from their comparison table because it was labeling papers from before 2022 as AI-written.

1:43Finn: Right, so the field pivots, and the pivot is the interesting part. You stop doing forensics and you start doing epidemiology. Public health people estimate what share of a city has a virus from wastewater samples, with no ability to point at a house. It's the same posture here: you give up on "which paper" to get "how many."

2:03Cassidy: Mm-hm. And the standard way to do that is word counting. Chatbots overuse a small, specific set of unremarkable words.

2:10Finn: So the recipe writes itself, right? You take those words, you measure how much more often they show up now than they used to, and you report the excess as the usage rate — fifteen percent excess, fifteen percent of papers. No classifier, no black box, just arithmetic on a corpus.

2:27Cassidy: And that last step is the mistake. The excess frequency is not the usage rate. It's a floor — a guaranteed minimum — and as it turns out, a very loose one. What makes that beat land is who's saying it. The senior author on this paper, Dmitry Kobak, is the person who pioneered the excess-vocabulary method in the first place, in a 2025 paper in Science Advances. This new paper is the same lineage telling you their own earlier number was too low.

2:54Finn: Okay, so before the fix — say back the failure in one line. Why does counting the excess undercount?

3:00Cassidy: Because it tells you how often the fingerprint shows up. It doesn't tell you how many papers have a fingerprint on them.

3:08Finn: Right. Those are different questions.

3:10Cassidy: So, the marker words first, because everything rests on them. They inherit a list of roughly 380 words from that earlier work, and the list is deliberately boring — not technical terms, not topic words, but words like "these," and "potential," and "invaluable," words you could use in a paper about mitochondria or a paper about hospital scheduling. That means their frequency tracks how someone writes, not what they're writing about. That insulation matters. If you used topic words, a shift in what biomedicine happens to study this year would masquerade as chatbot usage.

3:45Finn: And the baseline? Because "more often than they used to" needs a "used to."

3:50Cassidy: So this is excess mortality, exactly. During the pandemic nobody adjudicated individual death certificates. You take the trend from the previous several years, project what a normal year would have produced, and count the overshoot. Here, you take each word's frequency from 2018 through 2022, fit a line, and extend it through 2025. That dashed line is your counterfactual — the world where ChatGPT never shipped.

4:16Finn: And everything downstream rides on that dashed line being right. Which I want to come back to, because the way the math works, an error in that line gets multiplied on its way to the answer — roughly threefold. That's the soft spot in the whole paper, and it's not small.

4:33Cassidy: It isn't, and we'll get there. But watch the worked example first, because it's the best teaching moment in the paper, and it uses the most boring word in English. The word is "these." Just "these." On screen, that's the curve climbing gently from 2018. The dashed projection carries it to about thirty-three percent of abstracts by December 2025 — and then the real line comes in at fifty percent.

4:57Finn: Half of all abstracts. On "these."

5:00Cassidy: Half. Seventeen points of excess, on a word nobody has ever noticed in their life. Now here's the arithmetic that changes everything. The old method reports seventeen percent and stops. But think about who was available to change. Thirty-three percent of papers were predicted to contain the word already. Those papers can't tell you anything — they were going to say "these" regardless. The room available for change is the other sixty-seven percent, the papers that should have been silent. Seventeen points out of sixty-seven points of room is a quarter. So that single word, on its own, says at least twenty-five percent of papers had LLM help.

5:37Finn: From one word. And that's a floor, still.

5:40Cassidy: That's still a floor. If you want every major AI paper taken apart like this, daily, subscribing is the way to get them.

5:47Finn: So, why is it still a floor? Because "these" is a single fine thread in the water. Plenty of chatbot-edited papers swim right past it — the model rewrote your discussion section and never happened to reach for that particular word. To turn a floor into an estimate, you'd need a test that fires on essentially every LLM-touched paper.

6:07Cassidy: Which is the technical core, and it pays off in something nobody else in this literature has done — a version of the method you can check against a known answer. So the move is: stop tracking one word. Pool hundreds of them, and ask a single yes-or-no question of each paper. Does it contain any of these marker words, anywhere?

6:26Finn: Widen the net.

6:27Cassidy: Widen the net. And now the thread becomes a mesh. A language model that touched a paper anywhere almost certainly deposited at least one of its several hundred favorite words. So the detection rate on LLM-assisted papers climbs toward one hundred percent — which is exactly the assumption the old floor was quietly making. The floor and the true estimate collapse into each other.

6:49Finn: Except a net that catches everything tells you nothing about the water.

6:53Cassidy: And that's the tension that makes this non-trivial. Widen far enough and human papers all trip it too. Everybody writes "these." Then your baseline is ninety-nine point nine percent, your headroom is basically zero. You're dividing a tiny number by a tiny number, and the estimate explodes into noise. So there's a mesh size in the middle: wide enough that LLM-touched papers essentially always get caught, narrow enough that a real share of human papers still slips through. That gap is where the information lives.

7:25Finn: And how do they pick it?

7:27Cassidy: They sweep it across nineteen settings for how rare a word has to be to make the list. They throw out the settings where the estimate goes unstable, and they take the largest stable answer. Hold that last step in mind, Finn's going to have something to say about it.

7:43Finn: Oh, I am.

7:44Cassidy: So now run the same silence arithmetic on the pooled set, and this is the moment the whole paper turns. Across the corpus, the share of papers containing at least one marker word went from about eighty-three percent in 2022 to about ninety-five percent in 2025.

8:00Finn: Twelve points, which sounds like nothing.

8:03Cassidy: It sounds like nothing. But picture a concert hall that was already eighty-three percent full. There were only seventeen empty seats in a hundred to begin with, and twelve of them filled. You didn't sell twelve percent more tickets. You sold most of the tickets that were left. Same twelve points, completely different story — and that's why the estimate comes out near seventy percent instead of twelve.

8:27Finn: Huh. So the tiny-looking number is tiny because there was almost no room left to move.

8:33Cassidy: Right. And with that, here's the trajectory. About a fifth of full papers in 2023. About half in 2024. Seventy-seven percent across 2025. And for December of 2025 alone — eighty-nine percent.

8:45Finn: Okay, and this is where I'd normally ask how they know the estimator isn't just producing a big number because it's built to produce big numbers. And they answer it. They simulate a hundred thousand documents a year with five hundred marker words and a fraction of LLM-written papers that they choose. So they know the true answer, and then they run the entire pipeline blind at it.

9:07Cassidy: And the prediction, if the method is sound, is that it recovers the planted fraction anywhere across the range.

9:14Finn: It does, within two percentage points, from zero percent all the way to one hundred. Run the old excess-frequency measure on the same synthetic corpus and it undershoots badly, exactly as predicted. That figure is the only place in this literature where one of these estimators gets checked against ground truth, and it's the reason to take the number seriously at all.

9:36Cassidy: So here's the checkpoint so far. Detectors can't do individuals. Word-frequency gaps only give a floor. And pooling marker words plus the headroom arithmetic turns that floor into an estimate that survives a test with a known answer. Which is when the method starts paying dividends, because it works on any slice of text you hand it.

9:55Finn: And this is my favorite part, because the internal structure of the result argues against the scary reading of the headline. Papers have a standard skeleton — introduction, methods, results, discussion. Methods is a protocol list — reagents, doses, statistical tests — written in near-telegraphic prose. Discussion is the essay: what it means, what the limitations are, why anyone should care.

10:18Cassidy: Two stylistic opposite ends of the same document.

10:22Finn: Right, and they compared equal-length chunks, because a longer section gives a model more chances to leave a fingerprint. Length-controlled for December 2025, Discussion sits at sixty-eight percent, Methods at thirty-two. Discussion is twice as likely to be touched. And whole Methods sections score much higher than equal-length crops of Methods, which tells you the usage inside Methods is patchy — concentrated in a couple of prose-heavy passages rather than spread through the protocol.

10:51Cassidy: Which is the signature of polishing.

10:53Finn: It's the signature of polishing and translating. And the geography says the same thing. For full papers in 2025, South Korea sits at eighty-five percent, China at eighty-two, the United States at thirty-nine, and the United Kingdom at twenty-eight. Aggregate the majority-native-English countries and you get thirty-seven percent. Everyone else comes in at seventy-two. I'll flag that the country analysis is the weakest-supported part of the paper. It uses one fixed word list rather than tuning per country, so some of those numbers are explicitly lower bounds — and the authors say so.

11:28Cassidy: Still, the direction is hard to argue with, and there's a reading of it that isn't about misconduct at all. English is the working language of international science. If your first language isn't English, you've been paying an invisible tax for decades — good science, prose that gets called unclear, sometimes rejected for it. Copy-editing services have charged money for that fix forever. A free universal editor showed up.

11:53Finn: And that sets up the number I think is the most under-sold thing in this whole paper, Cassidy. In 2022, marker-word frequency was ninety-two percent in native-English-speaking countries and seventy-nine percent everywhere else — a real, measurable gap in stylistic vocabulary between scientists in Manchester and scientists in Seoul.

12:14Cassidy: And by 2025?

12:14Finn: Ninety-six and ninety-five, so the gap is gone — in three years.

12:19Cassidy: That's a measurement of cultural change, not tool adoption.

12:23Finn: And the thing is, both readings of it are supported by the exact same number. A language barrier that stood for a century fell in three years, which is an equity win. And the world's scientific prose converged on one polished register in three years, which is homogenization. Take your pick, the data doesn't.

12:41Cassidy: Before the caveat, one thing has to be said plainly about what's being measured. This estimate covers LLM involvement anywhere in a paper. "Some help" includes a grammar pass on one paragraph. It says nothing about fabricated results, invented citations, or fake science.

12:57Finn: Which is my first objection, actually. The results section is careful about that. The discussion then wanders through hallucinated citations, paper mills, degraded thinking, and lost creativity — all cited from other people's work, none of it measured here. The eighty-nine will get quoted in service of that framing, not in service of the evidence.

13:18Cassidy: That's fair.

13:19Finn: And then there's the assumption. The whole estimate hangs on one dashed line: five yearly points, 2018 through 2022, extended three years forward. The estimator divides the excess by the headroom, and the headroom is small — marker-word presence was already around ninety percent before any of this. Work it through with their own corpus figures and a one-point error in the baseline moves the answer about three points. So a three-point drift in how humans write moves the estimate by roughly nine.

13:48Cassidy: And the mechanism that would cause that drift is the obvious one.

13:52Finn: People read this prose all day. There's published evidence that LLM style is bleeding into human speech, and the authors cite it themselves. There's also a counter-effect — people deliberately purging marker words from their drafts — and the authors say plainly that it's unclear which is stronger. The simulation doesn't rescue this, either. That synthetic corpus was built to obey the estimator's own premises, and human drift is absent from it by construction. They stress-tested the ruler. They didn't test the assumption about the object.

14:22Cassidy: No, they didn't, and nobody in this literature has. I'd add the guardrail question you were sitting on, too — taking the maximum over nineteen noisy settings selects for whichever one happened to read high.

14:33Finn: Nineteen thermometers, report the hottest. And the eighty-nine specifically is the thinnest slice of data in the paper — one month, the last month before the snapshot was taken, which contains whatever journals deposit fastest.

14:46Cassidy: And I'll concede the number on that. The defensible claim is about three-quarters across 2025, trending up toward nine in ten. The precision here is real, but the accuracy rests on a straight line, and that's a different kind of confidence.

15:00Finn: So where does that leave the headline?

15:02Cassidy: It's roughly where we started, but decodable now. That opening question — what fraction of new papers are written with a chatbot — turned out to be unanswerable in the form everyone was asking it, because you cannot check any single paper. What you can do is count who stopped being silent. Most of the papers that should have avoided these words didn't. Whether that's three-quarters or nine in ten, LLM-assisted writing is not a fringe behavior to be policed. It's the default condition of scientific prose, and it's the training data and the retrieval substrate for everything being built on top of it.

15:36Finn: So here's the question worth arguing about in the comments. That convergence number, ninety-two and seventy-nine becoming ninety-six and ninety-five — is that a language barrier falling, or the world's scientific writing flattening into one register? You almost certainly lean one way. Say which, and why.

15:54Cassidy: The full annotated version of this episode is on paperdive dot AI — every technical term tap-to-define, with links to the related papers grouped by theme.

16:03Finn: Quick housekeeping: the script was written by Anthropic's Claude Opus 5, Cassidy and I are AI voices from Eleven Labs, and we're not affiliated with either company. The paper is "Most biomedical publications show signs of LLM-assisted writing," by Lena Holzwarth and their colleagues, posted August 11th, 2026.

16:22Cassidy: If nine in ten new papers already sound like the model, what exactly is the next model learning from?