← All notes

The arms race

The humanizers lost

Every tool in this category sells one number. A Chicago Booth audit and the detectors' own benchmarks now say that number is going the wrong way, and it was always the wrong thing to buy.

August 24, 20264 min readMichael Lynn

Search for an AI humanizer and you get a page of companies reviewing each other. Every one of them leads with the same number. Ninety-seven percent past GPTZero. Ninety-four past Turnitin. Undetectable, guaranteed, this month.

That number is the whole product. It is what gets bought.

And it is falling.

What the benchmarks say

Start with the number the detector publishes about itself, and treat it as what it is. Pangram says its humanizer-trained model catches 93.66% of high-quality humanized text where GPTZero catches 34.53%. That is a vendor scoring its own exam, and a vendor with an obvious reason to like the result.

So go to somebody with no stake in it.

Brian Jabarian and Alex Imas at Chicago Booth built a corpus of about two thousand human-written passages — blogs, reviews, news, novels, restaurant write-ups, résumés — and ran three commercial detectors and an open-source model over it. The paper is Artificial Writing and Automated Detection, NBER working paper 34223.

Pangram came out with essentially zero false positives and essentially zero false negatives on medium-length and long passages. And it held up when the text had been through a humanizer.

That is an independent audit saying the laundering step did not work.

Why the race was never winnable

The two sides of this are not doing symmetric work.

A humanizer has to move text past a classifier without wrecking it. A detector only has to notice that the text was moved. Every humanizer that ships becomes training data for the next detector, because its output is abundant, labelled, and free to collect. WriteHuman pushed an update aimed at GPTZero in July. That is the loop, and the loop runs faster on the detector's side, because there are thousands of paying humanizer users generating samples and the detector needs no permission to read them.

Then there is what the laundering does to the writing. The tools that score best are the ones that scramble hardest. Odd word choices, inserted errors, sentences bent out of shape. The output is text nobody would choose to send. You bought a number and paid for it in prose.

The part that looks like a contradiction

We published a piece a few weeks ago arguing that a detector score is not evidence of who wrote something, and that these tools flag careful and non-native writing at rates nobody should be comfortable with. Now here is a paper saying detection works well.

Both are true, and they are about different failures.

A detector can be very good at spotting machine text that has been run through a laundering tool, and still be bad at proving a particular person did not write a particular paragraph. Those are separate questions. One is about text in bulk. The other is about a named human on a Thursday morning with their degree in the balance.

The same Chicago paper is careful about this. Every detector it tested lost accuracy on passages under fifty words. Short samples are where the false accusations live, and short samples are most of what a person actually gets judged on.

So: reliable enough to catch a laundering pipeline. Not reliable enough to convict anybody. Anyone quoting one of those findings to settle the other is selling something.

What this leaves

If the reason to rewrite is a number, the ground under that is moving and it is moving away from you.

If the reason is that a draft came back sounding like nobody, and you have to send it under your name, that has not changed at all. Readers are not running a classifier. They are noticing that the third sentence hedges the way the first one did, and that you have never once said delve out loud.

That reader is not going to retrain in July.

What we do about it

We do not publish a bypass rate, and we would not stand behind one if we did. Our page about Turnitin says so where the search traffic lands, which is the least convenient place to say it.

What the checker gives you is a list of what is in the text and why it is marked. It runs in your browser. The words do not leave the tab. The score is a description, and you can go and read the rule behind every mark.

Then you fix the writing, or you decide it reads fine and send it. The point is that you looked, and that what you send sounds like you.

The tools selling the number are going to keep selling it, and the number is going to keep sliding. Write for the person reading it. That target holds still.

Say the thing you meant.
Just not the way a model would say it.

Paste a draft, see every tell named, and take the rewrite. No card to start.