Skip to content

The writing our detector scores highest is not AI writing

We sorted every sample in our benchmark by the score our own detector gave it, and published the ranking. It is not the ranking you would want from a product you are selling.

This page is written by people who build one of these tools, which is a reason to check it as well as to read it. Every figure comes from the benchmark we publish, regenerated whenever the scoring changes, with a build that fails if the numbers on this page stop matching the engine.

Mean score by register

48 samples, 24 human and 24 AI. Higher means "reads more like AI to our engine". The threshold at which we would call something AI is 50.

AI AI, asked plainly · n=7 44.3
Human A human academic abstract · n=4 37.3
Human Human, authors at non-anglophone institutions · n=5 23.8
Human Human technical documentation · n=4 21.8
AI AI, writing on a technical subject · n=4 21
Human A human government notice · n=4 17.8
Human Human business copy · n=3 14.3
AI AI, run through our own humanizer · n=3 12.7
Human A human writing casually · n=4 12
AI AI, told to write casually · n=5 10.2
AI AI, told to vary its rhythm · n=5 9.8

Sample sizes are small — between 3 and 7 per register — so read these as directions, not as precise figures. That limitation is the reason our headline false positive rate carries a confidence interval of 1–20% rather than a single number.

What the ranking says

Only one register scores above 37.3, and it is ai, asked plainly — AI produced with no attempt to disguise it. Immediately below it is a human academic abstract, written by a person, which outscores 4 of the 5 AI registers in the benchmark.

At the bottom of the table is ai, told to vary its rhythm, averaging 9.8. The instruction that produced it was one line long. Below the human casual writing, below the government notices, below almost everything a person wrote.

Put plainly: on our benchmark, a student writing a careful academic abstract is scored higher than someone who typed one extra sentence into a chatbot. Our single false positive — the one piece of human writing we flagged — is an academic abstract. The better your formal register, the more exposed you are, and the person actually using AI has to do almost nothing to fall below you.

Why formal prose scores high

Detectors do not read meaning. They measure shape: how much sentence lengths vary, how often a sentence opens the same way as the last one, how wide the vocabulary spread is, how many stock connectives appear. Formal writing is trained toward evenness — consistent sentence length, conventional transitions, a restricted register. Machine writing is even for a different reason. The measurements cannot tell the reasons apart.

Which is also why the disguise is cheap. Ask a model to vary its rhythm and the shape stops matching, while nothing about the text's origin has changed. That asymmetry is not a bug we are about to fix — it is what style-based detection is.

One row you should not read too much into

The table includes a row for authors at non-anglophone institutions, and the honest thing to say about it is that it does not measure what you would expect it to. It is shown rather than hidden, with the corpus's own note reproduced as written:

READ THIS BEFORE PUBLISHING ANY ESL NUMBER. The five "esl-nonnative" samples are NOT learner-corpus essays. No learner corpus with a license compatible with commercial use could be sourced: PELIC is CC BY-NC-SA (NC excludes us), the Cambridge Learner Corpus FCE set, ICLE, ICNALE, NUCLE and the BEA-2019 W&I+LOCNESS data are all behind restrictive or registration-gated licenses, and the ETS TOEFL11 corpus is a paid LDC product. WHO and FAO publications were also rejected as CC BY-NC-SA. Rather than substitute something unlicensed, these five are CC BY / CC BY-SA scholarly prose written by authors whose institutional affiliations are all in non-anglophone countries (Indonesia, Turkey, Iran, China), selected because visible L2-transfer features survived peer review uncorrected. That is evidence of non-native authorship, not proof of any individual author's first language, and the register is academic rather than the undergraduate essay that produces our worst false positives. Treat any FPR computed on this slice as a floor, not as the ESL-essay false-positive rate.

If you have been accused

The table above is context, not a defence. What actually settles it is how the draft was written — version history, notes, the paragraph you deleted and rewrote. We wrote a separate page about that, including what to say and what not to do.