Skip to content

AI Detector / Accuracy

How Accurate Is This AI Detector?

Every detector is sometimes wrong. Here is exactly how often we are, and where.

4%

False positive rate

95% CI 1–20%

1 of 24 human samples flagged as AI

96%

False negative rate

95% CI 80–99%

23 of 24 AI samples passed as human

0.6

Separation

Mean AI score 22.2 minus mean human score 21.5, both rounded here; the separation is computed from the unrounded means.

0.41

ROC AUC

1.00 is perfect, 0.50 is a coin flip. On this adversarial corpus the score does not separate AI from human writing.

48

Corpus size

Labelled adversarial samples, threshold 50

Measured on our adversarial benchmark · last regenerated 2026-08-28 · updated with every scoring change

Read these numbers for what they are. 48 samples is a small corpus. We publish it anyway because a small measured number beats a large invented one — but a single additional false positive would move that 4% by 4 points. These are lab-benchmark results on deliberately hostile text, not a guarantee about your document.

What this tool can and cannot tell you

It can tell you this

Whether your draft still reads as machine-written: the stock transitions, the uniform sentence rhythm, the em-dash habit, the tidy rule-of-three lists. On our benchmark it did not flag a single one of the 5 texts written by non-native English speakers, which is the accusation we care most about never making.

It cannot tell you this

Whether a person wrote it. Ask any model to write casually, or to vary its sentence lengths, and the tells disappear along with our ability to see them — the 10 samples we prompted that way top out at 13, far under the 50 we call AI. Treat a low score as "no obvious tells", never as "a human wrote this".

This is a limit of the method, not a setting we can turn up. Our engine is statistical and runs on our own servers with no language model involved — that is what makes it free with no daily limit, and it is also why it sees style and nothing else.

Where it holds and where it fails

The two headline rates average over very different situations. Broken out by the kind of writing, from a corpus of 24 human texts published before 2020 and 24 AI texts with recorded model and prompt:

Kind of writing Samples Mean score We got wrong
Academic abstract (human) 4 37.3 1 of 4 wrongly flagged
Esl nonnative (human) 5 23.8 0 of 5 — none accused
Government bureaucratic (human) 4 17.8 0 of 4 — none accused
Technical documentation (human) 4 21.8 0 of 4 — none accused
Casual first person (human) 4 12 0 of 4 — none accused
Business marketing (human) 3 14.3 0 of 3 — none accused
Instructed casual (AI) 5 10.2 5 of 5 missed
Instructed varied rhythm (AI) 5 9.8 5 of 5 missed
Humanizer output (AI) 3 12.7 3 of 3 missed
Domain specific (AI) 4 21 4 of 4 missed
Default assistant (AI) 7 44.3 6 of 7 missed

"Humanizer output" is text from our own AI Humanizer. Our detector does not catch it. We would rather you learn that here than discover it yourself.

Use these numbers

Everything here is free to quote, reproduce and check, including the parts that make us look bad — a number nobody can verify is worth nothing. The raw benchmark output is one file, and the corpus records every sample's source and licence so you can re-run it against your own detector.

Raw data

detector-accuracy.json — every sample, its score, our call, and whether we got it right. Regenerated with each scoring change.

Citation

Coda One, AI Detector Accuracy Report, 2026-08-28. codaone.ai/ai-detector/accuracy — CC BY 4.0.

If you re-run it and get different results, we want to know — that is the only way a published number stays worth publishing. Tell us what you found.

What is wrong with our own corpus

Start with the sample size

The false positive rate is 1 of 24. At that sample size the 95% confidence interval is 1–20% — so the honest reading of "4%" is "somewhere in 1–20%, and we cannot narrow it without more samples." If you are here because a detector flagged something you wrote, the upper end of that range is the number to argue with, not the headline. The false negative rate, 96%, has a 95% interval of 80–99%. The interval above is on 48 samples. Pangram publishes 95% intervals too — on 1,000,000 human examples for its false positive rate and 519,993 generations across 26 models for its false negative rate, Wilson-derived and five orders of magnitude larger than ours. This paragraph used to say no competitor published an interval. It was wrong three times in one day, each version narrower than the last. What is left is not a claim about them: our numbers are the worst on the table and we print them anyway.

Every benchmark has a shape that flatters or punishes it. These are the limitations recorded by whoever built the corpus, reproduced as written rather than summarised by us:

How we know the human samples are human

Every sample was published between 2010 and 2019, i.e. before any general-purpose LLM was publicly available. Human authorship here is a date fact, not a stylistic judgement. Each sample records how the date was verified, preferring identifiers that cannot be back-dated: DOIs (checked against Crossref), PMC ids and Europe PMC firstPublicationDate, Federal Register document numbers, MediaWiki revision ids, git tags, and Stack Exchange server-assigned creation timestamps.

On the non-native English samples

READ THIS BEFORE PUBLISHING ANY ESL NUMBER. The five "esl-nonnative" samples are NOT learner-corpus essays. No learner corpus with a license compatible with commercial use could be sourced: PELIC is CC BY-NC-SA (NC excludes us), the Cambridge Learner Corpus FCE set, ICLE, ICNALE, NUCLE and the BEA-2019 W&I+LOCNESS data are all behind restrictive or registration-gated licenses, and the ETS TOEFL11 corpus is a paid LDC product. WHO and FAO publications were also rejected as CC BY-NC-SA. Rather than substitute something unlicensed, these five are CC BY / CC BY-SA scholarly prose written by authors whose institutional affiliations are all in non-anglophone countries (Indonesia, Turkey, Iran, China), selected because visible L2-transfer features survived peer review uncorrected. That is evidence of non-native authorship, not proof of any individual author's first language, and the register is academic rather than the undergraduate essay that produces our worst false positives. Treat any FPR computed on this slice as a floor, not as the ESL-essay false-positive rate.

Which models the AI samples came from

Read this before quoting any FNR computed from this file. WHICH MODELS ARE ACTUALLY IN HERE (24 samples): - claude-opus-5 (Anthropic), generated 2026-08-26 for this benchmark: 18 samples. - OpenAI gpt-4o-mini, as the engine of the Codaone production humanizer, rewriting claude-opus-5 text: 3 samples. Mixed lineage — Claude wrote it, GPT rewrote it. - OpenAI ChatGPT, December 2022 web release (gpt-3.5 era), via the HC3 dataset: 3 samples. WHICH ARE NOT: Google Gemini. Meta Llama. Mistral. DeepSeek. Qwen. xAI Grok. Cohere. Every current-generation OpenAI model (GPT-4o, GPT-5 family) except gpt-4o-mini in its humanizer role. Every current-generation Claude except opus-5. Nothing from a consumer "undetectable AI" service (StealthGPT, Undetectable.ai, Phrasly) — those are the strongest evasion tools in the wild and we have zero samples of their output. THEREFORE: an FNR computed over this corpus is an FNR *for Claude Opus 5 output, for our own humanizer's output, and for three-year-old ChatGPT output*. It is not an estimate of how often we miss AI text in general, and it must not be published as one. If the number is quoted on /ai-detector/accuracy, this sentence has to travel with it. A SECOND BIAS, LESS OBVIOUS: 18 of the samples were written by the same model that assembled this corpus, while it knew it was building a detector benchmark. That is a demand-characteristic risk in both directions — the model may have written more stereotypically "AI" prose for the default-register samples, and may have tried harder than a naive user would on the evasion samples. The prompts are recorded verbatim so a third party can rerun them on any model and check. Doing that on a non-Claude model is the cheapest available fix for both biases. THIRD, THE REGISTER MIX IS NOT A USAGE DISTRIBUTION: {"default-assistant":7,"instructed-casual":5,"instructed-varied-rhythm":5,"humanizer-output":3,"domain-specific":4}. It is deliberately weighted toward evasion (13 of 24 samples are instructed-casual, instructed-varied-rhythm, or humanizer output) because §1.4 of AI-DETECTOR-COMPETITIVE-2026-08-25 measured that all of our false negatives come from disguised AI. A corpus weighted this way will report a WORSE FNR than a corpus of default-register text would. That is intentional and honest; it is not comparable to a competitor's headline accuracy number, which is measured on whatever mix flatters them. FOURTH, ON THE HUMANIZER SAMPLES: the free anonymous quota is 3/day/IP and all three were used on 2026-08-26, so there are exactly three and no reruns. They are the highest-value rows in the file — a false negative on any of them means our own paid product defeats our own free product, which is the central question the accuracy page exists to answer honestly.

On how few voices some registers have

Registers are unevenly hard to source under an open license. Government and technical-documentation samples come from a small number of institutional voices (three Federal Register rules, two Kubernetes docs pages, three English Wikipedia revisions), so within-register variance understates the real world.

What is wrong inside the detector

The score above is an ensemble of four signals. Measured separately on the same 48 samples, two of them beat the blend they feed, and the blend is beaten again by its own score floors. 0.5 is a coin flip.

Signal ROC-AUC
statistical alone 0.407 ↓
pattern alone 0.560
dashDensity alone 0.522
tricolon alone 0.477 ↓
pattern + dash 50/50 0.563
weighted, before floors 0.443 ↓
SHIPPED (aiScore) 0.407 ↓

The score floors make it worse

Five rules can raise a score above what the weighted blend produced, to catch documents the blend misses. On this corpus they cost 0.036 AUC — they misfire more often than they fire.

The benchmark does not exercise the engine

The phrase table is 70 patterns across 22 named categories. 17 of those categories never fire on any of the 48 samples. Each of them does fire on text matching its own pattern, so this is a gap in the corpus rather than a dead rule — but it means the error rates above were measured without two thirds of what the detector looks for ever appearing.

Where the ranking inverts

Mean score by register. The registers most likely to be accused sit near the top, and AI given a one-line instruction to vary its rhythm sits at the bottom.

ai default-assistant n=7 44.3
human academic-abstract n=4 37.3
human esl-nonnative n=5 23.8
human technical-documentation n=4 21.8
ai domain-specific n=4 21
human government-bureaucratic n=4 17.8
human business-marketing n=3 14.3
ai humanizer-output n=3 12.7
human casual-first-person n=4 12
ai instructed-casual n=5 10.2
ai instructed-varied-rhythm n=5 9.8

What each error costs, and to whom

False positive — the expensive error

A false positive accuses a human of using AI. In a classroom that is an integrity case; for a freelancer it is a withheld payment; for an ESL writer it is both, plus the suggestion that their English is not their own. This error has a victim.

False negative — the cheap error

A false negative lets AI text pass as human. Nobody is wrongly accused, but the tool failed to do the job you came for. Ours is 96% — the number is that bad because our engine reads style, and AI that has been asked to avoid AI-sounding style leaves nothing for it to read. We would rather publish that sentence than a number that flatters us.

Every sample, every score

The full benchmark run behind the numbers above — including the ones we got wrong. Scores at or above 50 count as an AI call.

Truth Score Our call Sample
Human 24 correct esl-nonnative: Indonesian authors, education research abstract. L1-transfer markers: "The subject of this research are", "the interaction that occurs in students", "This interaction is built with heterogeneous student skills." Katarina Tri Utaminingtyas, Rachmadina Eka Herdianti, Inti Hayatul Fitria & Anton Prayitno, "Small Groups: Student Productive Interactions in Learning Cooperative (Case Study of Mathematics Learning at Junior High School in Pakis, Malang)", Educational Process: International Journal 6(2), 2017 — https://doi.org/10.22521/edupij.2017.62.3
Human 12 correct esl-nonnative: Turkish nursing academics, structured abstract. L2 markers: "The scale was carried out to the students", "While analyzing the research data;", "Fisher' Exact". Ebru Özen Bekar, Dilek Konuk Şener, Çetin Yılmaz & Şengül Cangür, "The Evaluation of Professional Self-esteem of Nurses and Social Workers Before and After Graduation", Sağlık ve Hemşirelik Yönetimi Dergisi (Journal of Health and Nursing Management) 4(3), 2017 — https://doi.org/10.5222/SHYD.2017.050
Human 31 correct esl-nonnative: Iranian-Persian L1 authors at Turkish/Iranian institutions; discursive, argumentative register — the closest thing in this set to a student essay. Sahar Pouya & Homa Irani Behbahani, "Landscape visual assessment: A case of Iran-Iraq war memorial garden", Turkish Journal of Forestry / Türkiye Ormancılık Dergisi 18(4), 2017 — https://doi.org/10.18182/tjf.294916
Human 10 correct esl-nonnative: Iranian clinical-trial abstract. Article omission throughout ("in treatment of patients", "in case they had"), a classic L1-Persian marker. Shakiba M, Moazen-Zadeh E, Noorbala AA, Jafarinia M, Divsalar P, Kashani L, "Saffron (Crocus sativus) versus duloxetine for treatment of patients with fibromyalgia: A randomized double-blind clinical trial", Avicenna Journal of Phytomedicine, 2018 — https://europepmc.org/article/PMC/PMC6235666
Human 42 correct esl-nonnative: Chinese-hospital author team, medical case series. Semicolon-splice and "which may demand an additional approach to the ongoing practice" are non-native constructions that survived peer review. Pandey S, Li L, Deng XY, Cui DM, Gao L, "Outcome Following the Treatment of Ventriculitis Caused by Multi/Extensive Drug Resistance Gram Negative Bacilli; Acinetobacter baumannii and Klebsiella pneumonia", Frontiers in Neurology 9:1174, 2019 — https://doi.org/10.3389/fneur.2018.01174
Human 42 correct academic-abstract: Review abstract, UK. Uniform long sentences, heavy formal connectors ("Furthermore", "Overall") — the canonical AI-lookalike register. Barton AJ, Hill J, Pollard AJ, Blohmke CJ, "Transcriptomics in Human Challenge Models", Frontiers in Immunology 8:1839, 2017 (Oxford Vaccine Group, University of Oxford) — https://doi.org/10.3389/fimmu.2017.01839
Human 15 correct academic-abstract: Structured systematic-review abstract with labelled sections. Extremely uniform sentence length and near-zero first-person voice. Cairns AE, Pealing L, Duffy JMN, Roberts N, Tucker KL, Leeson P, "Postpartum management of hypertensive disorders of pregnancy: a systematic review", BMJ Open 7(11):e018696, 2017 (University of Oxford) — https://doi.org/10.1136/bmjopen-2017-018696
Human 42 correct academic-abstract: Plant-genetics abstract, US. Dense nominalisation and hedged claims; the kind of prose humans write that scores as "too uniform". PLOS ONE 13(12):e0207723, 2018 — MSU-DOE Plant Research Laboratory, Michigan State University — https://doi.org/10.1371/journal.pone.0207723
Human 50 false positive academic-abstract: Statistical-ecology abstract, US. Long subordinate clauses, "However"/"In these settings" transitions. PLOS ONE 13(12):e0204150, 2018 — Department of Forestry, Michigan State University — https://doi.org/10.1371/journal.pone.0204150
Human 27 correct government-bureaucratic: FAA final rule summary. "Additionally", "Finally", "These actions are necessary to" — the exact connector profile our detector punishes. Federal Aviation Administration, "Regulatory Relief: Aviation Training Devices; Pilot Certification, Training, and Pilot Schools; and Other Provisions", final rule, Federal Register document 2018-12800 — https://www.federalregister.gov/documents/2018/06/27/2018-12800
Human 18 correct government-bureaucratic: FDA final rule summary — one 100-word sentence built out of semicolon-separated clauses. Human, and maximally machine-like. Food and Drug Administration, "Food Labeling: Revision of the Nutrition and Supplement Facts Labels", final rule, Federal Register document 2016-11867 — https://www.federalregister.gov/documents/2016/05/27/2016-11867
Human 11 correct government-bureaucratic: Department of Education interim final rule. Statutory cross-references and self-referential procedural language. Department of Education, "Student Assistance General Provisions, Federal Perkins Loan Program, Federal Family Education Loan Program, William D. Ford Federal Direct Loan Program, and Teacher Education Assistance for College and Higher Education Grant Program", interim final rule, Federal Register document 2017-22851 — https://www.federalregister.gov/documents/2017/10/24/2017-22851
Human 15 correct government-bureaucratic: CDC surveillance report opening. Numbered citations, passive voice, agency-as-actor ("CDC analyzed data from..."). Cree RA, Bitsko RH, Robinson LR, Holbrook JR, Danielson ML, Smith C, et al., "Health Care, Family, and Community Factors Associated with Mental, Behavioral, and Developmental Disorders and Poverty Among Children Aged 2-8 Years — United States, 2016", MMWR Morbidity and Mortality Weekly Report 67(50), 2018 — https://europepmc.org/article/PMC/PMC6342550
Human 16 correct technical-documentation: Encyclopedic network-protocol description. Flat declaratives, list-like enumeration, zero first person. Wikipedia contributors, "Transmission Control Protocol", English Wikipedia, revision 843439827 (2018-05-29T05:04:10Z) — https://en.wikipedia.org/w/index.php?title=Transmission_Control_Protocol&oldid=843439827
Human 42 correct technical-documentation: Cryptography explainer. Long conditional sentences and "For this to work it must be" constructions. Wikipedia contributors, "Public-key cryptography", English Wikipedia, revision 736786168 (2016-08-29T20:54:56Z) — https://en.wikipedia.org/w/index.php?title=Public-key_cryptography&oldid=736786168
Human 16 correct technical-documentation: Product documentation, procedural voice. Repetitive parallel clauses ("does not kill... does not create...") read as templated. Kubernetes documentation, "Deployments" (content/en/docs/concepts/workloads/controllers/deployment.md), kubernetes/website at tag release-1.16 — https://github.com/kubernetes/website/blob/release-1.16/content/en/docs/concepts/workloads/controllers/deployment.md
Human 13 correct technical-documentation: Networking internals documentation. Comparative technical claims with hedging, written by contributors for whom English varies. Kubernetes documentation, "Service" (content/en/docs/concepts/services-networking/service.md), kubernetes/website at tag release-1.16 — https://github.com/kubernetes/website/blob/release-1.16/content/en/docs/concepts/services-networking/service.md
Human 17 correct casual-first-person: Blunt forum answer. Sentence fragments, an em dash, a typo ("polystychrene"), and a joke — high burstiness. Stack Exchange user "Criggie", answer 51225 on Bicycles Stack Exchange, 2017-12-01 — https://bicycles.stackexchange.com/a/51225
Human 10 correct casual-first-person: Advice post with imperatives, a parenthetical aside and an editorialising last line. Wildly uneven sentence lengths. Stack Exchange user "keshlam", answer 82857 on The Workplace Stack Exchange, 2017-01-12 — https://workplace.stackexchange.com/a/82857
Human 10 correct casual-first-person: Reflective first-person explanation of social convention, with quoted dialogue and a self-deprecating closing parenthesis. Stack Exchange user "Max", answer 38180 on Travel Stack Exchange, 2014-11-03 — https://travel.stackexchange.com/a/38180
Human 11 correct casual-first-person: Wikipedia talk-page comment. Enumerated grievance, slang ("crabon"), signed and timestamped by the editor. Wikipedia editor Dennis Bratland, comment dated 22:30, 20 July 2014 (UTC) on Talk:Bicycle; captured in English Wikipedia revision 664757797 (2015-05-30T21:01:42Z) — https://en.wikipedia.org/w/index.php?title=Talk%3ABicycle&oldid=664757797
Human 17 correct business-marketing: Enforcement press release. Announcement-lede structure, superlatives ("record", "by far the largest"), third-person institutional voice. Federal Trade Commission, press release "Google and YouTube Will Pay Record $170 Million for Alleged Violations of Children's Privacy Law", September 4, 2019 — https://www.ftc.gov/news-events/news/press-releases/2019/09/google-youtube-will-pay-record-170-million-alleged-violations-childrens-privacy-law
Human 14 correct business-marketing: Press release with an executive quote — the register that reads most like generated corporate copy. Federal Trade Commission, press release "Uber Agrees to Expanded Settlement with FTC Related to Privacy, Security Claims", April 12, 2018 (quote from Acting FTC Chairman Maureen K. Ohlhausen) — https://www.ftc.gov/news-events/news/press-releases/2018/04/uber-agrees-expanded-settlement-ftc-related-privacy-security-claims
Human 12 correct business-marketing: Product release announcement. Feature-benefit sentences and capability claims — vendor marketing prose written by engineers. Kubernetes 1.16 Release Team, "Kubernetes 1.16: Custom Resources, Overhauled Metrics, and Volume Extensions", Kubernetes blog, 2019-09-18 — https://kubernetes.io/blog/2019/09/18/kubernetes-1-16-release-announcement/
AI 42 missed — passed as human default-assistant: Classic essay register: abstract subject, "Furthermore", uniform long sentences. The easy case — if we miss this we have nothing. Also the input to humanizer sample #1. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 30 missed — passed as human default-assistant: Default register with one em-dash. Included deliberately: commit ba4f0aef demoted em-dashes from conviction evidence to corroboration, and this sample is the regression guard for that decision on the AI side. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 70 correct default-assistant: SEO-blog intro register — the highest-volume real-world use of an LLM and the text most likely to be pasted into a detector by an editor. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 42 missed — passed as human default-assistant: Default register, informational/advisory. Also the input to humanizer sample #2, so the pair isolates what the humanizer actually changes. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-casual: Texting register: lowercase, no terminal punctuation on the last line, fragments. Direct replication of the §1.4 miss that scored 15. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 11 missed — passed as human instructed-casual: Reddit-comment imitation — the §1.4 miss that scored 9, and the single cheapest evasion a real user can perform (one sentence of prompt). claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-casual: Fake user review in a chatty voice — the commercial evasion case (review farms), and a register where sentence-capitalization is preserved so the detector cannot key on lowercase alone. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 13 missed — passed as human instructed-casual: Self-deprecating personal confession — tests whether emotional first-person content plus low lexical formality is enough to push the score under threshold. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-casual: Personal-blog voice with the transition words explicitly banned in the prompt — isolates how much of our AI signal is carried by "moreover / additionally" alone. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 11 missed — passed as human instructed-varied-rhythm: Extreme length variance (2-word sentences next to 45-word sentences). This is the direct attack on a CV-of-sentence-length feature. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-varied-rhythm: Varied rhythm carrying an anecdote with reported speech — the register a student actually gets when they ask for "a personal essay that does not sound like AI". claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-varied-rhythm: Argumentative op-ed with varied rhythm and a concrete verifiable claim (Buffalo 2017). Tests whether specific facts and dates read as human to us. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 11 missed — passed as human instructed-varied-rhythm: Literary reflective memoir with varied rhythm — sits directly on top of the human Woolf/Fitzgerald samples in the human half. If this scores like they do, the score is not separating authorship, it is separating genre. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 9 missed — passed as human instructed-varied-rhythm: Technical explainer written with varied rhythm and a second-person analogy — the hardest combination for us, because it is also exactly how a good human technical writer writes. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 10 missed — passed as human humanizer-output: Our production humanizer applied to default-assistant sample #1 (remote work). API metrics on the call: aiScoreBefore 24, aiScoreAfterPass1 26, aiScoreAfter 8, iterations 2, 120→141 words. Note the humanizer's own estimator already scored the untouched Claude essay at only 24 — that estimator disagreeing with detect.js is its own finding. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (default-assistant sample #1, remote-work essay) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize
AI 12 missed — passed as human humanizer-output: Our production humanizer applied to default-assistant sample #4 (small-business cybersecurity). API metrics: aiScoreBefore 24, aiScoreAfterPass1 18, aiScoreAfter 14, iterations 2, 114→152 words. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (default-assistant sample #4, small-business cybersecurity paragraph) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize
AI 16 missed — passed as human humanizer-output: Our production humanizer applied to a formal academic abstract (domain-specific sample #1, lightly shortened to fit the free word budget). API metrics: aiScoreBefore 16, aiScoreAfter 20, iterations 1, 123→141 words. This is the one call where our estimator scored the output HIGHER than the input (16→20) — the humanizer made formal text read more like generic AI, not less. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (domain-specific sample #1, an academic abstract, with the "(n = 4,812)" parenthetical and the final clause about student effort removed so the input fit the 300-word free budget cleanly) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize
AI 12 missed — passed as human domain-specific: AI-written empirical abstract. Pairs with the false-positive side: if human academic prose scores high AND this scores high, the feature is "academic", not "AI". Also the stage-1 input for humanizer sample #3. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 42 missed — passed as human domain-specific: AI-written contract-law analysis. Direct counterpart to the human Marbury v. Madison sample in the human half — same register, opposite label. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 13 missed — passed as human domain-specific: AI-written API reference. Counterpart to the human Wright brothers patent specification in the human half: both are uniform procedural prose with near-zero sentence-length variance, which is precisely where a variance-based score has no information. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 17 missed — passed as human domain-specific: AI-written internal corporate memo. The register a real employee most plausibly delegates to an LLM, and the register a manager most plausibly runs through a detector. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
AI 42 missed — passed as human default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). Encyclopedic explainer register. Three years older than our own samples, so it also probes whether we only detect current-generation phrasing. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "wiki_csai", row_idx 11 (record id 11), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3
AI 42 missed — passed as human default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). Financial explainer, domain register. Contains a real generation artifact — a missing space at "potential investors.Stock splits" — which is preserved verbatim rather than cleaned up. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "finance", row_idx 15 (record id 15), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3
AI 42 missed — passed as human default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). ELI5 prompt, but note the answer is still in default assistant register — the casual PROMPT did not produce a casual REGISTER. That contrast with our instructed-casual block is the point of including it. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "reddit_eli5", row_idx 10 (record id 10), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3

Where we know we get it wrong

Seven categories of text where this detector — and, in most cases, every statistical detector — is systematically wrong or refuses to answer. For each: what happens, why, and what we do about it.

Correction, 2026-08-27: the "what we do" line on three of these cards described a mitigation this scoring does not implement — a verdict held at "Possibly" with low confidence for formal, ESL and uniform-technical prose — and quoted two verdict strings, "Mixed or AI-assisted" and "Possibly AI-generated", that this tool has never returned. There is no register-aware cap anywhere in the code. The lines below describe what actually runs.

ESL and non-native English writing

False positive risk

Why it fails: Statistical detection reads simpler vocabulary and regular sentence rhythm as machine-like. This is the best-documented bias in the entire detector category, and it lands on the people with the most at stake. In the run above, the 5 samples written by non-native English speakers averaged 23.8, the highest of them reached 42, and none was flagged.

What we do: Nothing in the score is register-aware. We cannot tell that English is your second language, so we do not claim to adjust for it — the number you get is the same number any other writer would get for the same prose. One rule does read register, and it lowers our stated confidence rather than your score: prose with contractions and lowercase sentence starts caps confidence at low below 40, because our own feature screen showed we cannot separate casual-prompted AI from casual humans. What the code does do: scores from 40 to 59 return "Mixed — some AI-sounding passages" at low confidence, scores under 20 return "No obvious AI tells" at low confidence, and any input under 100 words drops one confidence notch. The per-sentence view names the exact phrases that moved the score. If you write in English as a second language and got flagged, that context outweighs our number.

Formal, legal, and corporate human prose

False positive risk

Why it fails: Connectors like "Furthermore" and "In conclusion" are also among the most reliable LLM tells, so a regulation, a contract, or a press release can out-score actual AI text on the pattern model. In the run above the 4 government-bureaucratic samples averaged 17.8 and the 3 business-marketing samples 14.3, none of them flagged — but the 4 academic abstracts averaged 37.3 and produced the single false positive in the whole run, at 50.

What we do: Nothing caps this band, and we should not have implied otherwise. Four or more stacked formal connectors in one document trip a floor that sets the score to at least 78 whatever the other signals say: a 203-word human regulatory passage returns 78, "Reads like AI in places", medium confidence. What we can offer is visibility — the deep report now prints the weighted blend of the four models and then names the rule that overrode it ("5 stacked formal connectors raised the score to 78%"), so you can see that the number came from counting connectors rather than from four signals agreeing.

Uniform technical documentation

False positive risk

Why it fails: Good documentation is deliberately uniform: same sentence shape, same structure, restrained vocabulary. A variance-based model reads that discipline as generation. In the run above the 4 technical-documentation samples averaged 21.8 and the highest reached 42.

What we do: Uniform rhythm has a floor of its own, and it is not capped either: at four or more sentences, a sentence length that varies by under 15% of the average forces the score to at least 50, and under 25% to at least 42. That is why disciplined documentation drifts upward even with no AI phrasing anywhere in it. The deep report names that rule when it fires, and the per-sentence breakdown will show you an empty flag column — which is the tell that no actual phrasing was found.

Punctuation-styled writing (em-dash-heavy prose)

False positive risk (fixed, still watched)

Why it fails: We used to over-weight punctuation habits as an AI tell: an em-dash-heavy human essay scored 80 on an earlier version of the scoring. Plenty of humans write like that; so do LLMs.

What we do: We downgraded punctuation from primary evidence to corroboration: dashes now need an independent signal before they can raise a verdict. The em-dash-heavy human sample kept for this scores 12 today. It is not in the table below — it is one of 11 regression cases, each encoding a specific defect this detector was caught making, which run on every scoring change but are excluded from the rates above because they were written to trap a bug rather than sampled from anywhere. They are in the raw data file under "regressionResults".

AI written to sound casual or literary

False negative risk

Why it fails: Prompt an LLM to write casually — or run its output through a paraphraser — and the formal tells our pattern model keys on are stripped away. 13 of the 24 AI samples above are evasion attempts of that kind; every one of them passed as human, and the highest any of them scored was 16. This is the main way to beat us, and every other statistical detector.

What we do: Mostly, we cannot fix this within our approach, and we say so. Treat a low score on text you already suspect as weak evidence of anything. Uneven per-sentence scores on a polished piece are sometimes the residue of partial rewriting — a lead worth following, not a verdict.

Short samples

Unreliable both ways

Why it fails: Statistics computed over a handful of sentences are close to noise. Sentence-length variance, vocabulary richness, and starter diversity all need material to measure.

What we do: Under 30 words we refuse to give a verdict at all. Just above that line, read the score as a hint, not a finding.

Non-English text

Refused, or weaker

Why it fails: Every signal here is English-shaped: words split on spaces, sentences end on a full stop, and the phrase list is English. In a script that does not space its words, a whole paragraph tokenises to roughly one "word", so vocabulary richness pins at 1 and length variance at 0 — and the arithmetic still produces a confident-looking number. It was not merely weak, it was inverted: AI-written Chinese scored 9 ("No obvious AI tells") while human-written Chinese scored 55.

What we do: Chinese, Japanese, Korean, Thai, Lao, Khmer and Burmese input is now refused before scoring — the tool returns an error explaining why, not a number. Other space-separated scripts (Arabic, Hebrew, Cyrillic, Greek, Devanagari) still score, because their structural signals survive; those results carry a note saying the phrase-level checks are English-only and the score is weaker evidence than usual.

Methodology

Where the samples come from

The 24 human samples are open-licence prose published between 2010 and 2019, each with its source and licence recorded in the file: peer-reviewed papers under CC BY (PLOS ONE, BMJ Open, Frontiers), US federal government text in the public domain, Wikipedia and Stack Exchange under CC BY-SA, and Kubernetes documentation under CC BY. Human authorship is a date fact, not a judgement — every one predates instruction-tuned models. By register: 5 written by non-native English speakers, 4 academic abstracts, 4 government-bureaucratic, 4 technical documentation, 4 casual first-person, 3 business marketing. The 24 AI samples are genuine model output with the prompt recorded verbatim — 18 from claude-opus-5, 3 from gpt-4o-mini via our own humanizer, 3 real ChatGPT from December 2022. By register: 7 default assistant voice, 5 instructed to write casually, 5 instructed to vary sentence rhythm, 3 humanizer output, 4 domain-specific. That weighting is deliberate and it makes the numbers above look worse than a default-register corpus would: 13 of the 24 AI samples are evasion attempts, which is what a student told to “make it sound like me” actually produces. Correction, 2026-08-27: this section previously described a different set of samples — Darwin, Marbury v. Madison, the Wright brothers’ patent, Twain, Woolf, Fitzgerald. Those are regression fixtures kept elsewhere in the harness and they are not in the corpus these figures come from. The page that exists to be honest about method was describing the wrong method.

What runs

The benchmark feeds every sample through the same scoring engine that answers the live tool's detection API. There is no separate "benchmark mode" — what we measure is what you get.

The corpus

48 labelled English samples, adversarial in both directions: human writing that looks AI-ish to a statistical model (an academic abstract, legal prose, technical documentation, em-dash-heavy essay writing) and AI text prompted to sound casual and human. Uniform-but-human prose is over-represented on purpose, because that is exactly what a variance-based detector gets wrong. Each sample records why it is in the set, so a regression tells us which property broke — not just which string.

The threshold

Scores at or above 50 count as an AI call. False positive rate is the share of human samples at or over that line; false negative rate is the share of AI samples under it. Separation — mean AI score minus mean human score — matters more than either: if it is small, the score is noise no matter where the threshold sits.

Regeneration

Every change to the scoring reruns this benchmark, and the run rewrites the data file this page is built from. The numbers above cannot drift away from the code, because they are produced by it.

The promise

We will never quietly change these numbers. If a scoring change makes them worse, this page publishes worse numbers. The benchmark data lives in version control next to the scoring code, so the history of every published figure is inspectable.

Why publish this at all

AI detectors are evidence, not proof. The measurable signals in a piece of text — sentence rhythm, vocabulary spread, phrase patterns — overlap between careful human writers and language models, and no amount of engineering makes that overlap disappear. Anyone claiming 99% accuracy is selling certainty they cannot deliver, because the certainty does not exist to sell.

We would rather show you a small, honest benchmark than a big, unverifiable claim. The numbers on this page are less flattering than the ones you will see advertised elsewhere, and that is precisely why you can trust them: they were measured on hostile inputs by the same code that scores your text, and they update automatically whether they improve or not.

Use the detector the way we use it ourselves: as one signal among several, weakest exactly where this page says it is weakest, and never as the sole basis for a decision about a person.

Frequently Asked Questions

How accurate is this AI detector really?
On our adversarial benchmark of 48 labelled samples, we flagged 1 of 24 human samples as AI (a 4% false positive rate) and missed 23 of 24 AI samples (a 96% false negative rate) at the default threshold of 50. The corpus is small and deliberately hostile, so these are benchmark numbers, not universal accuracy. No detector has universal accuracy numbers, whatever its marketing says.
Can an AI detector prove someone used AI?
No. A detector score is statistical evidence, not proof. It should start a conversation — alongside drafts, writing history, and context — never end one. This is true of every detector, including any that advertises 99% accuracy.
Why publish your error rates at all?
Because the honest answer to "how accurate is your detector" is a measured number on a corpus you can inspect, not a marketing claim. We have no institutional contracts built on an accuracy figure, so we can publish exactly what we measure — including the categories of text where we fail.
Will these numbers change?
Yes. Every time we change the scoring, the same benchmark reruns and this page republishes whatever it reports, better or worse. The numbers are generated by the script that gates our scoring changes, and the data file lives in version control next to the scoring code. We will not quietly edit them.

Back to AI Detector

More AI Tools: AI Humanizer ·AI Rewriter ·Grammar Checker