Detector tell · 16 patterns · weights 1–3
The biggest bucket in the table, and the least specific
Sixteen rows spanning weight 1 to weight 3 — everything from "delve into" to the word "significant". It is the largest single label in the engine and the one most likely to fire on ordinary prose.
Our detector calls this category “AI vocabulary”. The patterns below are
read out of functions/api/tools/detect.js when this page is
built, and the examples are produced by running the shipped scorer over the
48 licensed samples in our
benchmark. If a claim here stops matching the engine, the build fails.
What our own benchmark saw
Appeared in 2 of the human samples and 3 of the AI samples, 6 occurrences in total.
Matched: “holistic”, “paradigm”
-
“This research is a case of landscape evaluation based on the subjectivist or psychological paradigm.”
Our scorer gave this sentence 63 out of 100.
-
“In this article, three main holistic concepts formed during the war (war between Iran and Iraq) are determined, and then seniors majoring landscape architecture were asked to define these concepts for each the landscape elements.”
Our scorer gave this sentence 74 out of 100.
- Source:
- Sahar Pouya & Homa Irani Behbahani, "Landscape visual assessment: A case of Iran-Iraq war memorial garden", Turkish Journal of Forestry / Türkiye Ormancılık Dergisi 18(4), 2017
- Licence:
- CC BY (journal-level license recorded in DOAJ)
- How we know the date:
- Published 30 November 2017 — verified against Crossref (api.crossref.org/works/10.18182/tjf.294916 returns published date-parts [2017,11,30]). Five years before public LLMs. Authors are affiliated with İstanbul Teknik Üniversitesi and the University of Tehran (affiliations in the DOAJ record); both names are Persian, and the subject matter is the Iran-Iraq war. L2 features survive uncorrected: "This article emphasizes on", "have impressive role in making decision", "seniors majoring landscape architecture".
Matched: “significant”
-
“After the announcement of last year’s proposed settlement, the Commission learned that Uber had failed to disclose a significant breach of consumer data that occurred in 2016 -- in the midst of the FTC’s investigation that led to the August 2017 settlement announcement.”
Our scorer gave this sentence 41 out of 100.
- Source:
- Federal Trade Commission, press release "Uber Agrees to Expanded Settlement with FTC Related to Privacy, Security Claims", April 12, 2018 (quote from Acting FTC Chairman Maureen K. Ohlhausen)
- Licence:
- Public domain — work of the US federal government (17 U.S.C. §105)
- How we know the date:
- The release date April 12, 2018 appears on the FTC page itself and is encoded in the URL path (/2018/04/); the named official, Maureen K. Ohlhausen, was Acting FTC Chairman only until May 2018, which independently bounds the date. Four years before public LLMs.
Matched: “significant”
-
“This change addresses a recurring reconciliation issue: a significant share of submissions currently arrive more than a quarter after the fact, which complicates period-end close and makes budget variance difficult to interpret.”
Our scorer gave this sentence 41 out of 100.
- Source:
- claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming.
- Licence:
- generated for this benchmark
- How we know the date:
- No human wrote or edited any part of this text — it was produced token-by-token by the model in the session that wrote this file, and the session transcript is the primary record. There is no such policy, company, or October deadline; the memo is generated on instruction. Residual risk (stated, not dismissed): the model could in principle have reproduced memorized training text; that was not independently checked against a plagiarism index.
Matched: “revolutionize”
-
“Self-driving cars have the potential to revolutionize transportation by making it safer, more efficient, and more accessible.”
Our scorer gave this sentence 63 out of 100.
- Source:
- OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "wiki_csai", row_idx 11 (record id 11), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API.
- Licence:
- CC BY-SA 4.0 (per the HC3 dataset card on Hugging Face). Reproduced verbatim with attribution. NOTE for integration: ShareAlike attaches to adaptations of the dataset. Quoting a handful of records inside a larger benchmark collection is a collection, not an adaptation, so it does not relicense this repo — but the attribution and license line must travel with the samples wherever they are published.
- How we know the date:
- The AI label is the dataset's construction, not an inference from the text: Guo et al. built HC3 by putting each question to ChatGPT and recording its answer in `chatgpt_answers`, alongside separately-collected human answers in `human_answers`. The AI side is machine-generated by construction. Published Jan 2023 (arXiv:2301.07597) and widely cited since, so the labelling has had three years of public scrutiny. Caveat: OpenAI never disclosed the exact checkpoint behind the Dec-2022 ChatGPT web release, so the model id can only be given at family/date granularity.
Matched: “significantly”
-
“Companies typically announce stock splits when the market price of their stock has risen significantly, and they want to make it more affordable for individual investors to purchase shares.”
Our scorer gave this sentence 41 out of 100.
- Source:
- OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "finance", row_idx 15 (record id 15), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API.
- Licence:
- CC BY-SA 4.0 (per the HC3 dataset card on Hugging Face). Reproduced verbatim with attribution. NOTE for integration: ShareAlike attaches to adaptations of the dataset. Quoting a handful of records inside a larger benchmark collection is a collection, not an adaptation, so it does not relicense this repo — but the attribution and license line must travel with the samples wherever they are published.
- How we know the date:
- The AI label is the dataset's construction, not an inference from the text: Guo et al. built HC3 by putting each question to ChatGPT and recording its answer in `chatgpt_answers`, alongside separately-collected human answers in `human_answers`. The AI side is machine-generated by construction. Published Jan 2023 (arXiv:2301.07597) and widely cited since, so the labelling has had three years of public scrutiny. Caveat: OpenAI never disclosed the exact checkpoint behind the Dec-2022 ChatGPT web release, so the model id can only be given at family/date granularity.
2 of those samples were written by a person, years before any general-purpose model existed, with a licence and a publication date on record. 3 were machine-written. That is the whole reason these pages exist: a hit here is a fact about a phrase, not about an author.
About the non-anglophone sample above
READ THIS BEFORE PUBLISHING ANY ESL NUMBER. The five "esl-nonnative" samples are NOT learner-corpus essays. No learner corpus with a license compatible with commercial use could be sourced: PELIC is CC BY-NC-SA (NC excludes us), the Cambridge Learner Corpus FCE set, ICLE, ICNALE, NUCLE and the BEA-2019 W&I+LOCNESS data are all behind restrictive or registration-gated licenses, and the ETS TOEFL11 corpus is a paid LDC product. WHO and FAO publications were also rejected as CC BY-NC-SA. Rather than substitute something unlicensed, these five are CC BY / CC BY-SA scholarly prose written by authors whose institutional affiliations are all in non-anglophone countries (Indonesia, Turkey, Iran, China), selected because visible L2-transfer features survived peer review uncorrected. That is evidence of non-native authorship, not proof of any individual author's first language, and the register is academic rather than the undergraduate essay that produces our worst false positives. Treat any FPR computed on this slice as a floor, not as the ESL-essay false-positive rate.
In short: do not read this page as evidence about how detectors treat non-native English writers. It is one licensed, dated document that happens to contain the phrase.
A constructed example — written by us, not evidence
This sentence was written for this check. It is not from the benchmark, nobody wrote it as real prose, and it says nothing about how anyone writes. Its only job is to prove the rule still fires, and the generator fails the build if it stops.
“We will delve into a myriad of comprehensive scheduling options.”
-
/\ba myriad of\b/gi→ “a myriad of” -
/\bdelve into\b/gi→ “delve into” -
/\bcomprehensive\b/gi→ “comprehensive”
What it means when a person writes this way
"Significantly" is a required word in any paper reporting a statistical test. "Comprehensive" is what insurance is. "Robust" is what an engineer calls a system that survives bad input. This is one of the few labels our benchmark exercises at all, and it fires on both sides of it — the human samples and the machine ones. Read the counts above before you treat a hit here as meaning anything.
The patterns, as the engine holds them
16 of the 70 rows in our phrase table carry this label. They are printed here as they are written, because a paraphrase of a regular expression is a different regular expression.
/\ba myriad of\b/gi weight 3 Matches “a myriad of” , in any capitalisation .
/\bdelve into\b/gi weight 3 Matches “delve into” , in any capitalisation .
/\bembark on (?:a|this|the)\b/gi weight 3 Matches “embark on a”, “embark on this”, “embark on the” , in any capitalisation .
/\bholistic\b/gi weight 2 Matches “holistic” , in any capitalisation .
/\bcomprehensive\b/gi weight 2 Matches “comprehensive” , in any capitalisation .
/\bmultifaceted\b/gi weight 2 Matches “multifaceted” , in any capitalisation .
/\bparadigm\b/gi weight 2 Matches “paradigm” , in any capitalisation .
/\bstreamline\b/gi weight 2 Matches “streamline” , in any capitalisation .
/\brevolutionize\b/gi weight 2 Matches “revolutionize” , in any capitalisation .
/\bfoster a\b/gi weight 2 Matches “foster a” , in any capitalisation .
/\bseamless(?:ly)?\b/gi weight 2 Matches “seamlessly”, “seamless” , in any capitalisation .
/\bpivotal\b/gi weight 2 Matches “pivotal” , in any capitalisation .
/\bintricate\b/gi weight 2 Matches “intricate” , in any capitalisation .
/\brobust\b/gi weight 1 Matches “robust” , in any capitalisation .
/\bmeticulous(?:ly)?\b/gi weight 1 Matches “meticulously”, “meticulous” , in any capitalisation .
/\bsignificant(?:ly)?\b/gi weight 1 Matches “significantly”, “significant” , in any capitalisation .
What a hit does to your score
The phrase model is one of four in our ensemble, and its formula is
patternScore = min(98, max(5, 10 + 7 x total weight)). One
occurrence of the heaviest row in this category takes that model from 10 to
31 out of 100 —
and the pattern model carries 40% of the ensemble when the em-dash signal fires and 55% when it does not.
The sixteen rows are not equal: "a myriad of", "delve into" and "embark on a" carry weight 3, while "robust", "meticulously" and "significantly" carry weight 1. A single weight-1 hit moves the pattern model from 10 to 17 out of 100; the ensemble weights that model at 40-55%.
Check your own text
Our detector is free and shows you which sentences it flagged and why, including this category. It also gets things wrong, and we publish how often.
Other things the detector looks for
- Formal connector — 4 patterns
- GPT-4 closer — 8 patterns
- GPT-4 sycophant opener — 7 patterns
- GPT-4 labeled output — 4 patterns
- Structural marker — 4 patterns
- Hedging filler — 3 patterns
- All 22 categories