Pillar13 min readJul 31, 2026

Signs of AI Writing: 12 Patterns With Reproducible Thresholds

The signs of AI writing, each with a number you can measure: 12 patterns from em dash density to sentence-length variance, with reproducible thresholds.

TLToo Long; Didn't Read

  • Every sign has a number — a density, ratio, variance, or count you can check by hand or in a spreadsheet
  • We measure the writing, not the writer — these are AI writing characteristics that correlate with slop, not proof a machine was involved
  • No single number convicts — every pattern has an innocent explanation on its own
  • Convergence is the fingerprint — three or four signs failing at once in the same short passage is the signal
  • Substance is the most reliable sign — after a paragraph, can you restate one concrete fact?

Most lists of "signs of AI writing" hand you a feeling. It sounds too polished. It feels generic. Something's off. That's not something you can check, argue with, or defend when a teacher or a client pushes back. So we built this list differently.

Every one of the 12 AI writing patterns below comes with a number you can measure yourself — a density, a ratio, a variance, a count. Not "AI uses a lot of em dashes," but "more than 20 em dashes per 1,000 words in short modern prose." Not "the sentences feel samey," but "a burstiness score under 0.4." You can run these checks by hand, in a spreadsheet, or paste your text into a tool and watch them light up.

We build an AI slop detector that already measures most of these, so the thresholds here aren't guesses — they're the actual lines our engine draws. One caveat we'll repeat until it's annoying: no single number convicts. These are signals that get loud when they stack. Read the whole thing before you accuse anyone of anything.

Why is AI writing so bad — and what are we actually measuring?

Before the list, the honest question people keep asking: why is AI writing so bad in such a specific, recognizable way? The answer explains every pattern below.

A language model doesn't write to say something. It predicts the most probable next token given everything before it. Averaged over a planet's worth of text, the most probable next word is almost always the safest, blandest, most middle-of-the-road one. That's not a bug — it's the whole mechanism. The result reads like the statistical center of everything ever published: fluent, confident, grammatically spotless, and empty. It's the prose equivalent of a face generated by averaging ten thousand faces. Smooth. Featureless. Nobody's.

This is why we need to be precise about what we're detecting. We are not detecting whether a machine was involved. Origin detectors chase a moving target and routinely punish humans — a fact well documented in false-accusation cases. We're detecting a set of measurable AI writing characteristics that correlate with low-substance, templated text: what the internet now calls slop. Merriam-Webster made it official, naming "slop" a Word of the Year candidate and defining it as "digital content of low quality... produced usually in quantity by means of artificial intelligence" (Merriam-Webster).

A careless human can produce every pattern below. A careful prompt can dodge several. That's fine. We're measuring the writing, not the writer. Here are the twelve signs, grouped by the five dimensions our detector scores: Vocabulary, Cliché, Structure, Diversity, and Substance.

Vocabulary signs

1Excess "style words" (delve, tapestry, underscore)

The pattern: a specific set of words that were rare in everyday prose before late 2022 and then exploded. Not scientific terms — style words.

The threshold you can measure

Count the flagged words per 100. This is one of the best-measured signs in existence. In a peer-reviewed analysis of 14.2 million PubMed abstracts (2010–2024), Kobak and colleagues found the post-ChatGPT vocabulary spike was "almost entirely style words." Delves appeared at roughly 25× its pre-ChatGPT frequency; showcasing and underscores jumped about nine-fold. Their conservative estimate: at least 10% of 2024 abstracts were LLM-assisted, up to 30% in some subfields.

How to use it: one "delve" means nothing. A rate above roughly 3 flagged style words per 500 words, clustered, is a real signal. We keep the full tiered inventory in the AI words list, tagged by how badly each one gives you away.

When it lies: consultants and academics used some of these words for years, and writers who read the same viral lists now avoid them on purpose. After "delve" got called out in 2024, its rate in new papers dropped. The tells drift.

2Corporate verb inflation (utilize, leverage, facilitate)

The pattern: simple actions dressed in three-syllable Latinate verbs. Use becomes utilize. Help becomes facilitate. Boost becomes leverage.

The threshold

Flag any text where inflated verbs outnumber their plain equivalents. A quick proxy: if you see utilize used where use would work, and it happens more than once per 300 words, that's inflation, not vocabulary.

When it lies: legal, procurement, and academic registers legitimately run formal. Judge inflation against the genre.

Cliché signs

3Empty-opener clichés

The pattern: the throat-clearing phrase that opens on nothing. "In today's fast-paced world." "In the digital age." "In a world where." "Picture this."

The threshold

These are near-binary — the phrase is present or it isn't. One opener cliché in a piece is a yellow flag. Two or more in a 500-word text is a strong signal, because a human editor deletes the first one on sight and a model produces them as a default.

When it lies: motivational and ad copy lean on these by design and predate AI by a century.

4Fake-authority phrases

The pattern: the gesture at evidence with no evidence attached. "Studies have shown." "Experts agree." "Research suggests." "The data speaks for itself."

The threshold

Count claims-of-authority that carry no citation, name, number, or link. In genuinely researched writing, most authority claims come with a source. In slop, nearly all of them float free. A ratio above 50% uncited authority claims is the tell.

When it lies: casual writing skips citations too. Weigh this against Sign 11 (substance).

5Pseudo-wisdom filler

The pattern: the sentence that sounds profound and means nothing. "The key is to find balance." "True growth comes from within." "It's not about the destination; it's about the journey."

The threshold

The delete test. Remove the sentence. If the paragraph loses zero information, it was filler. A passage where more than a third of sentences survive deletion with no loss is running on filler.

Total cliché load matters more than any single phrase. A pile-up — an empty opener, then a fake-authority phrase, then pseudo-wisdom, all in two paragraphs — is the rhythm of generated prose. Three or more in a short passage is a strong signal on its own.

Structure signs

6Em dash density in the wrong register

The pattern: a steady drumbeat of clean em dashes joining clauses, in short modern prose where a comma or period would do.

The threshold — and this is the whole reason to measure instead of vibe

We counted 700,000+ words of published human prose and found human writers use dash constructions at roughly 3.7 to 10 per 1,000 words. Twain's Huckleberry Finn scores 10.13. Meanwhile a controlled study clocked GPT-4.1 at 10.62 em dashes per 1,000 words, about 3.3× a matched human baseline of 3.23. So presence proves nothing — Twain would fail a naive test. Our detector only starts adding points above 20 per 1,000 words (2 per 100), roughly double the most dash-happy human novel we measured. Below that, it counts for nothing.

When it lies: constantly, on its own. Careful human writers love the mark, and since late 2025 you can tell ChatGPT to stop using it — OpenAI's Sam Altman acknowledged the quirk publicly. We ran the full experiment in the em dash data study; the short version is that a dash is a keystroke, not a thought.

7Rule-of-three on autopilot

The pattern: everything arrives in tidy triplets. "Efficient, scalable, and reliable." "Plan, execute, and optimize." "More productive, more focused, more fulfilled."

The threshold

Count three-item parallel lists per 500 words. Humans use them occasionally. When a piece runs more than one polished triplet per 200 words, you're watching a pattern generator reach for the statistically safe balanced list.

When it lies: speechwriters and copywriters deploy the rule of three on purpose, and it works.

8The "not just X, it's Y" contrast tic

The pattern: the engagement-shaped reversal template. "It's not just about working harder — it's about working smarter." "The result? Transformation." "But here's the thing."

The threshold

One is fine. When the same contrast frame repeats three or more times in a single article, snapping into place paragraph after paragraph, that's a template, not a train of thought.

When it lies: skilled marketers use these frames deliberately. Pair with a failed substance check.

Diversity signs

9Low sentence-length variance (burstiness)

The pattern: sentence after sentence landing in the same 15-to-20-word range, each a complete balanced clause. Human writing is bumpy — a long sprawling sentence, then a short one. For emphasis. AI writing is a metronome.

The threshold you can actually compute

This one has a name and a formula. Burstiness is the standard deviation of sentence lengths divided by the mean. GPTZero, whose detector was built on this signal, reports that human writing tends to spread around 0.6–1.2, while model output clusters around 0.2–0.4. You can compute this in a spreadsheet: split on sentences, count words per sentence, take stdev ÷ mean. A result under 0.4 in ordinary prose is a real diversity flag.

When it lies: technical documentation, legal text, and translations are legitimately uniform. Low variance outside those genres is the signal — not uniformity as such.

10Narrow vocabulary range (low type-token ratio)

The pattern: the text keeps reaching for the same words. It sounds fluent, but the actual vocabulary is thin.

The threshold

Type-token ratio (TTR) — unique words divided by total words. Corpus studies across Wikipedia, blog posts, and academic abstracts find machine-generated text consistently shows lower TTR and a narrower lexical range than human writing, clustering around simpler, more uniform word choices. This one is noisy — TTR falls naturally as any text gets longer — so use it on matched lengths and as a supporting signal, not a headline.

When it lies: short texts and constrained topics naturally have lower diversity. Compare like with like.

11Transition-word stacking

The pattern: paragraphs that all open with the same formal connectors. "Furthermore." "Moreover." "Additionally." "Ultimately." Real writers vary how they join ideas; models overuse a small set.

The threshold

Count paragraph-opening transitions. When more than half of a piece's paragraphs open with a formal transition word, that's stacking.

When it lies: formal essays and structured reports use transitions heavily. Read alongside the burstiness check.

Substance signs

12Nothing you can restate (the deletion test)

The pattern: the single most reliable sign and the hardest to fake. After reading a paragraph, try to state one specific, concrete thing you now know — a name, a number, a date, a cause, a trade-off. AI text is engineered to sound informative while committing to nothing.

The threshold

The restatement test, scored across a piece. Read each paragraph and ask: can I name one concrete fact it gave me? If more than half the paragraphs fail — you finish them fluent but unable to say what they told you — the text is substance-poor, whoever wrote it.

AI VERSIONSurvives deletion; nothing is lost

"Nutrition plays a crucial role in overall wellness. By making mindful choices and understanding your body's needs, you can unlock a healthier lifestyle."

HUMAN VERSIONCan't be deleted without losing the fact

"Swap the 6 p.m. soda for water and you cut roughly 40,000 calories a year — about 11 pounds. That did more for my blood sugar than any app I tried."

You can strip every em dash, every "delve," and every rule-of-three from a slop paragraph and it stays slop — still vague, still substance-free, still confident about nothing. Punctuation is a keystroke. Substance is a thought.

When it lies: a nervous intern padding a report reads exactly this way, with no AI involved. Which is precisely the point: the flag is emptiness, not authorship.

How to read the twelve together

Here is the rule that separates a fair read from a witch hunt: no single sign convicts. Every pattern above has an innocent explanation on its own. A couple of em dashes, one rule of three, a lower TTR — each has a human cause.

What you want is convergence. Three or four signs failing at once in the same short passage is the fingerprint: hollow substance plus a style-word cluster plus a cliché pile-up plus burstiness under 0.4. Any one alone is noise. Four together is a signal.

And notice what you measured across all twelve: vocabulary, cliché, structure, diversity, substance. Whether a machine typed it never came up once. You judged the writing on quality — the only judgment that holds up, because a lazy human and a careless prompt produce the same slop. To see these patterns annotated on real specimens, our AI slop examples page breaks down actual cases so you can calibrate your own eye against the thresholds above.

Check your own text against all 12 at once

Running twelve checks by hand is slow, and your buzzword memory isn't as complete as a dictionary of 200+ flagged terms. That's what our detector automates.

Paste any text into the free, 100% local slop checker and it scores all five dimensions this article walks through — vocabulary, cliché, structure, diversity, substance — with the exact words, phrases, and dash densities that dragged the score down highlighted in place. Nothing leaves your browser. The analysis runs entirely on your machine; we never see your text. It won't tell you "a robot wrote this." It'll tell you something more useful: whether the writing is worth reading, and exactly which of these twelve signs it trips.

Keep reading in this series

FAQ

What are the clearest signs of AI writing?

The most reliable is substance: after reading a paragraph, can you restate one concrete fact? AI text is fluent but empty. Next come measurable patterns — excess style words (delve, tapestry), cliché pile-ups, low sentence-length variance (burstiness under 0.4), high em dash density (over 20 per 1,000 words in short prose), and formulaic structure like the rule of three. No single sign is proof; look for three or more together.

Why is AI writing so bad in the same recognizable way?

Because a language model predicts the most probable next word, and averaged across all text, the most probable word is the blandest, safest one. The output reads like the statistical center of everything ever published: grammatically perfect, confident, and featureless. The recognizable "AI voice" is that averaging showing through.

Can you measure AI writing patterns instead of guessing?

Yes — that's the point of this list. Em dash density, sentence-length burstiness (stdev ÷ mean), type-token ratio, cliché counts, and flagged-word rates are all computable by hand or in a spreadsheet. This article gives a reproducible threshold for each of the twelve patterns.

Is a single sign enough to call something AI-written?

No. Every individual sign has an innocent explanation — careful human writers use em dashes, formal essays stack transitions, short texts have low lexical diversity. The honest question isn't "was a machine involved" but "is this substantive or slop," and that verdict only holds when several signs converge.

Sources & Further Reading

Check Your Own Writing for Slop

Paste your content and it scores all five dimensions — vocabulary, cliché, structure, diversity, substance. 100% local. Nothing leaves your browser.

Try the Slop Detector