AI Slop Detection: Two-Tier Reflex Methodology skill
A phrase blocklist catches the obvious tells.
by AgriciDaniel·MIT license·★ 2,219 Stars on the repo·GitHub ↗
Files of AI Slop Detection: Two-Tier Reflex Methodology
AgriciDaniel/
Show the full text157 lines
AI Slop Detection: Two-Tier Reflex Methodology
A phrase blocklist catches the obvious tells. Most AI-generated prose passes that filter and still reads like AI. The structural tics, the rhythmic flatness, the "everything is a three-clause sentence" cadence: those survive the first pass.
This reference defines a two-tier reflex check for editorial review. Run both passes before declaring a draft human-natural. Adapted from the impeccable plugin's UI slop methodology (Paul Bakaus, Apache 2.0).
Why two tiers
LLMs converge on a small set of safe patterns. The first thing the model reaches for is the first-order reflex: the genre-obvious tell. Replace it and the model reaches for the second-order reflex, the next-most-trained pattern that survives anti-AI guidance.
Most AI-detection passes only check the first-order pattern. The result is "anti-AI" rewrites that still read like AI because the structural pass was never run.
Note on terminology: this file uses "first-order" and "second-order" for the two detection passes. Elsewhere in the project, "Tier 1 / Tier 2 / Tier 3" refers to source authority (Google Search Central = Tier 1, Ahrefs = Tier 2, reputable industry sources = Tier 3). The two namespaces are intentionally kept separate; do not call the second-order detection pass "Tier 2."
Examples of the same idea across both tiers:
| Topic | First-order tell | Second-order tell that survives |
|---|---|---|
| SEO blog | "In today's digital landscape..." | Every H2 ends with a rhetorical question |
| SaaS post | "Game-changer," "Revolutionize" | Three-clause sentence rhythm, "While X, also Y" framings |
| How-to guide | "Dive into," "Unlock the potential" | Every step opens with an imperative verb identical in length |
| Listicle | "Cutting-edge," numbered fluff | Every item is ~80 words, identical structure |
| Thought leadership | "Comprehensive guide," "harness the power" | Hedge stack: "often," "typically," "may" within 20 words |
The point: a draft can score zero on a phrase blocklist and still be obviously AI.
First-order reflex (phrase + lexical)
This is what the existing AI-detection in blog-analyze and blog-rewrite already covers. Documented here for completeness.
Trigger phrases (full list in agents/blog-reviewer.md and scripts/analyze_blog.py):
- "In today's digital landscape" / "In the ever-evolving"
- "It's important to note" / "It is worth mentioning"
- "Dive into" / "deep dive" / "delve"
- "Game-changer" / "Revolutionize" / "transformative"
- "Cutting-edge" / "state-of-the-art" / "robust"
- "Harness the power" / "Unlock the potential"
- "Leverage" (as a verb, non-financial)
- "Seamlessly" / "seamless integration"
- "Tapestry" / "rich tapestry" / "multifaceted"
- "Comprehensive guide" (in body text)
- "Furthermore" / "Moreover" (transition overload)
- Em dashes used as a stylistic flourish (any density)
Lexical signals:
- AI trigger-word density > 5 per 1,000 words
- Type-Token Ratio (TTR) below 0.40 on long-form
- Burstiness (sentence-length standard deviation / mean) below 0.3
Outcome of the first-order pass: a "phrase-clean" draft. Necessary, not sufficient.
Second-order reflex (structural + rhythmic)
These are the patterns LLMs default to after the obvious vocabulary is replaced. They are structural and rhythmic, so a vocabulary swap doesn't fix them. Run this pass on drafts that already passed the first-order check.
Structural tics to flag
Question-cadence H2s. Every section heading is phrased as a question. Real long-form mixes question, statement, and noun-phrase headings. Flag if > 70% of H2 headings end with a question mark.
The Here opener. A paragraph opens with the word "Here" ("Here's why...", "Here are five..."). Once is fine. Three or more in a 1,500-word post is an AI fingerprint.
Three-clause sentence rhythm. Most sentences in a paragraph follow the structure
[clause], [clause], [clause].The cadence is metronomic. Flag if > 50% of sentences in any 200-word window match this shape.False-balance framing. Repeated use of "While X, also Y" or "On one hand X, on the other Y" without a real contrast. The model uses it to feel even-handed but it adds no information. Flag if it appears more than twice per 1,000 words.
Hedge stacking. Three or more hedges in a 20-word span ("It may often be the case that..."). Flag any 20-word window with > 2 of: may, might, often, typically, generally, usually, tend to, perhaps, somewhat, likely.
Symmetric list bloat. Every item in a numbered or bulleted list is the same length within +/- 10 words and follows the same syntactic structure. Real lists vary; some items need one line, others need a paragraph. Flag if list-item length standard deviation < 5 words.
The wrap-up question. Section ends with "What does this mean for [audience]?" or "Why does this matter?" Once per post is rhetorical; three or more is filler.
Capsule transitions. Each H2 opener begins with a single-word transition ("First..." "Next..." "Additionally..." "Crucially..."). Real prose buries transitions inside sentences. Flag if > 50% of H2 openers start with a transition word.
The "key insight" tell. The phrase "The key insight is..." or "What's important here is..." appears as a sentence-opener. This is the model telegraphing that it's about to summarize. Cut and let the sentence stand.
Listicle introduction bloat. Before the actual list, three or more paragraphs of "context." Real listicles get to the list. Flag if > 250 words of pre-list intro.
Rhythmic signals to compute
- Sentence-length flatness within paragraphs. Compute SD of sentence length per paragraph; flag any paragraph with internal SD < 4.
- Opening-word repetition. Count first-word frequencies across all sentences. Flag if the top three first-words account for > 25% of all sentence openings.
- Paragraph-shape flatness. Compute SD of paragraph word counts across the post; flag if < 25 (real long-form varies dramatically).
How to run the two-tier check
For blog-rewrite and blog-reviewer:
- Run the first-order pass first (phrase + lexical). If it fails, fix and re-run before moving on.
- Once the first-order pass is clean, run the second-order pass. Report each second-order pattern with line numbers and an example.
- Do not declare "AI-detection passed" unless both passes are clean.
For blog-write (initial drafting):
- First-order is enforced at generation time via the persona's anti-phrase list.
- Second-order is checked once on the full draft before delivery.
Output format
When reporting findings, use:
## AI Slop Detection Report
### First-order (Phrase + Lexical)
- Trigger phrases: [N found] -> [list with line numbers]
- AI trigger words: [N/1K words], [pass/fail at ≤5]
- TTR: [score], [pass/fail at ≥0.40]
- Burstiness: [score], [pass/fail at ≥0.3]
### Second-order (Structural + Rhythmic)
- Question-cadence H2s: [X%], [pass/fail at ≤70%]
- "Here" openers: [N], [pass/fail at ≤2]
- Three-clause rhythm: [X%], [pass/fail at ≤50%]
- False-balance framings: [N/1K words], [pass/fail at ≤2]
- Hedge stacking: [N windows], [pass/fail at 0]
- Symmetric list bloat: [N lists], [pass/fail at 0]
- Wrap-up questions: [N], [pass/fail at ≤2]
- Capsule transitions on H2s: [X%], [pass/fail at ≤50%]
- "Key insight" sentence openers: [N], [pass/fail at 0]
- Listicle intro bloat: [pre-list words], [pass/fail at ≤250]
- Sentence-length flat paragraphs: [N], [pass/fail at 0]
- Opening-word repetition: [top-3 share], [pass/fail at ≤25%]
- Paragraph-shape SD: [value], [pass/fail at ≥25]
### Verdict
First-order: [PASS / FAIL]
Second-order: [PASS / FAIL]
Overall: [PASS only if both passes clean]
Why this matters for ranking + AI citations
- Editorial quality risk: content that lacks experience and original perspective is easier to classify as interchangeable consensus content. Second-order patterns are a warning sign, even when no specific Google update is cited.
- AI citations: ChatGPT and Perplexity reward citable, distinctive passages. Second-order tics produce interchangeable prose that no AI surface has reason to prefer over the source it was trained on.
The two-tier check is the editorial parallel to impeccable's "design slop" methodology: vocabulary-clean is necessary but not sufficient; structural distinctiveness is what separates citeable content from indexable filler.
Attribution
The two-tier first-order / second-order reflex methodology is adapted from the impeccable plugin v3.1.1 (Paul Bakaus, Apache 2.0, https://github.com/pbakaus/impeccable). The original applies it to UI design cliches ("observability -> dark blue"). This reference adapts the same mental model to prose.
| 1 | # AI Slop Detection: Two-Tier Reflex Methodology |
| 2 | |
| 3 | A phrase blocklist catches the obvious tells. Most AI-generated prose passes that filter and still reads like AI. The structural tics, the rhythmic flatness, the "everything is a three-clause sentence" cadence: those survive the first pass. |
| 4 | |
| 5 | This reference defines a **two-tier reflex check** for editorial review. Run both passes before declaring a draft human-natural. Adapted from the impeccable plugin's UI slop methodology (Paul Bakaus, Apache 2.0). |
| 6 | |
| 7 | |
| 8 | |
| 9 | ## Why two tiers |
| 10 | |
| 11 | LLMs converge on a small set of safe patterns. The first thing the model reaches for is the **first-order reflex**: the genre-obvious tell. Replace it and the model reaches for the **second-order reflex**, the next-most-trained pattern that survives anti-AI guidance. |
| 12 | |
| 13 | Most AI-detection passes only check the first-order pattern. The result is "anti-AI" rewrites that still read like AI because the structural pass was never run. |
| 14 | |
| 15 | **Note on terminology**: this file uses **"first-order"** and **"second-order"** for the two detection passes. Elsewhere in the project, "Tier 1 / Tier 2 / Tier 3" refers to *source authority* (Google Search Central = Tier 1, Ahrefs = Tier 2, reputable industry sources = Tier 3). The two namespaces are intentionally kept separate; do not call the second-order detection pass "Tier 2." |
| 16 | |
| 17 | Examples of the same idea across both tiers: |
| 18 | |
| 19 | | Topic | First-order tell | Second-order tell that survives | |
| 20 | |---|---|---| |
| 21 | | SEO blog | "In today's digital landscape..." | Every H2 ends with a rhetorical question | |
| 22 | | SaaS post | "Game-changer," "Revolutionize" | Three-clause sentence rhythm, "While X, also Y" framings | |
| 23 | | How-to guide | "Dive into," "Unlock the potential" | Every step opens with an imperative verb identical in length | |
| 24 | | Listicle | "Cutting-edge," numbered fluff | Every item is ~80 words, identical structure | |
| 25 | | Thought leadership | "Comprehensive guide," "harness the power" | Hedge stack: "often," "typically," "may" within 20 words | |
| 26 | |
| 27 | The point: **a draft can score zero on a phrase blocklist and still be obviously AI.** |
| 28 | |
| 29 | |
| 30 | |
| 31 | ## First-order reflex (phrase + lexical) |
| 32 | |
| 33 | This is what the existing AI-detection in `blog-analyze` and `blog-rewrite` already covers. Documented here for completeness. |
| 34 | |
| 35 | **Trigger phrases** (full list in `agents/blog-reviewer.md` and `scripts/analyze_blog.py`): |
| 36 | |
| 37 | "In today's digital landscape" / "In the ever-evolving" |
| 38 | "It's important to note" / "It is worth mentioning" |
| 39 | "Dive into" / "deep dive" / "delve" |
| 40 | "Game-changer" / "Revolutionize" / "transformative" |
| 41 | "Cutting-edge" / "state-of-the-art" / "robust" |
| 42 | "Harness the power" / "Unlock the potential" |
| 43 | "Leverage" (as a verb, non-financial) |
| 44 | "Seamlessly" / "seamless integration" |
| 45 | "Tapestry" / "rich tapestry" / "multifaceted" |
| 46 | "Comprehensive guide" (in body text) |
| 47 | "Furthermore" / "Moreover" (transition overload) |
| 48 | Em dashes used as a stylistic flourish (any density) |
| 49 | |
| 50 | **Lexical signals**: |
| 51 | |
| 52 | AI trigger-word density > 5 per 1,000 words |
| 53 | Type-Token Ratio (TTR) below 0.40 on long-form |
| 54 | Burstiness (sentence-length standard deviation / mean) below 0.3 |
| 55 | |
| 56 | **Outcome of the first-order pass**: a "phrase-clean" draft. Necessary, not sufficient. |
| 57 | |
| 58 | |
| 59 | |
| 60 | ## Second-order reflex (structural + rhythmic) |
| 61 | |
| 62 | These are the patterns LLMs default to **after** the obvious vocabulary is replaced. They are structural and rhythmic, so a vocabulary swap doesn't fix them. Run this pass on drafts that already passed the first-order check. |
| 63 | |
| 64 | ### Structural tics to flag |
| 65 | |
| 66 | **Question-cadence H2s.** Every section heading is phrased as a question. Real long-form mixes question, statement, and noun-phrase headings. Flag if > 70% of H2 headings end with a question mark. |
| 67 | |
| 68 | **The Here opener.** A paragraph opens with the word "Here" ("Here's why...", "Here are five..."). Once is fine. Three or more in a 1,500-word post is an AI fingerprint. |
| 69 | |
| 70 | **Three-clause sentence rhythm.** Most sentences in a paragraph follow the structure `[clause], [clause], [clause].` The cadence is metronomic. Flag if > 50% of sentences in any 200-word window match this shape. |
| 71 | |
| 72 | **False-balance framing.** Repeated use of "While X, also Y" or "On one hand X, on the other Y" without a real contrast. The model uses it to feel even-handed but it adds no information. Flag if it appears more than twice per 1,000 words. |
| 73 | |
| 74 | **Hedge stacking.** Three or more hedges in a 20-word span ("It may often be the case that..."). Flag any 20-word window with > 2 of: may, might, often, typically, generally, usually, tend to, perhaps, somewhat, likely. |
| 75 | |
| 76 | **Symmetric list bloat.** Every item in a numbered or bulleted list is the same length within +/- 10 words and follows the same syntactic structure. Real lists vary; some items need one line, others need a paragraph. Flag if list-item length standard deviation < 5 words. |
| 77 | |
| 78 | **The wrap-up question.** Section ends with "What does this mean for [audience]?" or "Why does this matter?" Once per post is rhetorical; three or more is filler. |
| 79 | |
| 80 | **Capsule transitions.** Each H2 opener begins with a single-word transition ("First..." "Next..." "Additionally..." "Crucially..."). Real prose buries transitions inside sentences. Flag if > 50% of H2 openers start with a transition word. |
| 81 | |
| 82 | **The "key insight" tell.** The phrase "The key insight is..." or "What's important here is..." appears as a sentence-opener. This is the model telegraphing that it's about to summarize. Cut and let the sentence stand. |
| 83 | |
| 84 | **Listicle introduction bloat.** Before the actual list, three or more paragraphs of "context." Real listicles get to the list. Flag if > 250 words of pre-list intro. |
| 85 | |
| 86 | ### Rhythmic signals to compute |
| 87 | |
| 88 | **Sentence-length flatness within paragraphs.** Compute SD of sentence length per paragraph; flag any paragraph with internal SD < 4. |
| 89 | **Opening-word repetition.** Count first-word frequencies across all sentences. Flag if the top three first-words account for > 25% of all sentence openings. |
| 90 | **Paragraph-shape flatness.** Compute SD of paragraph word counts across the post; flag if < 25 (real long-form varies dramatically). |
| 91 | |
| 92 | |
| 93 | |
| 94 | ## How to run the two-tier check |
| 95 | |
| 96 | For `blog-rewrite` and `blog-reviewer`: |
| 97 | |
| 98 | Run the first-order pass first (phrase + lexical). If it fails, fix and re-run before moving on. |
| 99 | Once the first-order pass is clean, run the second-order pass. Report each second-order pattern with line numbers and an example. |
| 100 | Do not declare "AI-detection passed" unless both passes are clean. |
| 101 | |
| 102 | For `blog-write` (initial drafting): |
| 103 | |
| 104 | First-order is enforced at generation time via the persona's anti-phrase list. |
| 105 | Second-order is checked once on the full draft before delivery. |
| 106 | |
| 107 | |
| 108 | |
| 109 | ## Output format |
| 110 | |
| 111 | When reporting findings, use: |
| 112 | |
| 113 | |
| 114 | ## AI Slop Detection Report |
| 115 | |
| 116 | ### First-order (Phrase + Lexical) |
| 117 | - Trigger phrases: [N found] -> [list with line numbers] |
| 118 | - AI trigger words: [N/1K words], [pass/fail at ≤5] |
| 119 | - TTR: [score], [pass/fail at ≥0.40] |
| 120 | - Burstiness: [score], [pass/fail at ≥0.3] |
| 121 | |
| 122 | ### Second-order (Structural + Rhythmic) |
| 123 | - Question-cadence H2s: [X%], [pass/fail at ≤70%] |
| 124 | - "Here" openers: [N], [pass/fail at ≤2] |
| 125 | - Three-clause rhythm: [X%], [pass/fail at ≤50%] |
| 126 | - False-balance framings: [N/1K words], [pass/fail at ≤2] |
| 127 | - Hedge stacking: [N windows], [pass/fail at 0] |
| 128 | - Symmetric list bloat: [N lists], [pass/fail at 0] |
| 129 | - Wrap-up questions: [N], [pass/fail at ≤2] |
| 130 | - Capsule transitions on H2s: [X%], [pass/fail at ≤50%] |
| 131 | - "Key insight" sentence openers: [N], [pass/fail at 0] |
| 132 | - Listicle intro bloat: [pre-list words], [pass/fail at ≤250] |
| 133 | - Sentence-length flat paragraphs: [N], [pass/fail at 0] |
| 134 | - Opening-word repetition: [top-3 share], [pass/fail at ≤25%] |
| 135 | - Paragraph-shape SD: [value], [pass/fail at ≥25] |
| 136 | |
| 137 | ### Verdict |
| 138 | First-order: [PASS / FAIL] |
| 139 | Second-order: [PASS / FAIL] |
| 140 | Overall: [PASS only if both passes clean] |
| 141 | |
| 142 | |
| 143 | |
| 144 | |
| 145 | ## Why this matters for ranking + AI citations |
| 146 | |
| 147 | **Editorial quality risk**: content that lacks experience and original perspective is easier to classify as interchangeable consensus content. Second-order patterns are a warning sign, even when no specific Google update is cited. |
| 148 | **AI citations**: ChatGPT and Perplexity reward citable, distinctive passages. Second-order tics produce interchangeable prose that no AI surface has reason to prefer over the source it was trained on. |
| 149 | |
| 150 | The two-tier check is the editorial parallel to impeccable's "design slop" methodology: vocabulary-clean is necessary but not sufficient; structural distinctiveness is what separates citeable content from indexable filler. |
| 151 | |
| 152 | |
| 153 | |
| 154 | ## Attribution |
| 155 | |
| 156 | The two-tier first-order / second-order reflex methodology is adapted from the impeccable plugin v3.1.1 (Paul Bakaus, Apache 2.0, https://github.com/pbakaus/impeccable). The original applies it to UI design cliches ("observability -> dark blue"). This reference adapts the same mental model to prose. |
| 157 |
Discussion
Alternatives
Browse more free Claude skills or everything in Content creator.