TikTok Humanizer V3
Remove the AI-script tells viewers hear in a TikTok spoken script and caption: 2026 vocabulary by density, reveal bridges, staccato stacks, stacked triads, performed sincerity, written-not-spoken phrasing, "hey guys" filler.
How to use it
- Hit Copy the whole skill.
- Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
ChatGPT: make a Project and paste it into Instructions.
Neither? Paste it at the top of a new chat — it works for that chat. - Describe your job in plain words. The AI follows the skill from there.
npx degit sergebulaev/tiktok-skills/.codex-marketplace/tiktok-skills/skills/tt-humanizer#main ~/.claude/skills/tt-humanizerFor one project only, change the path to .claude/skills/tt-humanizer.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Show the full text300 lines
TikTok Humanizer V3
Rewrites a spoken script (and caption) to remove the AI tells that viewers hear, and audits a finished draft against the 2026 TikTok checklist before you film. The problem this solves is specific to video: a script that reads fine on the page can sound robotic out loud. Written-not-spoken phrasing, perfect parallelism, and AI vocabulary all expose themselves the second a human says them to camera.
Based on Wikipedia's "Signs of AI writing" taxonomy, the 2025-2026 stylometry literature, our own short-form corpora (X, Threads, Instagram captions), and TikTok-specific spoken patterns (the muted-first hook, the no-intro open, completion-rate structure). V3 (2026-09): recalibrated on 2026 evidence. Vocabulary is scored by density, em dashes are capped instead of banned, forced rhythm is now a tell instead of a fix, and there is an over-correction guard.
What this skill does not do: it does not make text "pass" GPTZero, Pangram, Turnitin or Originality. Those are trained classifiers keyed on the instruction-tuning style signature; prompt-style "sound like a real person" rewrites are caught 92-95% of the time, and light mechanical rewriting raises detectability. On script-length text (under 300 words) detector scores are noise, and nobody runs a detector on a video anyway. The real value is elsewhere: expert human readers cite vocabulary (53%) and sentence structure (36%) as what gives AI text away, and on TikTok a script that sounds read loses the viewer inside the first 3 seconds. This skill removes what those viewers react to.
What changed in V3
Evidence tier in brackets: [strong] = replicated across 2+ independent 2025-2026 studies or our own corpora; [vendor] = single platform or vendor dataset; [weak] = one study or expert-panel report.
- Vocabulary moved from a delete-list to density scoring. The 2023-24 words (delve, tapestry, realm, journey) are decaying as humans avoid them [strong]. The durable 2026 markers are common words (significant, crucial, notably, comprehensive, insights, robust, leverage, foster, landscape, nuanced, streamline, elevate) plus grammar: nominalisations and "-ing" clause openers at 5.3x the human rate [strong]. Spoken, they are worse: nobody says "leveraging" to a camera. One marker in a script beat is not a verdict. Three is.
- Em dash is no longer a tell. GPT-5.4 emits 1.43 per 1,000 words, below
the 3.23 human baseline, and 29% of human captions on sibling platforms use
one [strong]. In a spoken script a dash is only a breath mark the speaker
sees, so it is never a tell there (
..reads better on a teleprompter). In the caption: cap at about 1 per 100 words. On an on-screen card (3-7 words): at most one, and a card rarely needs one. Replace the excess with a comma, colon,..or a line break. Never a period. - Forced burstiness is the #1 2026 tell, not the fix. Mechanical long/short alternation is a learnable humanizer fingerprint [weak], and "Short. Punchy. Done.", "No X. No Y. Just Z.", one-word lines for drama and "The result?" reveals are the current top reader-cited tells [strong]. Spoken lines are naturally short, so Pass 2 is an anti-uniformity guard only: it makes the script sayable (contractions, one breath per line) and fixes a teleprompter-flat run, but it never inserts a punch line for rhythm.
- Rule of three is still a tell, at density. Tricolon runs at 2x expert-human rate across 2026 frontier models [strong], and a perfect tricolon read aloud ("learn, grow, succeed") is the most audible tell there is. Stacked, perfectly parallel or hollow triads get scrubbed. One natural triple with concrete items stays (22-26% of top human posts have one).
- Fingerprint injection was half wrong. Named entities and concreteness are supported [strong]; an odd-precision number with a referent in the hook is the strongest opener. Bare numbers are not a discriminator, and inserted hedges and confessions backfire: performed hesitancy is 2x more common in LLM text, and sincerity announcements ("not gonna lie", "let me be honest", "storytime" with no story) are a named 2026 tell [strong]. Pass 3 asks for a flat, dated, uncomfortable fact instead.
- Over-correction guard. Humanizer output has its own fingerprint [weak]. Pass 4 checks whether Passes 1-3 introduced the very patterns they were meant to remove. Edits are proportional to real problems. When in doubt, leave it.
When to use
- Before filming any AI-drafted spoken script (rewrite mode)
- Pre-film review of a finished script + caption (audit mode, see
sub-skills/post-audit.md) - When a script "reads fine but sounds off" when you say it out loud
Input
A spoken script (the hook line plus the body), optionally the caption, and optionally voice samples (the user's past scripts or how they actually talk).
Output
- Rewritten script that sounds spoken, not written
- A diff showing what changed and why
- Caption char count (flagging over 2,200) when a caption is included
- Per-beat tell density (markers per script beat or caption paragraph; 3+ triggered a rewrite)
- Reader-read confidence: "sounds human", "mixed", "sounds read" (a viewer-tell estimate, not a detector score)
Modes
# Default: scrub AI tells (forensic + strict) and fix spoken-word issues
tt-humanizer <script>
# Forensic only - minimum touch, just kill model leakage
tt-humanizer --mode forensic <script>
# Audit - detection-only pass-fail review, no rewrite
# Runs the 2026 TikTok pre-film checklist: first 1-3 second hook strength,
# muted-first text, completion design, caption fit, hashtag and settings sanity.
# Returns Blockers + Warnings + suggested fixes. See sub-skills/post-audit.md.
tt-humanizer --mode audit <script>
# Profile - build/update the user's Voice & Brand Profile. See the section below.
tt-humanizer --mode profile
The four passes
Pass 1 - SCRUB (score, then delete or replace)
Apply the tiered catalogs in references/scrub-rules.md. The unit of
judgement is the script beat (or caption paragraph), not the word: count
markers per beat, rewrite the beat at 3+, leave a single marker alone unless
it is a reveal bridge, negative parallelism, a sincerity marker, dead filler,
or forensic leakage.
- Forensic (always on): real model leakage no human says. AI tool markers (oaicite, contentReference, turn0search0), knowledge-cutoff disclaimers ("As of my last update"), template blanks ([Your Name]), chat wrappers ("Certainly!", "I hope this helps"), and em dashes above the cap in the caption or on an on-screen card.
- Strict (default on): what viewers hear. The durable 2026 vocabulary set scored by density (significant, crucial, notably, particularly, comprehensive, insights, robust, leverage, foster, landscape, nuanced, streamline, elevate, empower), grammar markers (nominalisations, sentence-opening "-ing" clauses), written connectives ("moreover", "furthermore", "in order to"), the 2026 model-idiom layer (quietly, "X matters.", compound, "a signal", "the work", "built different", "let that sink in"), reveal bridges on a single hit ("The result?", "Here's what nobody tells you", "Stop X, start Y", "plot twist:"), all forms of negative parallelism, stacked or perfectly parallel triads, dead filler ("hey guys", "without further ado", "in this video I will"), and dead closers, both spoken ("thanks for watching", "don't forget to subscribe") and caption-level ("What do you think?", "Drop your thoughts below").
- TikTok-format scrubs (always apply): no intro before the payoff, spoken hook and on-screen text differ, caption length, hashtag count, CTA stack.
Pass 2 - RHYTHM (make it sayable, never manufactured)
Detectors do not score burstiness, and spoken lines are naturally short, so Pass 2 has nothing to "vary". Its jobs are: make the script sound spoken, remove manufactured drama-rhythm, and un-flatten only a run that reads teleprompter-flat. It never adds a punch line as a tactic.
- Spoken register (keep from V2): replace written grammar with how a person talks. Contractions, natural fragments, one breath per line. "It is something that you should consider" becomes "you should try this". This is register, not rhythm; it applies to every line.
- Read-aloud test: flag any line that needs two breaths or trips the tongue. Split at the natural breath, never at a dramatic pause.
- Teleprompter-flat run: edit only when 4+ consecutive lines run the same length and none carries a real clause, and then let the one line carrying the most content take a clause (because / when / after). Never insert a short punch line between long ones; the inserted punch is the humanizer fingerprint.
- Banned outright (rewrite as a spoken sentence): "The X? Y." reveals; "No X. No Y. Just Z."; "All the X. None of the Y."; "Simple. Effective. Easy." adjective stacks; one-word lines for drama ("Still." "Exactly."); pseudo- Socratic Q&A ("Why? Because..."); "Short. Punchy. Done." staccato runs. Fragment runs are the tell, on the page and out loud.
- Natural spoken fragments ("three takes. that's it.") are register and stay. A run of them staged for drama is the tell. In the caption, cap standalone fragments at 2.
- Never alternate long/short/long/short across the script. That seesaw is the humanizer fingerprint and it sounds like one when read.
The check is "would a person say this, and did I add a staccato pattern", not a variance number.
Pass 3 - ADD (human fingerprints)
Require where the content allows:
- One odd-precision number WITH a named referent in the hook: who, what, when, or what it cost ("47 minutes on the third take", "$12 at the hardware store", not "a few takes" and not "47"). A bare number is not a fingerprint; the referent carries the signal.
- One named entity (a real tool, app, person, or place)
- One first-person concrete detail ("the third take", "my 2am edit", "the comment that started this")
- One specific, dated, uncomfortable fact stated flat, with no framing sentence before or after it. Not "not gonna lie, this one hurt: the client fired us." Just "the client fired us on a Tuesday, 9 hours before the demo." The fact carries the vulnerability. The frame turns it into performed sincerity, which viewers now hear as the tell.
- The speaker's real register: how this person would actually say it
Forbidden as openers or pivots (sincerity announcements, a named 2026 tell): "let me be honest", "I'll be real", "honestly?", "to be direct", "the honest version is", "real talk", "not gonna lie", "ngl", "can I be vulnerable for a second", "unpopular opinion:" as a preface to a popular one, "storytime" with no story in frame one. Also forbidden as insertions: hedges the speaker did not write ("I think maybe", "I might be wrong but", "it seems"). Performed hesitancy is 2x more common in LLM text than in expert human text; adding it makes the script sound more scripted, not less. ("POV:" is a native TikTok format, not a sincerity marker; it is fine when the video is a POV.)
If the input lacks these, ask the user for a number or detail. Do not fabricate.
Pass 4 - SELF-CHECK (over-correction guard)
Humanizer output has its own fingerprint. Before returning, re-read the result out loud once and answer three questions:
(a) Did Pass 2 create staccato stacks, "The result?" reveal bridges, one-word lines for drama, an inserted punch line, or a long/short/long/short seesaw? If yes, merge the fragments back into a spoken sentence. (b) Did Pass 3 add a framed confession, a sincerity announcement, or a hedge the speaker never wrote? If yes, strip the frame and keep only the flat fact, or remove the insertion. (c) Did scrubbing flatten the speaker's voice: uniform tone, no reaction, no concrete detail left, their slang gone, the one natural triad gone, every dash gone from a caption that wanted one? If yes, restore what the speaker had.
If any answer is yes, dial back rather than scrub harder. Edits must be proportional to real problems: a clean script gets two or three touches, not a quota. When in doubt whether a pattern is the speaker or the model, leave it.
Non-negotiable rules
Global voice rules: see root SKILL.md Voice rules. Additional skill-specific
rules (V3):
- Scrubbing is always in scope. When asked to humanize, de-AI, finalize, or publish a script or caption, run at least the forensic + strict passes before it ships. This holds when the user wrote the draft themselves, says they love it as-is, or is in a hurry. Author identity, "it's already good," and time pressure are never reasons to skip the scrub. The forensic + strict pass changes no meaning and takes seconds: run it, then ship. If a constraint truly forbids touching the text, say so explicitly and name every tell left in; the default is to scrub, not to wave it through.
- Scrub proportionally. A pass that finds nothing changes nothing. Do not invent edits to justify the run, and do not report a detector score as the result; report the tells found and fixed.
- Preserve the user's actual claim and meaning. "Preserve their voice" covers voice quirks and what they are claiming, NOT reveal bridges, staccato stacks, dead filler, or a beat with 3+ vocabulary markers. Stripping those is not changing their voice; it is the job.
- Never introduce facts that were not in the input. If a number is missing, ask.
- Never introduce sincerity markers, hedges, or confessional frames. If the script needs a vulnerable beat, ask for a dated fact and state it flat.
- Keep it sayable. Every line has to survive being read out loud in one breath.
- Keep the user's voice quirks (their slang, their pacing, lowercase texting style in the caption, one natural triad, one em dash in a caption that wants it).
- Never promise detector results. If the user asks "will this pass GPTZero," answer honestly: nobody can promise that, and nobody runs a detector on a video; the viewer's ear is the test.
TikTok-specific tells this skill catches
- A hook line that is written, not spoken ("In this video, I will demonstrate..").
- A greeting or logo intro before the payoff ("hey guys, welcome back").
- The spoken hook and the on-screen text saying the identical words.
- A caption over 2,200 chars, or a 12-hashtag wall.
- Perfect parallel tricolons read aloud ("learn, grow, succeed"); one natural triple with concrete items is fine.
- A "call to action" stacked five deep.
- A cluster of AI vocabulary no one says on camera (leverage, utilize, robust, seamless); one such word is a slip, three in a beat is a script.
- Staccato drama ("No script. No plan. Just vibes.") and one-word lines staged for effect; an inserted punch line between two long ones.
- "Not gonna lie" / "storytime" framing around what should be a plain fact.
Example
See references/examples.md for worked before/after rewrites of spoken scripts.
Files
SKILL.md- this file (rewrite scrubber + audit-mode entry)references/scrub-rules.md- V3 catalogs by tier, density scoring, em dash cap, spoken-word fixes, rhythm rules, forbidden insertionsreferences/examples.md- worked before/after script rewritesreferences/audit-checklist.md- the pre-film checklist with thresholdssub-skills/post-audit.md- pre-film audit workflow (detection-only, no rewrite)sub-skills/voice-profile.md- build/update the user's Voice & Brand Profile (--mode profile)sub-skills/illustration.md- optional Pixfaro image workflow
Voice profile mode (--mode profile)
tt-humanizer --mode profile builds or updates the user's Voice & Brand Profile at ../../references/voice-profile.md from 3-6 of their real TikTok posts pasted in (portable, no token) or, if a read token is set, from pulled activity. Once filled, every writing skill in this bundle drafts in the user's voice automatically. See sub-skills/voice-profile.md. Triggers: "build my voice profile", "learn my voice".
Related skills
tt-hook-scripter- generates hooks that already pass the humanizertt-caption-writer- generates captions that already pass the humanizer
| 1 | |
| 2 | name tt-humanizer |
| 3 | description "Remove the AI-script tells viewers hear in a TikTok spoken script and caption: 2026 vocabulary by density, reveal bridges, staccato stacks, stacked triads, performed sincerity, written-not-spoken phrasing, \"hey guys\" filler; caps em dashes. Includes --mode audit pre-film check (hook, completion design, caption fit) and --mode profile. Not for beating AI detectors (no edit reliably does). Not for writing from scratch (use tt-hook-scripter). Keywords: humanize script, de-AI, audit before filming." |
| 4 | |
| 5 | |
| 6 | # TikTok Humanizer V3 |
| 7 | |
| 8 | Rewrites a spoken script (and caption) to remove the AI tells that viewers |
| 9 | hear, and audits a finished draft against the 2026 TikTok checklist before you |
| 10 | film. The problem this solves is specific to video: a script that reads fine |
| 11 | on the page can sound robotic out loud. Written-not-spoken phrasing, perfect |
| 12 | parallelism, and AI vocabulary all expose themselves the second a human says |
| 13 | them to camera. |
| 14 | |
| 15 | Based on Wikipedia's "Signs of AI writing" taxonomy, the 2025-2026 stylometry |
| 16 | literature, our own short-form corpora (X, Threads, Instagram captions), and |
| 17 | TikTok-specific spoken patterns (the muted-first hook, the no-intro open, |
| 18 | completion-rate structure). **V3 (2026-09):** recalibrated on 2026 evidence. |
| 19 | Vocabulary is scored by density, em dashes are capped instead of banned, |
| 20 | forced rhythm is now a tell instead of a fix, and there is an over-correction |
| 21 | guard. |
| 22 | |
| 23 | **What this skill does not do:** it does not make text "pass" GPTZero, |
| 24 | Pangram, Turnitin or Originality. Those are trained classifiers keyed on the |
| 25 | instruction-tuning style signature; prompt-style "sound like a real person" |
| 26 | rewrites are caught 92-95% of the time, and light mechanical rewriting raises |
| 27 | detectability. On script-length text (under 300 words) detector scores are |
| 28 | noise, and nobody runs a detector on a video anyway. The real value is |
| 29 | elsewhere: expert human readers cite vocabulary (53%) and sentence structure |
| 30 | (36%) as what gives AI text away, and on TikTok a script that sounds read |
| 31 | loses the viewer inside the first 3 seconds. This skill removes what those |
| 32 | viewers react to. |
| 33 | |
| 34 | ## What changed in V3 |
| 35 | |
| 36 | Evidence tier in brackets: [strong] = replicated across 2+ independent |
| 37 | 2025-2026 studies or our own corpora; [vendor] = single platform or vendor |
| 38 | dataset; [weak] = one study or expert-panel report. |
| 39 | |
| 40 | **Vocabulary moved from a delete-list to density scoring.** The 2023-24 words |
| 41 | (delve, tapestry, realm, journey) are decaying as humans avoid them [strong]. |
| 42 | The durable 2026 markers are common words (significant, crucial, notably, |
| 43 | comprehensive, insights, robust, leverage, foster, landscape, nuanced, |
| 44 | streamline, elevate) plus grammar: nominalisations and "-ing" clause openers |
| 45 | at 5.3x the human rate [strong]. Spoken, they are worse: nobody says |
| 46 | "leveraging" to a camera. One marker in a script beat is not a verdict. |
| 47 | Three is. |
| 48 | **Em dash is no longer a tell.** GPT-5.4 emits 1.43 per 1,000 words, below |
| 49 | the 3.23 human baseline, and 29% of human captions on sibling platforms use |
| 50 | one [strong]. In a spoken script a dash is only a breath mark the speaker |
| 51 | sees, so it is never a tell there (`..` reads better on a teleprompter). In |
| 52 | the caption: cap at about 1 per 100 words. On an on-screen card (3-7 words): |
| 53 | at most one, and a card rarely needs one. Replace the excess with a comma, |
| 54 | colon, `..` or a line break. Never a period. |
| 55 | **Forced burstiness is the #1 2026 tell, not the fix.** Mechanical |
| 56 | long/short alternation is a learnable humanizer fingerprint [weak], and |
| 57 | "Short. Punchy. Done.", "No X. No Y. Just Z.", one-word lines for drama and |
| 58 | "The result?" reveals are the current top reader-cited tells [strong]. |
| 59 | Spoken lines are naturally short, so Pass 2 is an anti-uniformity guard |
| 60 | only: it makes the script sayable (contractions, one breath per line) and |
| 61 | fixes a teleprompter-flat run, but it never inserts a punch line for |
| 62 | rhythm. |
| 63 | **Rule of three is still a tell, at density.** Tricolon runs at 2x |
| 64 | expert-human rate across 2026 frontier models [strong], and a perfect |
| 65 | tricolon read aloud ("learn, grow, succeed") is the most audible tell there |
| 66 | is. Stacked, perfectly parallel or hollow triads get scrubbed. One natural |
| 67 | triple with concrete items stays (22-26% of top human posts have one). |
| 68 | **Fingerprint injection was half wrong.** Named entities and concreteness are |
| 69 | supported [strong]; an odd-precision number with a referent in the hook is |
| 70 | the strongest opener. Bare numbers are not a discriminator, and inserted |
| 71 | hedges and confessions backfire: performed hesitancy is 2x more common in |
| 72 | LLM text, and sincerity announcements ("not gonna lie", "let me be honest", |
| 73 | "storytime" with no story) are a named 2026 tell [strong]. Pass 3 asks for |
| 74 | a flat, dated, uncomfortable fact instead. |
| 75 | **Over-correction guard.** Humanizer output has its own fingerprint [weak]. |
| 76 | Pass 4 checks whether Passes 1-3 introduced the very patterns they were meant |
| 77 | to remove. Edits are proportional to real problems. When in doubt, leave it. |
| 78 | |
| 79 | ## When to use |
| 80 | |
| 81 | Before filming any AI-drafted spoken script (rewrite mode) |
| 82 | Pre-film review of a finished script + caption (audit mode, see |
| 83 | `sub-skills/post-audit.md`) |
| 84 | When a script "reads fine but sounds off" when you say it out loud |
| 85 | |
| 86 | ## Input |
| 87 | |
| 88 | A spoken script (the hook line plus the body), optionally the caption, and |
| 89 | optionally voice samples (the user's past scripts or how they actually talk). |
| 90 | |
| 91 | ## Output |
| 92 | |
| 93 | Rewritten script that sounds spoken, not written |
| 94 | A diff showing what changed and why |
| 95 | Caption char count (flagging over 2,200) when a caption is included |
| 96 | Per-beat tell density (markers per script beat or caption paragraph; 3+ |
| 97 | triggered a rewrite) |
| 98 | Reader-read confidence: "sounds human", "mixed", "sounds read" (a |
| 99 | viewer-tell estimate, not a detector score) |
| 100 | |
| 101 | ## Modes |
| 102 | |
| 103 | |
| 104 | # Default: scrub AI tells (forensic + strict) and fix spoken-word issues |
| 105 | tt-humanizer <script> |
| 106 | |
| 107 | # Forensic only - minimum touch, just kill model leakage |
| 108 | tt-humanizer --mode forensic <script> |
| 109 | |
| 110 | # Audit - detection-only pass-fail review, no rewrite |
| 111 | # Runs the 2026 TikTok pre-film checklist: first 1-3 second hook strength, |
| 112 | # muted-first text, completion design, caption fit, hashtag and settings sanity. |
| 113 | # Returns Blockers + Warnings + suggested fixes. See sub-skills/post-audit.md. |
| 114 | tt-humanizer --mode audit <script> |
| 115 | |
| 116 | # Profile - build/update the user's Voice & Brand Profile. See the section below. |
| 117 | tt-humanizer --mode profile |
| 118 | |
| 119 | |
| 120 | ## The four passes |
| 121 | |
| 122 | ### Pass 1 - SCRUB (score, then delete or replace) |
| 123 | |
| 124 | Apply the tiered catalogs in `references/scrub-rules.md`. The unit of |
| 125 | judgement is the **script beat (or caption paragraph), not the word**: count |
| 126 | markers per beat, rewrite the beat at 3+, leave a single marker alone unless |
| 127 | it is a reveal bridge, negative parallelism, a sincerity marker, dead filler, |
| 128 | or forensic leakage. |
| 129 | |
| 130 | **Forensic** (always on): real model leakage no human says. AI tool markers |
| 131 | (oaicite, contentReference, turn0search0), knowledge-cutoff disclaimers ("As of |
| 132 | my last update"), template blanks ([Your Name]), chat wrappers ("Certainly!", |
| 133 | "I hope this helps"), and em dashes above the cap in the caption or on an |
| 134 | on-screen card. |
| 135 | **Strict** (default on): what viewers hear. The durable 2026 vocabulary set |
| 136 | scored by density (significant, crucial, notably, particularly, |
| 137 | comprehensive, insights, robust, leverage, foster, landscape, nuanced, |
| 138 | streamline, elevate, empower), grammar markers (nominalisations, |
| 139 | sentence-opening "-ing" clauses), written connectives ("moreover", |
| 140 | "furthermore", "in order to"), the 2026 model-idiom layer (quietly, "X |
| 141 | matters.", compound, "a signal", "the work", "built different", "let that |
| 142 | sink in"), reveal bridges on a single hit ("The result?", "Here's what |
| 143 | nobody tells you", "Stop X, start Y", "plot twist:"), all forms of negative |
| 144 | parallelism, stacked or perfectly parallel triads, dead filler ("hey guys", |
| 145 | "without further ado", "in this video I will"), and dead closers, both |
| 146 | spoken ("thanks for watching", "don't forget to subscribe") and |
| 147 | caption-level ("What do you think?", "Drop your thoughts below"). |
| 148 | **TikTok-format scrubs** (always apply): no intro before the payoff, spoken |
| 149 | hook and on-screen text differ, caption length, hashtag count, CTA stack. |
| 150 | |
| 151 | ### Pass 2 - RHYTHM (make it sayable, never manufactured) |
| 152 | |
| 153 | Detectors do not score burstiness, and spoken lines are naturally short, so |
| 154 | Pass 2 has nothing to "vary". Its jobs are: make the script sound spoken, |
| 155 | remove manufactured drama-rhythm, and un-flatten only a run that reads |
| 156 | teleprompter-flat. It never adds a punch line as a tactic. |
| 157 | |
| 158 | **Spoken register (keep from V2):** replace written grammar with how a |
| 159 | person talks. Contractions, natural fragments, one breath per line. "It is |
| 160 | something that you should consider" becomes "you should try this". This is |
| 161 | register, not rhythm; it applies to every line. |
| 162 | **Read-aloud test:** flag any line that needs two breaths or trips the |
| 163 | tongue. Split at the natural breath, never at a dramatic pause. |
| 164 | **Teleprompter-flat run:** edit only when 4+ consecutive lines run the same |
| 165 | length and none carries a real clause, and then let the one line carrying |
| 166 | the most content take a clause (because / when / after). Never insert a |
| 167 | short punch line between long ones; the inserted punch is the humanizer |
| 168 | fingerprint. |
| 169 | Banned outright (rewrite as a spoken sentence): "The X? Y." reveals; "No X. |
| 170 | No Y. Just Z."; "All the X. None of the Y."; "Simple. Effective. Easy." |
| 171 | adjective stacks; one-word lines for drama ("Still." "Exactly."); pseudo- |
| 172 | Socratic Q&A ("Why? Because..."); "Short. Punchy. Done." staccato runs. |
| 173 | Fragment runs are the tell, on the page and out loud. |
| 174 | Natural spoken fragments ("three takes. that's it.") are register and stay. |
| 175 | A run of them staged for drama is the tell. In the caption, cap standalone |
| 176 | fragments at 2. |
| 177 | Never alternate long/short/long/short across the script. That seesaw is the |
| 178 | humanizer fingerprint and it sounds like one when read. |
| 179 | |
| 180 | The check is "would a person say this, and did I add a staccato pattern", |
| 181 | not a variance number. |
| 182 | |
| 183 | ### Pass 3 - ADD (human fingerprints) |
| 184 | |
| 185 | Require where the content allows: |
| 186 | One odd-precision number WITH a named referent in the hook: who, what, |
| 187 | when, or what it cost ("47 minutes on the third take", "$12 at the hardware |
| 188 | store", not "a few takes" and not "47"). A bare number is not a fingerprint; |
| 189 | the referent carries the signal. |
| 190 | One named entity (a real tool, app, person, or place) |
| 191 | One first-person concrete detail ("the third take", "my 2am edit", "the |
| 192 | comment that started this") |
| 193 | One specific, dated, uncomfortable fact stated flat, with no framing |
| 194 | sentence before or after it. Not "not gonna lie, this one hurt: the client |
| 195 | fired us." Just "the client fired us on a Tuesday, 9 hours before the demo." |
| 196 | The fact carries the vulnerability. The frame turns it into performed |
| 197 | sincerity, which viewers now hear as the tell. |
| 198 | The speaker's real register: how this person would actually say it |
| 199 | |
| 200 | Forbidden as openers or pivots (sincerity announcements, a named 2026 tell): |
| 201 | "let me be honest", "I'll be real", "honestly?", "to be direct", "the honest |
| 202 | version is", "real talk", "not gonna lie", "ngl", "can I be vulnerable for a |
| 203 | second", "unpopular opinion:" as a preface to a popular one, "storytime" with |
| 204 | no story in frame one. Also forbidden as insertions: hedges the speaker did |
| 205 | not write ("I think maybe", "I might be wrong but", "it seems"). Performed |
| 206 | hesitancy is 2x more common in LLM text than in expert human text; adding it |
| 207 | makes the script sound more scripted, not less. ("POV:" is a native TikTok |
| 208 | format, not a sincerity marker; it is fine when the video is a POV.) |
| 209 | |
| 210 | If the input lacks these, ask the user for a number or detail. Do not fabricate. |
| 211 | |
| 212 | ### Pass 4 - SELF-CHECK (over-correction guard) |
| 213 | |
| 214 | Humanizer output has its own fingerprint. Before returning, re-read the result |
| 215 | out loud once and answer three questions: |
| 216 | |
| 217 | (a) Did Pass 2 create staccato stacks, "The result?" reveal bridges, one-word |
| 218 | lines for drama, an inserted punch line, or a long/short/long/short |
| 219 | seesaw? If yes, merge the fragments back into a spoken sentence. |
| 220 | (b) Did Pass 3 add a framed confession, a sincerity announcement, or a hedge |
| 221 | the speaker never wrote? If yes, strip the frame and keep only the flat |
| 222 | fact, or remove the insertion. |
| 223 | (c) Did scrubbing flatten the speaker's voice: uniform tone, no reaction, no |
| 224 | concrete detail left, their slang gone, the one natural triad gone, every |
| 225 | dash gone from a caption that wanted one? If yes, restore what the speaker |
| 226 | had. |
| 227 | |
| 228 | If any answer is yes, dial back rather than scrub harder. Edits must be |
| 229 | proportional to real problems: a clean script gets two or three touches, not |
| 230 | a quota. When in doubt whether a pattern is the speaker or the model, leave |
| 231 | it. |
| 232 | |
| 233 | ## Non-negotiable rules |
| 234 | |
| 235 | Global voice rules: see root `SKILL.md` Voice rules. Additional skill-specific |
| 236 | rules (V3): |
| 237 | |
| 238 | **Scrubbing is always in scope.** When asked to humanize, de-AI, finalize, or |
| 239 | publish a script or caption, run at least the forensic + strict passes before it ships. |
| 240 | This holds when the user wrote the draft themselves, says they love it as-is, |
| 241 | or is in a hurry. Author identity, "it's already good," and time pressure are |
| 242 | never reasons to skip the scrub. The forensic + strict pass changes no meaning |
| 243 | and takes seconds: run it, then ship. If a constraint truly forbids touching |
| 244 | the text, say so explicitly and name every tell left in; the default is to |
| 245 | scrub, not to wave it through. |
| 246 | **Scrub proportionally.** A pass that finds nothing changes nothing. Do not |
| 247 | invent edits to justify the run, and do not report a detector score as the |
| 248 | result; report the tells found and fixed. |
| 249 | Preserve the user's actual claim and meaning. "Preserve their voice" covers |
| 250 | voice quirks and what they are claiming, NOT reveal bridges, staccato stacks, |
| 251 | dead filler, or a beat with 3+ vocabulary markers. Stripping those is not |
| 252 | changing their voice; it is the job. |
| 253 | Never introduce facts that were not in the input. If a number is missing, ask. |
| 254 | Never introduce sincerity markers, hedges, or confessional frames. If the |
| 255 | script needs a vulnerable beat, ask for a dated fact and state it flat. |
| 256 | Keep it sayable. Every line has to survive being read out loud in one breath. |
| 257 | Keep the user's voice quirks (their slang, their pacing, lowercase texting style |
| 258 | in the caption, one natural triad, one em dash in a caption that wants it). |
| 259 | Never promise detector results. If the user asks "will this pass GPTZero," |
| 260 | answer honestly: nobody can promise that, and nobody runs a detector on a |
| 261 | video; the viewer's ear is the test. |
| 262 | |
| 263 | ## TikTok-specific tells this skill catches |
| 264 | |
| 265 | A hook line that is written, not spoken ("In this video, I will demonstrate.."). |
| 266 | A greeting or logo intro before the payoff ("hey guys, welcome back"). |
| 267 | The spoken hook and the on-screen text saying the identical words. |
| 268 | A caption over 2,200 chars, or a 12-hashtag wall. |
| 269 | Perfect parallel tricolons read aloud ("learn, grow, succeed"); one natural |
| 270 | triple with concrete items is fine. |
| 271 | A "call to action" stacked five deep. |
| 272 | A cluster of AI vocabulary no one says on camera (leverage, utilize, robust, |
| 273 | seamless); one such word is a slip, three in a beat is a script. |
| 274 | Staccato drama ("No script. No plan. Just vibes.") and one-word lines |
| 275 | staged for effect; an inserted punch line between two long ones. |
| 276 | "Not gonna lie" / "storytime" framing around what should be a plain fact. |
| 277 | |
| 278 | ## Example |
| 279 | |
| 280 | See `references/examples.md` for worked before/after rewrites of spoken scripts. |
| 281 | |
| 282 | ## Files |
| 283 | |
| 284 | `SKILL.md` - this file (rewrite scrubber + audit-mode entry) |
| 285 | `references/scrub-rules.md` - V3 catalogs by tier, density scoring, em dash cap, spoken-word fixes, rhythm rules, forbidden insertions |
| 286 | `references/examples.md` - worked before/after script rewrites |
| 287 | `references/audit-checklist.md` - the pre-film checklist with thresholds |
| 288 | `sub-skills/post-audit.md` - pre-film audit workflow (detection-only, no rewrite) |
| 289 | `sub-skills/voice-profile.md` - build/update the user's Voice & Brand Profile (`--mode profile`) |
| 290 | `sub-skills/illustration.md` - optional Pixfaro image workflow |
| 291 | |
| 292 | ## Voice profile mode (`--mode profile`) |
| 293 | |
| 294 | `tt-humanizer --mode profile` builds or updates the user's Voice & Brand Profile at `../../references/voice-profile.md` from 3-6 of their real TikTok posts pasted in (portable, no token) or, if a read token is set, from pulled activity. Once filled, every writing skill in this bundle drafts in the user's voice automatically. See `sub-skills/voice-profile.md`. Triggers: "build my voice profile", "learn my voice". |
| 295 | |
| 296 | ## Related skills |
| 297 | |
| 298 | `tt-hook-scripter` - generates hooks that already pass the humanizer |
| 299 | `tt-caption-writer` - generates captions that already pass the humanizer |
| 300 |