Voice Fingerprint — Preserving the user's voice while scrubbing skill
The humanizer is destructive by design.
by sergebulaev·MIT license·★ 3,059 Stars on the repo·GitHub ↗
Files of Voice Fingerprint — Preserving the user's voice while scrubbing
sergebulaev/
Show the full text119 lines
Voice Fingerprint — Preserving the user's voice while scrubbing
The humanizer is destructive by design. Every pass deletes or substitutes tokens. That's fine for AI tells. It's a bug for the user's actual voice.
This file lists the signals to preserve, even when they overlap with rules in scrub-rules.md.
Contents
- Preserve unconditionally (do NOT scrub these)
- Preserve when sample voice is provided
- Conflict resolution
- How to build a voice fingerprint from samples (sketch)
- Examples
- Don't fabricate
Preserve unconditionally (do NOT scrub these)
These are voice signatures, not AI tells. Leave them alone in every tier including --mode all.
| Pattern | Why it's voice, not AI |
|---|---|
Lowercase sentence starts (closed our seed on a tuesday...) |
Users like Serge use this as a deliberate cadence cue. Capitalizing flattens their voice. |
.. as a soft pause |
This is the humanizer's officially-blessed alternative to em dash. Removing it has nowhere to go. |
Sentence fragments (Worth it., Every time., Not even close.) |
Pass 2 ADDS fragments. Don't remove the ones already there. |
Contractions (don't, it's, you're, we're) |
Mandatory for natural rhythm. Scrubbing curly apostrophes is fine; expanding contractions is not. |
First-person sensory detail (my hands shook, the room went quiet) |
Pass 3 demands these. Never strip. |
Specific numbers ($47k, 9:14am, 47 days) |
Pass 3 demands these. Never strip. |
Named entities (HubSpot, Tuesday morning, brand names) |
Pass 3 demands these. Capitalize properly per non-negotiable rule. |
Self-correction within a paragraph (actually no, correction:) |
Burstiness signal. Real humans circle back. |
Preserve when sample voice is provided
If the user passes optional target_voice_samples (their last 5-10 LinkedIn posts), extract:
- Sentence-length distribution. If they routinely write 4-6 word sentences, don't force 12+ word "minimum lengths" on Pass 2.
- Vocabulary fingerprint. Words they use 3+ times across samples are part of their voice — even if those words appear on the strict blacklist. Flag for user review rather than auto-substituting.
- Punctuation habits. Some users use
...instead of.., or unbroken comma chains. Match the dominant pattern. - Opener patterns. If they always start with a number (
47 days ago,$2M ARR) or a name (Jake said), preserve that template. - Closer patterns. If they always close with a single fragment + period (no question), don't force a question CTA.
Conflict resolution
When a scrub rule fires on a token that's also in the user's voice fingerprint:
| Tier | Behavior |
|---|---|
| Forensic | Always scrub. Forensic rules catch model leakage; if the user's voice fingerprint contains oaicite it's because they pasted AI output. |
| Strict | Flag for user review. Don't auto-substitute. The user gets to decide. |
| Aesthetic | Skip the rule entirely. Aesthetic rules already explicitly tolerate human-writer defenses. |
How to build a voice fingerprint from samples (sketch)
from collections import Counter
import re
def build_voice_fingerprint(samples: list[str]) -> dict:
text = "\n".join(samples)
sentences = re.split(r'(?<=[.!?])\s+', text)
return {
"sentence_lengths": [len(s.split()) for s in sentences],
"vocab_freq": Counter(re.findall(r"\b[a-z][a-z']{2,}\b", text.lower())),
"starts_lowercase_pct": sum(1 for s in sentences if s and s[0].islower()) / max(len(sentences), 1),
"uses_double_dot": ".." in text,
"uses_triple_dot": "..." in text,
"punctuation_freq": Counter(c for c in text if c in ".!?,;:"),
"fragment_pct": sum(1 for s in sentences if len(s.split()) <= 4) / max(len(sentences), 1),
}
The skill should call this on target_voice_samples before running Pass 1.
Examples
Example 1 — .. as soft pause (preserve)
Input: closed our seed.. then everything broke
Wrong (scrubs the ..): closed our seed. then everything broke
Right (preserve): closed our seed.. then everything broke
The .. is on the explicit preserve list. Period substitution is for --, not ...
Example 2 — lowercase start (preserve)
Input: closed our seed on a tuesday morning at 9:14am
Wrong (capitalizes): Closed our seed on a Tuesday morning at 9:14am
Right (preserve closed, capitalize Tuesday): closed our seed on a Tuesday morning at 9:14am
The non-negotiable rule says capitalize NAMES — Tuesday is a proper noun in date context, but the sentence-initial closed stays lowercase per voice rule.
Example 3. Voice-fingerprint vocabulary collision (flag, don't substitute)
User samples contain harness 4 times across 6 posts (clearly part of their voice — they work in horse-training tech).
Strict tier scrub rule says: harness → use.
Right behavior: flag for user review. Output: [VOICE-CONFLICT: 'harness' is in your voice fingerprint (4 uses in past samples) but matches strict-tier scrub. Keep or substitute?]
Don't fabricate
The non-negotiable rule (SKILL.md line: "Never introduce facts that weren't in the input") overrides voice-fingerprint matching. If a sample contains specific numbers, do NOT carry those numbers into a different post. Only use numbers the current input already supplies.
| 1 | # Voice Fingerprint — Preserving the user's voice while scrubbing |
| 2 | |
| 3 | The humanizer is destructive by design. Every pass deletes or substitutes tokens. That's fine for AI tells. It's a bug for the user's actual voice. |
| 4 | |
| 5 | This file lists the signals to preserve, even when they overlap with rules in `scrub-rules.md`. |
| 6 | |
| 7 | ## Contents |
| 8 | |
| 9 | Preserve unconditionally (do NOT scrub these) |
| 10 | Preserve when sample voice is provided |
| 11 | Conflict resolution |
| 12 | How to build a voice fingerprint from samples (sketch) |
| 13 | Examples |
| 14 | Don't fabricate |
| 15 | |
| 16 | |
| 17 | |
| 18 | ## Preserve unconditionally (do NOT scrub these) |
| 19 | |
| 20 | These are voice signatures, not AI tells. Leave them alone in every tier including `--mode all`. |
| 21 | |
| 22 | | Pattern | Why it's voice, not AI | |
| 23 | |---|---| |
| 24 | | Lowercase sentence starts (`closed our seed on a tuesday...`) | Users like Serge use this as a deliberate cadence cue. Capitalizing flattens their voice. | |
| 25 | | `..` as a soft pause | This is the humanizer's officially-blessed alternative to em dash. Removing it has nowhere to go. | |
| 26 | | Sentence fragments (`Worth it.`, `Every time.`, `Not even close.`) | Pass 2 ADDS fragments. Don't remove the ones already there. | |
| 27 | | Contractions (`don't`, `it's`, `you're`, `we're`) | Mandatory for natural rhythm. Scrubbing curly apostrophes is fine; expanding contractions is not. | |
| 28 | | First-person sensory detail (`my hands shook`, `the room went quiet`) | Pass 3 demands these. Never strip. | |
| 29 | | Specific numbers (`$47k`, `9:14am`, `47 days`) | Pass 3 demands these. Never strip. | |
| 30 | | Named entities (`HubSpot`, `Tuesday morning`, brand names) | Pass 3 demands these. Capitalize properly per non-negotiable rule. | |
| 31 | | Self-correction within a paragraph (`actually no`, `correction:`) | Burstiness signal. Real humans circle back. | |
| 32 | |
| 33 | |
| 34 | |
| 35 | ## Preserve when sample voice is provided |
| 36 | |
| 37 | If the user passes optional `target_voice_samples` (their last 5-10 LinkedIn posts), extract: |
| 38 | |
| 39 | **Sentence-length distribution.** If they routinely write 4-6 word sentences, don't force 12+ word "minimum lengths" on Pass 2. |
| 40 | **Vocabulary fingerprint.** Words they use 3+ times across samples are part of their voice — even if those words appear on the strict blacklist. Flag for user review rather than auto-substituting. |
| 41 | **Punctuation habits.** Some users use `...` instead of `..`, or unbroken comma chains. Match the dominant pattern. |
| 42 | **Opener patterns.** If they always start with a number (`47 days ago`, `$2M ARR`) or a name (`Jake said`), preserve that template. |
| 43 | **Closer patterns.** If they always close with a single fragment + period (no question), don't force a question CTA. |
| 44 | |
| 45 | |
| 46 | |
| 47 | ## Conflict resolution |
| 48 | |
| 49 | When a scrub rule fires on a token that's also in the user's voice fingerprint: |
| 50 | |
| 51 | | Tier | Behavior | |
| 52 | |---|---| |
| 53 | | Forensic | Always scrub. Forensic rules catch model leakage; if the user's voice fingerprint contains `oaicite` it's because they pasted AI output. | |
| 54 | | Strict | Flag for user review. Don't auto-substitute. The user gets to decide. | |
| 55 | | Aesthetic | Skip the rule entirely. Aesthetic rules already explicitly tolerate human-writer defenses. | |
| 56 | |
| 57 | |
| 58 | |
| 59 | ## How to build a voice fingerprint from samples (sketch) |
| 60 | |
| 61 | |
| 62 | from collections import Counter |
| 63 | import re |
| 64 | |
| 65 | def build_voice_fingerprint(samples: list[str]) -> dict: |
| 66 | text = "\n".join(samples) |
| 67 | sentences = re.split(r'(?<=[.!?])\s+', text) |
| 68 | |
| 69 | return { |
| 70 | "sentence_lengths": [len(s.split()) for s in sentences], |
| 71 | "vocab_freq": Counter(re.findall(r"\b[a-z][a-z']{2,}\b", text.lower())), |
| 72 | "starts_lowercase_pct": sum(1 for s in sentences if s and s[0].islower()) / max(len(sentences), 1), |
| 73 | "uses_double_dot": ".." in text, |
| 74 | "uses_triple_dot": "..." in text, |
| 75 | "punctuation_freq": Counter(c for c in text if c in ".!?,;:"), |
| 76 | "fragment_pct": sum(1 for s in sentences if len(s.split()) <= 4) / max(len(sentences), 1), |
| 77 | } |
| 78 | |
| 79 | |
| 80 | The skill should call this on `target_voice_samples` before running Pass 1. |
| 81 | |
| 82 | |
| 83 | |
| 84 | ## Examples |
| 85 | |
| 86 | ### Example 1 — `..` as soft pause (preserve) |
| 87 | |
| 88 | Input: `closed our seed.. then everything broke` |
| 89 | |
| 90 | Wrong (scrubs the `..`): `closed our seed. then everything broke` |
| 91 | |
| 92 | Right (preserve): `closed our seed.. then everything broke` |
| 93 | |
| 94 | The `..` is on the explicit preserve list. Period substitution is for `--`, not `..`. |
| 95 | |
| 96 | ### Example 2 — lowercase start (preserve) |
| 97 | |
| 98 | Input: `closed our seed on a tuesday morning at 9:14am` |
| 99 | |
| 100 | Wrong (capitalizes): `Closed our seed on a Tuesday morning at 9:14am` |
| 101 | |
| 102 | Right (preserve `closed`, capitalize `Tuesday`): `closed our seed on a Tuesday morning at 9:14am` |
| 103 | |
| 104 | The non-negotiable rule says capitalize NAMES — Tuesday is a proper noun in date context, but the sentence-initial `closed` stays lowercase per voice rule. |
| 105 | |
| 106 | ### Example 3. Voice-fingerprint vocabulary collision (flag, don't substitute) |
| 107 | |
| 108 | User samples contain `harness` 4 times across 6 posts (clearly part of their voice — they work in horse-training tech). |
| 109 | |
| 110 | Strict tier scrub rule says: `harness → use`. |
| 111 | |
| 112 | Right behavior: flag for user review. Output: `[VOICE-CONFLICT: 'harness' is in your voice fingerprint (4 uses in past samples) but matches strict-tier scrub. Keep or substitute?]` |
| 113 | |
| 114 | |
| 115 | |
| 116 | ## Don't fabricate |
| 117 | |
| 118 | The non-negotiable rule (SKILL.md line: "Never introduce facts that weren't in the input") overrides voice-fingerprint matching. If a sample contains specific numbers, do NOT carry those numbers into a different post. Only use numbers the current input already supplies. |
| 119 |
Discussion
Browse more free Claude skills.