25 voice clone podcast global skill

Use when a PERSONAL brand needs AUDIO — voice cloning with ElevenLabs, Murf, or PlayHT, podcast production, audiobooks, and voiceover: short voiceover for TikTok and Reels, a 30 to 60 minute podcast format, and a 1-to-10 repurpose turning one episode into ten clips, in English with US, UK, AU, and SG accents.

by minhnv0807·MIT license·★ 599 Stars on the repo·GitHub ↗

Use now

Files of 25 voice clone podcast global

minhnv0807/master1 file shown
SKILL.md
Show the full text386 lines

Voice Clone & Podcast — Audio AI for Personal Brand (Global)

This skill focuses on audio AI — voice clone, podcast, audiobook, voiceover. Pairs with 24-ai-avatar-production-global (video) — combine both for full content stack coverage.


1. Newbie Guide

What is audio AI and how is it different from video AI?

Audio AI is the tech behind synthetic voices that sound nearly human — from a sample of your voice, AI learns and produces a synthetic clone (voice clone). You write text -> AI reads it back (Text-to-Speech).

Differences vs video AI:

  • Video AI (skill 24): produces video with face + voice -> talking head, social video
  • Audio AI (this skill): produces voice only -> podcast, audiobook, voiceover, narration
When to use audio AI instead of video?
Situation Pick audio AI Pick video AI
Long-form content (>10 min) YES — podcast format NO — too long for video
Don't want to be on camera YES NO
Need volume content fast YES — 1 podcast = 10 shorts YES but more expensive
Audience listens while driving / at gym YES NO
Need visuals to demo NO YES
Personal brand thought leader YES — podcast = authority YES — if face brand exists
Main tools (international)
  • ElevenLabs: Best in class for voice clone — top-tier English voices (US/UK/AU/IN), 30+ languages
  • Murf: 120+ voice library, strong for corporate voiceover, multilingual
  • PlayHT: API-friendly, instant clone, 800+ voices
  • HeyGen Voice: Bundles with HeyGen avatars — seamless voice + video pipeline
  • Descript: AI editing — cut audio by editing text, voice clone (Overdub)
  • Resemble.ai: Custom emotion control, brand-grade APIs
  • Riverside: Studio-quality podcast recording with AI Magic Clips repurpose
Time and cost
Task Time Cost (USD/mo)
Voice clone setup 30-60 min $5-22 (ElevenLabs Starter/Pro)
60s voiceover (TikTok) 5-10 min $5-22
30 min solo podcast 1-2 hrs $22-99 (ElevenLabs + Riverside)
Audiobook chapter (15 min) 30-45 min $22-99
1 podcast -> 10 clips 1-2 hrs $0-30 (Descript/Opus Clip)
5 common mistakes
  1. AI voice sounds robotic: sample too short or monotonic. Fix: re-record 3-5 minutes with varied emotions (happy, serious, sad).
  2. Mispronounced names/jargon: TTS engines mishandle proper nouns. Fix: use phonetic spelling (e.g., "Anthropic" -> "an-THROW-pic") in the script.
  3. Audio clipping: levels too hot. Fix: target -3dB peak, -16 LUFS loudness.
  4. Background noise/echo: untreated room. Fix: small room with curtains and rugs, or apply NVIDIA Broadcast / Krisp / Adobe Enhance Speech.
  5. Boring podcast: no editing, too many "ums". Fix: Descript auto-removes filler words, add light background music (-25dB).

2. Information collection

Ask up to 4 questions before starting:

  1. Main use case? Short voiceover (TikTok/Reels) / Podcast 30-60 min / Audiobook?
  2. Language(s)? English (US/UK/AU/IN) / multilingual / single non-English?
  3. Total length? <60s / 5-30 min / 30-60 min / >60 min (audiobook)?
  4. Budget tier? Free ($0) / Starter ($5-22) / Pro ($22-99) / Business ($99+)?

Based on the answers, pick the appropriate use case + tool stack.


3. Voice clone setup

Sample requirements
Criterion Minimum Optimal
Length 1 min (Free tier) 3-5 min (Pro tier)
Room Quiet, no echo Acoustic treatment, rugs, curtains
Mic Phone + headset mic Condenser mic (AT2020, $80-100)
Distance 20-30 cm 15-20 cm with pop filter
Format MP3 128 kbps WAV 44.1 kHz
Content One pre-written passage Three passages: business / casual / emotional

Full reference: references/voice-clone-prompts-global.md — sample scripts across English variants (US/UK/AU/SG/IN) and 3 topics (business / lifestyle / educational).

Tool comparison (global)
Tool English clone quality Price/mo Setup time Best for
ElevenLabs Pro Excellent (10/10) $22 30 min Multilingual, content creator
HeyGen Voice Good (8/10) Bundled with avatar 15 min Combo with video AI
Murf Excellent (9/10) $29-79 30 min Corporate voiceover, e-learning
PlayHT Excellent (9.5/10) $39-99 30 min API-driven, instant clone
Descript Overdub Good (8/10) $24 (Hobbyist) 30 min Podcast editing
Resemble.ai Excellent (9/10) $30-99 1 hr Brand custom voice, emotion control

Recommendations:

  • English-only creator: ElevenLabs Pro ($22) — best balance of quality and price
  • Multilingual creator: ElevenLabs Pro (30+ languages built in)
  • Combo with video: HeyGen (single platform — voice + avatar)
  • Brand/agency at scale: Resemble.ai or PlayHT (API + custom emotion)
VOICE CLONE LICENSE AGREEMENT

I, [Full name], ID/passport: [number], grant [Brand/Company]:
1. Permission to use samples of my voice to create an AI voice clone.
2. Use of the voice clone in [scope: internal / advertising / podcast / etc.].
3. Term: from [DD/MM/YYYY] to [DD/MM/YYYY].
4. Right of withdrawal: I may request deletion of the voice clone at any time
   in writing; the brand has 7 days to fully remove it.
5. Disclosure: the brand commits to disclose "AI-generated voice" wherever
   required by applicable law (FTC, EU AI Act, etc.).

Signed: ____________   Date: ____________

4. Three use cases

Use case A: Short voiceover for TikTok/Reels (Energetic)

Spec:

  • Length: 15-60s
  • Pace: fast (180-220 wpm) — younger English-speaking audience
  • Tone: energetic, slightly higher pitch, exciting
  • Audio levels: -14 LUFS (TikTok), peak -1 dB
  • CTA: clear in the last 5 seconds

Script template (30s):

[HOOK 0-3s] "Did you know [shocking stat]?"
[PROBLEM 3-10s] "Most people are still stuck in [wrong loop]"
[SOLUTION 10-22s] "I tried [method], and here are 3 things..."
[PAYOFF 22-27s] "Result: [specific number]"
[CTA 27-30s] "Comment 'YES' to get the full breakdown"

Voice settings (ElevenLabs):

  • Stability: 35-45 (low — allows variation)
  • Similarity: 75-85
  • Style: 50-65 (boost expressiveness)
  • Speaker Boost: ON
Use case B: Podcast 30-60 min (Conversational)

Structure:

  • Intro (1-2 min): hook + introduce topic + welcome listeners
  • Body (25-50 min): 3-5 main segments, each 5-10 min
  • Ad slot (optional): 3-5 min after intro, or mid-body
  • Outro (1-2 min): recap + CTA + thanks

Pacing:

  • Conversational pace: 140-160 wpm
  • 1-2s pause after important sentences
  • Segment transitions: 2-3s pause + audio sting

Sound design:

  • Background music: -25 to -30 dB (very subtle)
  • Stings/transitions: -15 dB, 1-2s
  • Voice levels: -16 LUFS (podcast standard), peak -1 dB

Voice settings (ElevenLabs):

  • Stability: 60-75 (high — consistent across 30+ minutes)
  • Similarity: 85-95
  • Style: 30-40 (natural, not over-expressive)
  • Speaker Boost: ON
Use case C: Audiobook (Mid-tempo)

Structure:

  • Chapter intro: "Chapter [X]: [Title]" — 2s pause
  • Chapter body: 10-20 min/chapter, 1s pause between paragraphs
  • Chapter end: 3s pause before next chapter

Pacing:

  • Mid-tempo: 150-170 wpm
  • Natural breath every 2-3 sentences
  • Dialogue: subtle voice shifts per character (fiction)

Consistency check (most important):

  • Render Chapter 1 and Chapter 5 -> compare voice -> must match 95%+
  • If voices drift: re-clone with a longer sample (5+ min)
  • Pronunciation guide: build a database of proper nouns + custom phonetics

Voice settings (ElevenLabs):

  • Stability: 70-85 (very high — consistent for hours)
  • Similarity: 90-95
  • Style: 20-30 (calm, even)
  • Speaker Boost: ON

5. Tool comparison (global)

Tool Price/mo English quality Multilingual Setup Pros Cons Best for
ElevenLabs $5-99 10/10 30+ langs 30 min Best clone, multilingual Pricier high tiers Multilingual creator
HeyGen Voice Bundle w/ avatar 8/10 40+ langs 15 min Combo with avatar Voice clone less expressive Combo with video
Descript $24-30 9/10 EN focus 30 min Audio editing first Multilingual weaker Podcast editing
Riverside $19-29 n/a (recording) n/a 5 min Studio recording Not TTS Live podcast
Murf $29-79 9/10 20+ langs 30 min 120+ voice library Voice clone limited tier Corporate voiceover
PlayHT $39-99 9.5/10 100+ langs 30 min Strong API, instant clone UI dense Developer/API
Resemble.ai $30-99 9/10 60+ langs 1 hr Custom emotion control Steep learning curve Brand custom voice

Recommended combos 2025-2026:

  • English solo creator: ElevenLabs Pro ($22) + Riverside Free + Descript Hobbyist ($24)
  • Multilingual creator: ElevenLabs Pro ($22) + Riverside Standard ($19) + Descript Pro ($30)
  • Brand/agency: ElevenLabs Creator ($99) + Resemble.ai + Riverside Pro ($29)

6. 1-on-1 podcast with an AI co-host

Use case: solo podcaster who wants conversational format but can't find a co-host. AI co-host = a second AI voice that asks questions while you answer.

Setup — prompt-engineering the AI personality

Step 1: Define the AI co-host's personality

Name: [AI co-host name]
Personality: curious, asks deep follow-ups, occasionally light humor
Role: asks the host questions, doesn't talk too much
Speaking style: casual, natural, addresses the host by first name
Knowledge level: average — asks questions like a listener would
Catchphrases: "Wow, that's wild." / "What does that mean exactly?" / "Can you go deeper?"

Step 2: Create a separate voice clone for the AI co-host

  • Use a different voice than the host (e.g., woman vs man, or different accent)
  • Clone from a consenting friend, or use one of ElevenLabs' built-in voices

Step 3: Tool stack

  • ElevenLabs: generate the AI co-host voice
  • Riverside: record the host live
  • Descript: edit + splice in the AI co-host (text-to-audio)
Q&A script template
[INTRO]
Host: Hey everyone, today [AI co-host] and I are diving into...
AI co-host: Hi all, I'm [name]. Today I want to dig into [topic] from [host]'s
            point of view. Let's go!

[BODY — 5-7 Q&A pairs]
AI co-host: [Broad opening question]
Host: [Answers 2-3 minutes]
AI co-host: [Deeper follow-up]
Host: [Answers with a concrete example]
... repeat 5-7 times ...

[OUTRO]
AI co-host: Thanks [host] for sharing. The biggest thing I learned was...
Host: Thanks [AI co-host]. If you have questions, drop them in the comments...

Tip: pre-write 7-10 AI co-host questions in a doc, record host responses in one go. Then generate AI co-host audio in ElevenLabs and splice in via Descript.


7. Repurpose pipeline 1:10 (1 podcast -> 10 short clips)

Workflow overview
[1] Record 60-min podcast (Riverside)
        v
[2] Auto-transcript (Descript / Riverside)
        v
[3] Identify hooks (10-15 quotable lines)
        v
[4] Cut 30-60s clips per quote (Opus Clip / Descript)
        v
[5] Add captions (auto-caption)
        v
[6] Distribute across 4 platforms
How to identify hooks

Find moments in the transcript with these traits:

  • Bold statement: "I think 90% of founders are doing this wrong"
  • Counter-intuitive: "Raising prices actually grew revenue"
  • Specific number: "Went from $0 to $1M in 6 months"
  • Personal story: "On my first startup I lost $200K"
  • Actionable tip: "3 specific steps you can take today"

Target: 10-15 hooks per 60-min podcast. Pick the 10 best.

Tool stack
  • Descript: auto-clip — select sentences, export short clips (free tier 1 hr/month)
  • Opus Clip: AI auto-finds viral moments + auto-format vertical/horizontal ($19-99)
  • Riverside Magic Clips: built-in to Riverside Pro ($29)
  • CapCut + ChatGPT: manual but free — paste transcript, ChatGPT extracts hook candidates
Distribution across 4 platforms
Platform Format Length Caption Bonus
TikTok 9:16 (1080×1920) 30-60s Bold caption on top Trend audio overlay (low volume)
Instagram Reels 9:16 15-90s Clean subtitle, sans-serif font Strong cover image
YouTube Shorts 9:16 <60s Auto-caption Title with target keyword
LinkedIn audio 1:1 (square video w/ audio) 60-120s Subtitle below Long-form thread (carousel)

Pro tip: each clip should target one platform with platform-specific captions and cover image. Maximizes reach.


8. Audio QA + Disclosure

5 QA criteria
  1. Clarity (10 pts): voice clear, no rasp, no stuttering. Test: play on phone speaker, still intelligible.
  2. No clipping (10 pts): peak below -1 dB. Tools: Audacity, Adobe Audition, Reaper.
  3. No background noise (10 pts): no fans, traffic, neighbors. Tools: Krisp, NVIDIA Broadcast, Adobe Enhance Speech.
  4. Consistent loudness (10 pts): stable -16 LUFS (podcast) or -14 LUFS (TikTok). Tool: loudness meter in DAW.
  5. Natural pauses (10 pts): human-like pacing. Manual review: listen 3 times.

Pass: 40+/50. Below 40 = re-render or re-record.

Global disclosure — when required
Situation Disclosure Placement
Commercial advertising REQUIRED Caption + end of audio ("This audio uses an AI voice clone")
Personal brand podcast RECOMMENDED — transparency Episode description
Fiction audiobook OPTIONAL Optional — credits at end
News/educational REQUIRED Beginning of audio + caption
Internal corporate content NOT REQUIRED n/a

Disclosure caption template:

This audio uses AI voice cloning technology
(ElevenLabs / Murf / [tool name]). Content was written and reviewed by [Name].

Full reference: references/ai-video-disclosure-global.md — FTC, EU AI Act, FCC, and OFCOM requirements; 3-tier disclosure framework, situational templates (also applies to audio).


9. Quality checklist

Before publishing audio:

  • Voice-clone sample 3-5 min, quiet room
  • Signed consent form (if cloning someone else's voice)
  • Use case matches: voiceover (energetic) / podcast (conversational) / audiobook (mid-tempo)
  • Voice settings appropriate for use case (Stability/Similarity/Style)
  • Loudness correct: -14 LUFS (TikTok) / -16 LUFS (podcast/audiobook)
  • Peak below -1 dB (no clipping)
  • No background noise (Krisp/NVIDIA Broadcast pass)
  • Pacing correct: 180-220 wpm (TikTok) / 140-160 wpm (podcast) / 150-170 wpm (audiobook)
  • QA Score 40+/50
  • Disclosure caption (if commercial use)
  • Repurpose plan: 1 podcast -> 10 clips across 4 platforms

Skill 25 (Global) | v1.0.0

1---
2name: 25-voice-clone-podcast-global
3description: "Use when a PERSONAL brand needs AUDIO — voice cloning with ElevenLabs, Murf, or PlayHT, podcast production, audiobooks, and voiceover: short voiceover for TikTok and Reels, a 30 to 60 minute podcast format, and a 1-to-10 repurpose turning one episode into ten clips, in English with US, UK, AU, and SG accents. Trigger on 'voice clone', 'ElevenLabs', 'start a podcast', 'audiobook narration', 'AI voiceover for my videos', 'I hate re-recording the same intro'. Not for — video with a talking-head avatar, see `24-ai-avatar-production-global`; the script being read aloud, see `04-script-video-global`; written long-form posts, see `26-thought-leadership-content-global`."
4metadata:
5 version: 1.0.1
6 category: content
7license: MIT
8triggers:
9 - "voice clone global"
10 - "ElevenLabs"
11 - "podcast AI"
12 - "audiobook AI"
13 - "voiceover AI"
14related:
15 - 24-ai-avatar-production-global
16 - 26-thought-leadership-content-global
17 - references/voice-clone-prompts-global
18 - references/ai-video-disclosure-global
19---
20 
21# Voice Clone & Podcast — Audio AI for Personal Brand (Global)
22 
23> **This skill focuses on audio AI** — voice clone, podcast, audiobook, voiceover.
24> Pairs with `24-ai-avatar-production-global` (video) — combine both for full content stack coverage.
25 
26---
27 
28## 1. Newbie Guide
29 
30### What is audio AI and how is it different from video AI?
31 
32Audio AI is the tech behind synthetic voices that sound nearly human — from a sample of your voice, AI learns and produces a synthetic clone (voice clone). You write text -> AI reads it back (Text-to-Speech).
33 
34**Differences vs video AI:**
35- Video AI (skill 24): produces video with face + voice -> talking head, social video
36- Audio AI (this skill): produces voice only -> podcast, audiobook, voiceover, narration
37 
38### When to use audio AI instead of video?
39 
40| Situation | Pick audio AI | Pick video AI |
41|-----------|---------------|---------------|
42| Long-form content (>10 min) | YES — podcast format | NO — too long for video |
43| Don't want to be on camera | YES | NO |
44| Need volume content fast | YES — 1 podcast = 10 shorts | YES but more expensive |
45| Audience listens while driving / at gym | YES | NO |
46| Need visuals to demo | NO | YES |
47| Personal brand thought leader | YES — podcast = authority | YES — if face brand exists |
48 
49### Main tools (international)
50 
51- **ElevenLabs:** Best in class for voice clone — top-tier English voices (US/UK/AU/IN), 30+ languages
52- **Murf:** 120+ voice library, strong for corporate voiceover, multilingual
53- **PlayHT:** API-friendly, instant clone, 800+ voices
54- **HeyGen Voice:** Bundles with HeyGen avatars — seamless voice + video pipeline
55- **Descript:** AI editing — cut audio by editing text, voice clone (Overdub)
56- **Resemble.ai:** Custom emotion control, brand-grade APIs
57- **Riverside:** Studio-quality podcast recording with AI Magic Clips repurpose
58 
59### Time and cost
60 
61| Task | Time | Cost (USD/mo) |
62|------|------|---------------|
63| Voice clone setup | 30-60 min | $5-22 (ElevenLabs Starter/Pro) |
64| 60s voiceover (TikTok) | 5-10 min | $5-22 |
65| 30 min solo podcast | 1-2 hrs | $22-99 (ElevenLabs + Riverside) |
66| Audiobook chapter (15 min) | 30-45 min | $22-99 |
67| 1 podcast -> 10 clips | 1-2 hrs | $0-30 (Descript/Opus Clip) |
68 
69### 5 common mistakes
70 
711. **AI voice sounds robotic:** sample too short or monotonic. Fix: re-record 3-5 minutes with varied emotions (happy, serious, sad).
722. **Mispronounced names/jargon:** TTS engines mishandle proper nouns. Fix: use phonetic spelling (e.g., "Anthropic" -> "an-THROW-pic") in the script.
733. **Audio clipping:** levels too hot. Fix: target -3dB peak, -16 LUFS loudness.
744. **Background noise/echo:** untreated room. Fix: small room with curtains and rugs, or apply NVIDIA Broadcast / Krisp / Adobe Enhance Speech.
755. **Boring podcast:** no editing, too many "ums". Fix: Descript auto-removes filler words, add light background music (-25dB).
76 
77---
78 
79## 2. Information collection
80 
81Ask up to 4 questions before starting:
82 
831. **Main use case?** Short voiceover (TikTok/Reels) / Podcast 30-60 min / Audiobook?
842. **Language(s)?** English (US/UK/AU/IN) / multilingual / single non-English?
853. **Total length?** <60s / 5-30 min / 30-60 min / >60 min (audiobook)?
864. **Budget tier?** Free ($0) / Starter ($5-22) / Pro ($22-99) / Business ($99+)?
87 
88> Based on the answers, pick the appropriate use case + tool stack.
89 
90---
91 
92## 3. Voice clone setup
93 
94### Sample requirements
95 
96| Criterion | Minimum | Optimal |
97|-----------|---------|---------|
98| Length | 1 min (Free tier) | 3-5 min (Pro tier) |
99| Room | Quiet, no echo | Acoustic treatment, rugs, curtains |
100| Mic | Phone + headset mic | Condenser mic (AT2020, $80-100) |
101| Distance | 20-30 cm | 15-20 cm with pop filter |
102| Format | MP3 128 kbps | WAV 44.1 kHz |
103| Content | One pre-written passage | Three passages: business / casual / emotional |
104 
105> **Full reference:** `references/voice-clone-prompts-global.md` — sample scripts across English variants (US/UK/AU/SG/IN) and 3 topics (business / lifestyle / educational).
106 
107### Tool comparison (global)
108 
109| Tool | English clone quality | Price/mo | Setup time | Best for |
110|------|----------------------|----------|------------|----------|
111| **ElevenLabs Pro** | Excellent (10/10) | $22 | 30 min | Multilingual, content creator |
112| **HeyGen Voice** | Good (8/10) | Bundled with avatar | 15 min | Combo with video AI |
113| **Murf** | Excellent (9/10) | $29-79 | 30 min | Corporate voiceover, e-learning |
114| **PlayHT** | Excellent (9.5/10) | $39-99 | 30 min | API-driven, instant clone |
115| **Descript Overdub** | Good (8/10) | $24 (Hobbyist) | 30 min | Podcast editing |
116| **Resemble.ai** | Excellent (9/10) | $30-99 | 1 hr | Brand custom voice, emotion control |
117 
118**Recommendations:**
119- **English-only creator:** ElevenLabs Pro ($22) — best balance of quality and price
120- **Multilingual creator:** ElevenLabs Pro (30+ languages built in)
121- **Combo with video:** HeyGen (single platform — voice + avatar)
122- **Brand/agency at scale:** Resemble.ai or PlayHT (API + custom emotion)
123 
124### Consent form template
125 
126```
127VOICE CLONE LICENSE AGREEMENT
128 
129I, [Full name], ID/passport: [number], grant [Brand/Company]:
1301. Permission to use samples of my voice to create an AI voice clone.
1312. Use of the voice clone in [scope: internal / advertising / podcast / etc.].
1323. Term: from [DD/MM/YYYY] to [DD/MM/YYYY].
1334. Right of withdrawal: I may request deletion of the voice clone at any time
134 in writing; the brand has 7 days to fully remove it.
1355. Disclosure: the brand commits to disclose "AI-generated voice" wherever
136 required by applicable law (FTC, EU AI Act, etc.).
137 
138Signed: ____________ Date: ____________
139```
140 
141---
142 
143## 4. Three use cases
144 
145### Use case A: Short voiceover for TikTok/Reels (Energetic)
146 
147**Spec:**
148- Length: 15-60s
149- Pace: fast (180-220 wpm) — younger English-speaking audience
150- Tone: energetic, slightly higher pitch, exciting
151- Audio levels: -14 LUFS (TikTok), peak -1 dB
152- CTA: clear in the last 5 seconds
153 
154**Script template (30s):**
155```
156[HOOK 0-3s] "Did you know [shocking stat]?"
157[PROBLEM 3-10s] "Most people are still stuck in [wrong loop]"
158[SOLUTION 10-22s] "I tried [method], and here are 3 things..."
159[PAYOFF 22-27s] "Result: [specific number]"
160[CTA 27-30s] "Comment 'YES' to get the full breakdown"
161```
162 
163**Voice settings (ElevenLabs):**
164- Stability: 35-45 (low — allows variation)
165- Similarity: 75-85
166- Style: 50-65 (boost expressiveness)
167- Speaker Boost: ON
168 
169### Use case B: Podcast 30-60 min (Conversational)
170 
171**Structure:**
172- **Intro (1-2 min):** hook + introduce topic + welcome listeners
173- **Body (25-50 min):** 3-5 main segments, each 5-10 min
174- **Ad slot (optional):** 3-5 min after intro, or mid-body
175- **Outro (1-2 min):** recap + CTA + thanks
176 
177**Pacing:**
178- Conversational pace: 140-160 wpm
179- 1-2s pause after important sentences
180- Segment transitions: 2-3s pause + audio sting
181 
182**Sound design:**
183- Background music: -25 to -30 dB (very subtle)
184- Stings/transitions: -15 dB, 1-2s
185- Voice levels: -16 LUFS (podcast standard), peak -1 dB
186 
187**Voice settings (ElevenLabs):**
188- Stability: 60-75 (high — consistent across 30+ minutes)
189- Similarity: 85-95
190- Style: 30-40 (natural, not over-expressive)
191- Speaker Boost: ON
192 
193### Use case C: Audiobook (Mid-tempo)
194 
195**Structure:**
196- **Chapter intro:** "Chapter [X]: [Title]" — 2s pause
197- **Chapter body:** 10-20 min/chapter, 1s pause between paragraphs
198- **Chapter end:** 3s pause before next chapter
199 
200**Pacing:**
201- Mid-tempo: 150-170 wpm
202- Natural breath every 2-3 sentences
203- Dialogue: subtle voice shifts per character (fiction)
204 
205**Consistency check (most important):**
206- Render Chapter 1 and Chapter 5 -> compare voice -> must match 95%+
207- If voices drift: re-clone with a longer sample (5+ min)
208- Pronunciation guide: build a database of proper nouns + custom phonetics
209 
210**Voice settings (ElevenLabs):**
211- Stability: 70-85 (very high — consistent for hours)
212- Similarity: 90-95
213- Style: 20-30 (calm, even)
214- Speaker Boost: ON
215 
216---
217 
218## 5. Tool comparison (global)
219 
220| Tool | Price/mo | English quality | Multilingual | Setup | Pros | Cons | Best for |
221|------|----------|-----------------|-------------|-------|------|------|----------|
222| **ElevenLabs** | $5-99 | 10/10 | 30+ langs | 30 min | Best clone, multilingual | Pricier high tiers | Multilingual creator |
223| **HeyGen Voice** | Bundle w/ avatar | 8/10 | 40+ langs | 15 min | Combo with avatar | Voice clone less expressive | Combo with video |
224| **Descript** | $24-30 | 9/10 | EN focus | 30 min | Audio editing first | Multilingual weaker | Podcast editing |
225| **Riverside** | $19-29 | n/a (recording) | n/a | 5 min | Studio recording | Not TTS | Live podcast |
226| **Murf** | $29-79 | 9/10 | 20+ langs | 30 min | 120+ voice library | Voice clone limited tier | Corporate voiceover |
227| **PlayHT** | $39-99 | 9.5/10 | 100+ langs | 30 min | Strong API, instant clone | UI dense | Developer/API |
228| **Resemble.ai** | $30-99 | 9/10 | 60+ langs | 1 hr | Custom emotion control | Steep learning curve | Brand custom voice |
229 
230**Recommended combos 2025-2026:**
231- **English solo creator:** ElevenLabs Pro ($22) + Riverside Free + Descript Hobbyist ($24)
232- **Multilingual creator:** ElevenLabs Pro ($22) + Riverside Standard ($19) + Descript Pro ($30)
233- **Brand/agency:** ElevenLabs Creator ($99) + Resemble.ai + Riverside Pro ($29)
234 
235---
236 
237## 6. 1-on-1 podcast with an AI co-host
238 
239> **Use case:** solo podcaster who wants conversational format but can't find a co-host. AI co-host = a second AI voice that asks questions while you answer.
240 
241### Setup — prompt-engineering the AI personality
242 
243**Step 1: Define the AI co-host's personality**
244```
245Name: [AI co-host name]
246Personality: curious, asks deep follow-ups, occasionally light humor
247Role: asks the host questions, doesn't talk too much
248Speaking style: casual, natural, addresses the host by first name
249Knowledge level: average — asks questions like a listener would
250Catchphrases: "Wow, that's wild." / "What does that mean exactly?" / "Can you go deeper?"
251```
252 
253**Step 2: Create a separate voice clone for the AI co-host**
254- Use a different voice than the host (e.g., woman vs man, or different accent)
255- Clone from a consenting friend, or use one of ElevenLabs' built-in voices
256 
257**Step 3: Tool stack**
258- **ElevenLabs:** generate the AI co-host voice
259- **Riverside:** record the host live
260- **Descript:** edit + splice in the AI co-host (text-to-audio)
261 
262### Q&A script template
263 
264```
265[INTRO]
266Host: Hey everyone, today [AI co-host] and I are diving into...
267AI co-host: Hi all, I'm [name]. Today I want to dig into [topic] from [host]'s
268 point of view. Let's go!
269 
270[BODY — 5-7 Q&A pairs]
271AI co-host: [Broad opening question]
272Host: [Answers 2-3 minutes]
273AI co-host: [Deeper follow-up]
274Host: [Answers with a concrete example]
275... repeat 5-7 times ...
276 
277[OUTRO]
278AI co-host: Thanks [host] for sharing. The biggest thing I learned was...
279Host: Thanks [AI co-host]. If you have questions, drop them in the comments...
280```
281 
282**Tip:** pre-write 7-10 AI co-host questions in a doc, record host responses in one go. Then generate AI co-host audio in ElevenLabs and splice in via Descript.
283 
284---
285 
286## 7. Repurpose pipeline 1:10 (1 podcast -> 10 short clips)
287 
288### Workflow overview
289 
290```
291[1] Record 60-min podcast (Riverside)
292 v
293[2] Auto-transcript (Descript / Riverside)
294 v
295[3] Identify hooks (10-15 quotable lines)
296 v
297[4] Cut 30-60s clips per quote (Opus Clip / Descript)
298 v
299[5] Add captions (auto-caption)
300 v
301[6] Distribute across 4 platforms
302```
303 
304### How to identify hooks
305 
306Find moments in the transcript with these traits:
307- **Bold statement:** "I think 90% of founders are doing this wrong"
308- **Counter-intuitive:** "Raising prices actually grew revenue"
309- **Specific number:** "Went from $0 to $1M in 6 months"
310- **Personal story:** "On my first startup I lost $200K"
311- **Actionable tip:** "3 specific steps you can take today"
312 
313> **Target:** 10-15 hooks per 60-min podcast. Pick the 10 best.
314 
315### Tool stack
316 
317- **Descript:** auto-clip — select sentences, export short clips (free tier 1 hr/month)
318- **Opus Clip:** AI auto-finds viral moments + auto-format vertical/horizontal ($19-99)
319- **Riverside Magic Clips:** built-in to Riverside Pro ($29)
320- **CapCut + ChatGPT:** manual but free — paste transcript, ChatGPT extracts hook candidates
321 
322### Distribution across 4 platforms
323 
324| Platform | Format | Length | Caption | Bonus |
325|----------|--------|--------|---------|-------|
326| **TikTok** | 9:16 (1080×1920) | 30-60s | Bold caption on top | Trend audio overlay (low volume) |
327| **Instagram Reels** | 9:16 | 15-90s | Clean subtitle, sans-serif font | Strong cover image |
328| **YouTube Shorts** | 9:16 | <60s | Auto-caption | Title with target keyword |
329| **LinkedIn audio** | 1:1 (square video w/ audio) | 60-120s | Subtitle below | Long-form thread (carousel) |
330 
331**Pro tip:** each clip should target one platform with platform-specific captions and cover image. Maximizes reach.
332 
333---
334 
335## 8. Audio QA + Disclosure
336 
337### 5 QA criteria
338 
3391. **Clarity (10 pts):** voice clear, no rasp, no stuttering. Test: play on phone speaker, still intelligible.
3402. **No clipping (10 pts):** peak below -1 dB. Tools: Audacity, Adobe Audition, Reaper.
3413. **No background noise (10 pts):** no fans, traffic, neighbors. Tools: Krisp, NVIDIA Broadcast, Adobe Enhance Speech.
3424. **Consistent loudness (10 pts):** stable -16 LUFS (podcast) or -14 LUFS (TikTok). Tool: loudness meter in DAW.
3435. **Natural pauses (10 pts):** human-like pacing. Manual review: listen 3 times.
344 
345> **Pass:** 40+/50. Below 40 = re-render or re-record.
346 
347### Global disclosure — when required
348 
349| Situation | Disclosure | Placement |
350|-----------|-----------|-----------|
351| Commercial advertising | REQUIRED | Caption + end of audio ("This audio uses an AI voice clone") |
352| Personal brand podcast | RECOMMENDED — transparency | Episode description |
353| Fiction audiobook | OPTIONAL | Optional — credits at end |
354| News/educational | REQUIRED | Beginning of audio + caption |
355| Internal corporate content | NOT REQUIRED | n/a |
356 
357**Disclosure caption template:**
358```
359This audio uses AI voice cloning technology
360(ElevenLabs / Murf / [tool name]). Content was written and reviewed by [Name].
361```
362 
363> **Full reference:** `references/ai-video-disclosure-global.md` — FTC, EU AI Act, FCC, and OFCOM requirements; 3-tier disclosure framework, situational templates (also applies to audio).
364 
365---
366 
367## 9. Quality checklist
368 
369Before publishing audio:
370 
371- [ ] Voice-clone sample 3-5 min, quiet room
372- [ ] Signed consent form (if cloning someone else's voice)
373- [ ] Use case matches: voiceover (energetic) / podcast (conversational) / audiobook (mid-tempo)
374- [ ] Voice settings appropriate for use case (Stability/Similarity/Style)
375- [ ] Loudness correct: -14 LUFS (TikTok) / -16 LUFS (podcast/audiobook)
376- [ ] Peak below -1 dB (no clipping)
377- [ ] No background noise (Krisp/NVIDIA Broadcast pass)
378- [ ] Pacing correct: 180-220 wpm (TikTok) / 140-160 wpm (podcast) / 150-170 wpm (audiobook)
379- [ ] QA Score 40+/50
380- [ ] Disclosure caption (if commercial use)
381- [ ] Repurpose plan: 1 podcast -> 10 clips across 4 platforms
382 
383---
384 
385*Skill 25 (Global) | v1.0.0*
386 

Discussion

Alternatives