Podcast-to-Everything Pipeline skill
Podcast-to-Everything content pipeline.
by ericosiu·MIT license·★ 3,615 Stars on the repo·GitHub ↗
npx degit ericosiu/ai-marketing-skills/podcast-ops#main ~/.claude/skills/podcast-opsChecked ·commit main
Files of Podcast-to-Everything Pipeline
Show the full text318 lines
Preamble (runs on skill start)
# Version check (silent if up to date)
python3 telemetry/version_check.py 2>/dev/null || true
# Telemetry opt-in (first run only, then remembers your choice)
python3 telemetry/telemetry_init.py 2>/dev/null || true
Privacy: This skill logs usage locally to
~/.ai-marketing-skills/analytics/. Remote telemetry is opt-in only. No code, file paths, or repo content is ever collected. Seetelemetry/README.md.
Podcast-to-Everything Pipeline
Turns podcast episodes into a full content calendar across every platform. One episode in, 15-20 content pieces out — scored, deduplicated, and scheduled.
Step 1: Ingest — Get the Transcript
Determine the input source and obtain a clean transcript.
Option A: RSS Feed (--rss <url>)
- Fetch the RSS feed XML
- Extract the latest episode's audio URL (or use
--episodes Nfor batch) - Download the audio file
- Transcribe via OpenAI Whisper API (with timestamps)
- Store transcript with episode metadata (title, date, description, duration)
Option B: Raw Transcript (--transcript <file>)
- Read the transcript file (plain text, SRT, or VTT)
- Parse timestamps if present
- Extract episode metadata from filename or prompt user
Option C: Batch Mode (--batch <rss_url> --episodes N)
- Fetch RSS feed
- Extract the last N episodes
- Process each through the full pipeline
- Deduplicate across all episodes in the batch
Transcript cleanup
- Remove filler words (um, uh, like, you know) for written content
- Preserve original with timestamps for video clip suggestions
- Split into logical segments by topic shift
Step 2: Editorial Brain — Deep Analysis
Feed the full transcript to the LLM with this extraction framework:
Extract these content atoms:
Narrative Arcs — Complete story segments with setup → tension → resolution. Tag with start/end timestamps.
Quotable Moments — Punchy, shareable statements. One-liners that stand alone. Must pass the "would someone screenshot this?" test.
Controversial Takes — Opinions that go against conventional wisdom. The stuff that makes people reply "hard disagree" or "finally someone said it."
Data Points — Specific numbers, percentages, dollar amounts, timeframes. Concrete proof points that add credibility.
Stories — Personal anecdotes, case studies, client examples. Must have a character, a problem, and an outcome.
Frameworks — Step-by-step processes, mental models, decision matrices. Anything structured that people would save or bookmark.
Predictions — Forward-looking claims about trends, markets, technology. Hot takes about where things are going.
Output format per atom:
- Type: [narrative_arc | quote | controversial_take | data_point | story | framework | prediction]
- Content: [extracted text]
- Timestamp: [start - end, if available]
- Context: [what was being discussed]
- Viral Score: [0-100, see Step 4]
- Suggested platforms: [where this atom works best]
Step 3: Content Generation — One Episode, Many Pieces
For each episode, generate ALL of these from the extracted atoms:
3a. Short-Form Video Clips (3-5 per episode)
- Hook: [First 3 seconds — pattern interrupt or bold claim]
- Clip segment: [Timestamp range from transcript]
- Caption overlay: [Text for the screen]
- Platform: [YouTube Shorts / TikTok / Instagram Reels]
- Why it works: [What makes this clippable]
Prioritize: controversial takes > stories with payoffs > surprising data points
3b. Twitter/X Threads (2-3 per episode)
- Thread hook (tweet 1): [Curiosity gap or bold opener]
- Thread body (5-10 tweets): [Each tweet is one complete thought]
- Thread closer: [CTA — follow, reply, retweet trigger]
- Source atoms: [Which content atoms feed this thread]
Rules: No tweet over 280 chars. Each tweet must stand alone. Use data points as proof.
3c. LinkedIn Article Draft (1 per episode)
- Headline: [Specific, benefit-driven]
- Hook paragraph: [Before the "see more" fold — must earn the click]
- Body: [3-5 sections with headers, 800-1200 words]
- CTA: [Engagement driver — question, not link]
- Hashtags: [3-5 relevant, not spammy]
Voice: Professional but not corporate. First-person. Story-driven.
3d. Newsletter Section (1 per episode)
- Section headline: [Scannable, specific]
- TL;DR: [One sentence, the core insight]
- Body: [3-5 bullet points, each with a takeaway]
- Pull quote: [The most shareable line from the episode]
- Link: [Back to full episode]
3e. Quote Cards (3-5 per episode)
- Quote text: [Max 20 words — must work as text overlay]
- Attribution: [Speaker name]
- Background suggestion: [Color/mood that matches the tone]
- Platform sizing: [1080x1080 for IG, 1200x675 for Twitter, 1080x1920 for Stories]
3f. Blog Post Outline (1 per episode)
- Title: [SEO-optimized, includes primary keyword]
- Primary keyword: [Search volume + difficulty estimate]
- Secondary keywords: [3-5 related terms]
- Meta description: [155 chars max]
- H2 sections: [5-7, each maps to a content atom]
- Internal linking opportunities: [Topics that connect to existing content]
- Estimated word count: [1500-2500]
3g. YouTube Shorts / TikTok Script (1 per episode)
- HOOK (0-3s): [Pattern interrupt — question, bold claim, or visual]
- SETUP (3-15s): [Context — why should they care]
- PAYOFF (15-45s): [The insight, data, or story resolution]
- CTA (45-60s): [Follow, comment prompt, or part 2 tease]
- On-screen text: [Key phrases to overlay]
- B-roll suggestions: [Visual ideas if not talking-head]
Step 4: Content Scoring — Viral Potential
Score every generated piece on three dimensions (each 0-100):
| Dimension | What It Measures | Signals |
|---|---|---|
| Novelty | Is this new or surprising? | Contrarian takes, unexpected data, first-to-say |
| Controversy | Will people argue about this? | Strong opinions, challenges norms, picks a side |
| Utility | Can someone use this immediately? | Frameworks, how-tos, templates, specific numbers |
Viral Score = (Novelty × 0.4) + (Controversy × 0.3) + (Utility × 0.3)
Score thresholds:
- 80+ → Priority publish. Schedule for peak engagement windows.
- 60-79 → Solid content. Fill the calendar.
- 40-59 → Filler. Use only if calendar has gaps.
- Below 40 → Cut it. Not worth the publish slot.
Step 5: Dedup Engine
Before finalizing, check all generated content against:
- This batch — No two pieces should cover the same angle
- Recent history — Compare against last N days of output (default: 30)
- Similarity threshold — Flag any pair with >70% semantic overlap
Dedup rules:
- If two pieces overlap >70%: keep the higher-scored one, cut the other
- If a piece overlaps with recently published content: flag with ⚠️ and suggest a differentiation angle
- Track all published content hashes in
output/content_history.json
Step 6: Calendar Generation (--calendar)
Assemble scored, deduplicated content into a weekly publish calendar.
Scheduling rules:
- Twitter/X: 1-2 per day, peak hours (8-10am, 12-1pm, 5-7pm ET)
- LinkedIn: 1 per day max, Tuesday-Thursday mornings
- YouTube Shorts/TikTok: 1 per day, evenings
- Newsletter: Weekly, same day each week
- Blog: 1-2 per week
- Quote cards: Intersperse on low-content days
Calendar output format:
{
"week_of": "2024-01-15",
"episode_source": "Episode Title - Guest Name",
"content_pieces": [
{
"date": "2024-01-15",
"time": "09:00 ET",
"platform": "twitter",
"type": "thread",
"content": "...",
"viral_score": 85,
"status": "draft"
}
],
"total_pieces": 18,
"avg_viral_score": 72,
"coverage": {
"twitter": 6,
"linkedin": 3,
"youtube_shorts": 3,
"newsletter": 1,
"blog": 1,
"quote_cards": 4
}
}
Step 7: Output
All output goes to output/ directory:
output/
├── episodes/
│ ├── YYYY-MM-DD-episode-slug/
│ │ ├── transcript.txt
│ │ ├── atoms.json # Extracted content atoms
│ │ ├── content_pieces.json # All generated content
│ │ └── calendar.json # Scheduled calendar
│ └── ...
├── calendar/
│ └── week-YYYY-WNN.json # Aggregated weekly calendar
├── content_history.json # Dedup tracking
└── pipeline_log.json # Run history and stats
CLI Reference
# Process latest episode from RSS feed
python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml"
# Process a local transcript
python podcast_pipeline.py --transcript episode-42.txt
# Batch process last 5 episodes
python podcast_pipeline.py --batch "https://feeds.example.com/podcast.xml" --episodes 5
# Generate weekly calendar from existing outputs
python podcast_pipeline.py --calendar
# Process with custom dedup window
python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml" --dedup-days 60
# Process and only keep 80+ viral score content
python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml" --min-score 80
Environment Variables
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY |
Yes (for Whisper) | OpenAI API key for audio transcription |
ANTHROPIC_API_KEY |
Yes (for generation) | Anthropic API key for content generation |
OPENAI_LLM_KEY |
Optional | Separate OpenAI key if using GPT for generation instead |
Reference Files
| File | Purpose |
|---|---|
podcast_pipeline.py |
Main pipeline script |
requirements.txt |
Python dependencies |
README.md |
Setup and usage guide |
| 1 | |
| 2 | name podcast-pipeline |
| 3 | description >- |
| 4 | Podcast-to-Everything content pipeline. Takes a podcast RSS feed or raw |
| 5 | transcript and generates a full cross-platform content calendar: short-form |
| 6 | video clips, Twitter/X threads, LinkedIn articles, newsletter sections, quote |
| 7 | cards, blog outlines with SEO keywords, and YouTube Shorts/TikTok scripts. |
| 8 | Scores each piece by viral potential (novelty × controversy × utility) and |
| 9 | deduplicates against recent output. Use when asked to: "repurpose this podcast", |
| 10 | "turn this episode into content", "podcast content calendar", "extract clips |
| 11 | from this episode", "podcast to social", "content from RSS feed", "batch |
| 12 | process episodes", or any request to turn podcast/audio content into a |
| 13 | multi-platform content plan. |
| 14 | |
| 15 | |
| 16 | |
| 17 | ## Preamble (runs on skill start) |
| 18 | |
| 19 | |
| 20 | # Version check (silent if up to date) |
| 21 | python3 telemetry/version_check.py 2>/dev/null || true |
| 22 | |
| 23 | # Telemetry opt-in (first run only, then remembers your choice) |
| 24 | python3 telemetry/telemetry_init.py 2>/dev/null || true |
| 25 | |
| 26 | |
| 27 | > **Privacy:** This skill logs usage locally to `~/.ai-marketing-skills/analytics/`. Remote telemetry is opt-in only. No code, file paths, or repo content is ever collected. See `telemetry/README.md`. |
| 28 | |
| 29 | |
| 30 | |
| 31 | # Podcast-to-Everything Pipeline |
| 32 | |
| 33 | Turns podcast episodes into a full content calendar across every platform. |
| 34 | One episode in, 15-20 content pieces out — scored, deduplicated, and scheduled. |
| 35 | |
| 36 | |
| 37 | |
| 38 | ## Step 1: Ingest — Get the Transcript |
| 39 | |
| 40 | Determine the input source and obtain a clean transcript. |
| 41 | |
| 42 | ### Option A: RSS Feed (`--rss <url>`) |
| 43 | Fetch the RSS feed XML |
| 44 | Extract the latest episode's audio URL (or use `--episodes N` for batch) |
| 45 | Download the audio file |
| 46 | Transcribe via OpenAI Whisper API (with timestamps) |
| 47 | Store transcript with episode metadata (title, date, description, duration) |
| 48 | |
| 49 | ### Option B: Raw Transcript (`--transcript <file>`) |
| 50 | Read the transcript file (plain text, SRT, or VTT) |
| 51 | Parse timestamps if present |
| 52 | Extract episode metadata from filename or prompt user |
| 53 | |
| 54 | ### Option C: Batch Mode (`--batch <rss_url> --episodes N`) |
| 55 | Fetch RSS feed |
| 56 | Extract the last N episodes |
| 57 | Process each through the full pipeline |
| 58 | Deduplicate across all episodes in the batch |
| 59 | |
| 60 | ### Transcript cleanup |
| 61 | Remove filler words (um, uh, like, you know) for written content |
| 62 | Preserve original with timestamps for video clip suggestions |
| 63 | Split into logical segments by topic shift |
| 64 | |
| 65 | |
| 66 | |
| 67 | ## Step 2: Editorial Brain — Deep Analysis |
| 68 | |
| 69 | Feed the full transcript to the LLM with this extraction framework: |
| 70 | |
| 71 | ### Extract these content atoms: |
| 72 | |
| 73 | **Narrative Arcs** — Complete story segments with setup → tension → resolution. |
| 74 | Tag with start/end timestamps. |
| 75 | |
| 76 | **Quotable Moments** — Punchy, shareable statements. One-liners that stand alone. |
| 77 | Must pass the "would someone screenshot this?" test. |
| 78 | |
| 79 | **Controversial Takes** — Opinions that go against conventional wisdom. |
| 80 | The stuff that makes people reply "hard disagree" or "finally someone said it." |
| 81 | |
| 82 | **Data Points** — Specific numbers, percentages, dollar amounts, timeframes. |
| 83 | Concrete proof points that add credibility. |
| 84 | |
| 85 | **Stories** — Personal anecdotes, case studies, client examples. |
| 86 | Must have a character, a problem, and an outcome. |
| 87 | |
| 88 | **Frameworks** — Step-by-step processes, mental models, decision matrices. |
| 89 | Anything structured that people would save or bookmark. |
| 90 | |
| 91 | **Predictions** — Forward-looking claims about trends, markets, technology. |
| 92 | Hot takes about where things are going. |
| 93 | |
| 94 | ### Output format per atom: |
| 95 | |
| 96 | - Type: [narrative_arc | quote | controversial_take | data_point | story | framework | prediction] |
| 97 | - Content: [extracted text] |
| 98 | - Timestamp: [start - end, if available] |
| 99 | - Context: [what was being discussed] |
| 100 | - Viral Score: [0-100, see Step 4] |
| 101 | - Suggested platforms: [where this atom works best] |
| 102 | |
| 103 | |
| 104 | |
| 105 | |
| 106 | ## Step 3: Content Generation — One Episode, Many Pieces |
| 107 | |
| 108 | For each episode, generate ALL of these from the extracted atoms: |
| 109 | |
| 110 | ### 3a. Short-Form Video Clips (3-5 per episode) |
| 111 | |
| 112 | - Hook: [First 3 seconds — pattern interrupt or bold claim] |
| 113 | - Clip segment: [Timestamp range from transcript] |
| 114 | - Caption overlay: [Text for the screen] |
| 115 | - Platform: [YouTube Shorts / TikTok / Instagram Reels] |
| 116 | - Why it works: [What makes this clippable] |
| 117 | |
| 118 | Prioritize: controversial takes > stories with payoffs > surprising data points |
| 119 | |
| 120 | ### 3b. Twitter/X Threads (2-3 per episode) |
| 121 | |
| 122 | - Thread hook (tweet 1): [Curiosity gap or bold opener] |
| 123 | - Thread body (5-10 tweets): [Each tweet is one complete thought] |
| 124 | - Thread closer: [CTA — follow, reply, retweet trigger] |
| 125 | - Source atoms: [Which content atoms feed this thread] |
| 126 | |
| 127 | Rules: No tweet over 280 chars. Each tweet must stand alone. Use data points as proof. |
| 128 | |
| 129 | ### 3c. LinkedIn Article Draft (1 per episode) |
| 130 | |
| 131 | - Headline: [Specific, benefit-driven] |
| 132 | - Hook paragraph: [Before the "see more" fold — must earn the click] |
| 133 | - Body: [3-5 sections with headers, 800-1200 words] |
| 134 | - CTA: [Engagement driver — question, not link] |
| 135 | - Hashtags: [3-5 relevant, not spammy] |
| 136 | |
| 137 | Voice: Professional but not corporate. First-person. Story-driven. |
| 138 | |
| 139 | ### 3d. Newsletter Section (1 per episode) |
| 140 | |
| 141 | - Section headline: [Scannable, specific] |
| 142 | - TL;DR: [One sentence, the core insight] |
| 143 | - Body: [3-5 bullet points, each with a takeaway] |
| 144 | - Pull quote: [The most shareable line from the episode] |
| 145 | - Link: [Back to full episode] |
| 146 | |
| 147 | |
| 148 | ### 3e. Quote Cards (3-5 per episode) |
| 149 | |
| 150 | - Quote text: [Max 20 words — must work as text overlay] |
| 151 | - Attribution: [Speaker name] |
| 152 | - Background suggestion: [Color/mood that matches the tone] |
| 153 | - Platform sizing: [1080x1080 for IG, 1200x675 for Twitter, 1080x1920 for Stories] |
| 154 | |
| 155 | |
| 156 | ### 3f. Blog Post Outline (1 per episode) |
| 157 | |
| 158 | - Title: [SEO-optimized, includes primary keyword] |
| 159 | - Primary keyword: [Search volume + difficulty estimate] |
| 160 | - Secondary keywords: [3-5 related terms] |
| 161 | - Meta description: [155 chars max] |
| 162 | - H2 sections: [5-7, each maps to a content atom] |
| 163 | - Internal linking opportunities: [Topics that connect to existing content] |
| 164 | - Estimated word count: [1500-2500] |
| 165 | |
| 166 | |
| 167 | ### 3g. YouTube Shorts / TikTok Script (1 per episode) |
| 168 | |
| 169 | - HOOK (0-3s): [Pattern interrupt — question, bold claim, or visual] |
| 170 | - SETUP (3-15s): [Context — why should they care] |
| 171 | - PAYOFF (15-45s): [The insight, data, or story resolution] |
| 172 | - CTA (45-60s): [Follow, comment prompt, or part 2 tease] |
| 173 | - On-screen text: [Key phrases to overlay] |
| 174 | - B-roll suggestions: [Visual ideas if not talking-head] |
| 175 | |
| 176 | |
| 177 | |
| 178 | |
| 179 | ## Step 4: Content Scoring — Viral Potential |
| 180 | |
| 181 | Score every generated piece on three dimensions (each 0-100): |
| 182 | |
| 183 | | Dimension | What It Measures | Signals | |
| 184 | |-----------|-----------------|---------| |
| 185 | | **Novelty** | Is this new or surprising? | Contrarian takes, unexpected data, first-to-say | |
| 186 | | **Controversy** | Will people argue about this? | Strong opinions, challenges norms, picks a side | |
| 187 | | **Utility** | Can someone use this immediately? | Frameworks, how-tos, templates, specific numbers | |
| 188 | |
| 189 | **Viral Score = (Novelty × 0.4) + (Controversy × 0.3) + (Utility × 0.3)** |
| 190 | |
| 191 | ### Score thresholds: |
| 192 | **80+** → Priority publish. Schedule for peak engagement windows. |
| 193 | **60-79** → Solid content. Fill the calendar. |
| 194 | **40-59** → Filler. Use only if calendar has gaps. |
| 195 | **Below 40** → Cut it. Not worth the publish slot. |
| 196 | |
| 197 | |
| 198 | |
| 199 | ## Step 5: Dedup Engine |
| 200 | |
| 201 | Before finalizing, check all generated content against: |
| 202 | **This batch** — No two pieces should cover the same angle |
| 203 | **Recent history** — Compare against last N days of output (default: 30) |
| 204 | **Similarity threshold** — Flag any pair with >70% semantic overlap |
| 205 | |
| 206 | ### Dedup rules: |
| 207 | If two pieces overlap >70%: keep the higher-scored one, cut the other |
| 208 | If a piece overlaps with recently published content: flag with ⚠️ and suggest a differentiation angle |
| 209 | Track all published content hashes in `output/content_history.json` |
| 210 | |
| 211 | |
| 212 | |
| 213 | ## Step 6: Calendar Generation (`--calendar`) |
| 214 | |
| 215 | Assemble scored, deduplicated content into a weekly publish calendar. |
| 216 | |
| 217 | ### Scheduling rules: |
| 218 | **Twitter/X:** 1-2 per day, peak hours (8-10am, 12-1pm, 5-7pm ET) |
| 219 | **LinkedIn:** 1 per day max, Tuesday-Thursday mornings |
| 220 | **YouTube Shorts/TikTok:** 1 per day, evenings |
| 221 | **Newsletter:** Weekly, same day each week |
| 222 | **Blog:** 1-2 per week |
| 223 | **Quote cards:** Intersperse on low-content days |
| 224 | |
| 225 | ### Calendar output format: |
| 226 | |
| 227 | { |
| 228 | "week_of": "2024-01-15", |
| 229 | "episode_source": "Episode Title - Guest Name", |
| 230 | "content_pieces": [ |
| 231 | { |
| 232 | "date": "2024-01-15", |
| 233 | "time": "09:00 ET", |
| 234 | "platform": "twitter", |
| 235 | "type": "thread", |
| 236 | "content": "...", |
| 237 | "viral_score": 85, |
| 238 | "status": "draft" |
| 239 | } |
| 240 | ], |
| 241 | "total_pieces": 18, |
| 242 | "avg_viral_score": 72, |
| 243 | "coverage": { |
| 244 | "twitter": 6, |
| 245 | "linkedin": 3, |
| 246 | "youtube_shorts": 3, |
| 247 | "newsletter": 1, |
| 248 | "blog": 1, |
| 249 | "quote_cards": 4 |
| 250 | } |
| 251 | } |
| 252 | |
| 253 | |
| 254 | |
| 255 | |
| 256 | ## Step 7: Output |
| 257 | |
| 258 | All output goes to `output/` directory: |
| 259 | |
| 260 | |
| 261 | output/ |
| 262 | ├── episodes/ |
| 263 | │ ├── YYYY-MM-DD-episode-slug/ |
| 264 | │ │ ├── transcript.txt |
| 265 | │ │ ├── atoms.json # Extracted content atoms |
| 266 | │ │ ├── content_pieces.json # All generated content |
| 267 | │ │ └── calendar.json # Scheduled calendar |
| 268 | │ └── ... |
| 269 | ├── calendar/ |
| 270 | │ └── week-YYYY-WNN.json # Aggregated weekly calendar |
| 271 | ├── content_history.json # Dedup tracking |
| 272 | └── pipeline_log.json # Run history and stats |
| 273 | |
| 274 | |
| 275 | |
| 276 | |
| 277 | ## CLI Reference |
| 278 | |
| 279 | |
| 280 | # Process latest episode from RSS feed |
| 281 | python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml" |
| 282 | |
| 283 | # Process a local transcript |
| 284 | python podcast_pipeline.py --transcript episode-42.txt |
| 285 | |
| 286 | # Batch process last 5 episodes |
| 287 | python podcast_pipeline.py --batch "https://feeds.example.com/podcast.xml" --episodes 5 |
| 288 | |
| 289 | # Generate weekly calendar from existing outputs |
| 290 | python podcast_pipeline.py --calendar |
| 291 | |
| 292 | # Process with custom dedup window |
| 293 | python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml" --dedup-days 60 |
| 294 | |
| 295 | # Process and only keep 80+ viral score content |
| 296 | python podcast_pipeline.py --rss "https://feeds.example.com/podcast.xml" --min-score 80 |
| 297 | |
| 298 | |
| 299 | |
| 300 | |
| 301 | ## Environment Variables |
| 302 | |
| 303 | | Variable | Required | Description | |
| 304 | |----------|----------|-------------| |
| 305 | | `OPENAI_API_KEY` | Yes (for Whisper) | OpenAI API key for audio transcription | |
| 306 | | `ANTHROPIC_API_KEY` | Yes (for generation) | Anthropic API key for content generation | |
| 307 | | `OPENAI_LLM_KEY` | Optional | Separate OpenAI key if using GPT for generation instead | |
| 308 | |
| 309 | |
| 310 | |
| 311 | ## Reference Files |
| 312 | |
| 313 | | File | Purpose | |
| 314 | |------|---------| |
| 315 | | `podcast_pipeline.py` | Main pipeline script | |
| 316 | | `requirements.txt` | Python dependencies | |
| 317 | | `README.md` | Setup and usage guide | |
| 318 |
Discussion
Alternatives
Browse more free Claude skills or everything in Content creator.