Article extractor skill

Extract clean article content from URLs (blog posts, articles, tutorials) and save as readable text.

by michalparkola·MIT license·★ 553 Stars on the repo·GitHub ↗

Use now

Files of Article extractor

michalparkola/main1 file shown
SKILL.md
Show the full text370 lines

Article Extractor

This skill extracts the main content from web articles and blog posts, removing navigation, ads, newsletter signups, and other clutter. Saves clean, readable text.

When to Use This Skill

Activate when the user:

  • Provides an article/blog URL and wants the text content
  • Asks to "download this article"
  • Wants to "extract the content from [URL]"
  • Asks to "save this blog post as text"
  • Needs clean article text without distractions

How It Works

Priority Order:
  1. Check if tools are installed (reader or trafilatura)
  2. Download and extract article using best available tool
  3. Clean up the content (remove extra whitespace, format properly)
  4. Save to file with article title as filename
  5. Confirm location and show preview

Installation Check

Check for article extraction tools in this order:

command -v reader

If not installed:

npm install -g @mozilla/readability-cli
# or
npm install -g reader-cli
Option 2: trafilatura (Python-based, very good)
command -v trafilatura

If not installed:

pip3 install trafilatura
Option 3: Fallback (curl + simple parsing)

If no tools available, use basic curl + text extraction (less reliable but works)

Extraction Methods

Method 1: Using reader (Best for most articles)
# Extract article
reader "URL" > article.txt

Pros:

  • Based on Mozilla's Readability algorithm
  • Excellent at removing clutter
  • Preserves article structure
Method 2: Using trafilatura (Best for blogs/news)
# Extract article
trafilatura --URL "URL" --output-format txt > article.txt

# Or with more options
trafilatura --URL "URL" --output-format txt --no-comments --no-tables > article.txt

Pros:

  • Very accurate extraction
  • Good with various site structures
  • Handles multiple languages

Options:

  • --no-comments: Skip comment sections
  • --no-tables: Skip data tables
  • --precision: Favor precision over recall
  • --recall: Extract more content (may include some noise)
Method 3: Fallback (curl + basic parsing)
# Download and extract basic content
curl -s "URL" | python3 -c "
from html.parser import HTMLParser
import sys

class ArticleExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_content = False
        self.content = []
        self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside'}
        self.current_tag = None

    def handle_starttag(self, tag, attrs):
        if tag not in self.skip_tags:
            if tag in {'p', 'article', 'main', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6'}:
                self.in_content = True
        self.current_tag = tag

    def handle_data(self, data):
        if self.in_content and data.strip():
            self.content.append(data.strip())

    def get_content(self):
        return '\n\n'.join(self.content)

parser = ArticleExtractor()
parser.feed(sys.stdin.read())
print(parser.get_content())
" > article.txt

Note: This is less reliable but works without dependencies.

Getting Article Title

Extract title for filename:

Using reader:
# reader outputs markdown with title at top
TITLE=$(reader "URL" | head -n 1 | sed 's/^# //')
Using trafilatura:
# Get metadata including title
TITLE=$(trafilatura --URL "URL" --json | python3 -c "import json, sys; print(json.load(sys.stdin)['title'])")
Using curl (fallback):
TITLE=$(curl -s "URL" | grep -oP '<title>\K[^<]+' | sed 's/ - .*//' | sed 's/ | .*//')

Filename Creation

Clean title for filesystem:

# Get title
TITLE="Article Title from Website"

# Clean for filesystem (remove special chars, limit length)
FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<' '' | tr '>' '' | tr '|' '-' | cut -c 1-100 | sed 's/ *$//')

# Add extension
FILENAME="${FILENAME}.txt"

Complete Workflow

ARTICLE_URL="https://example.com/article"

# Check for tools
if command -v reader &> /dev/null; then
    TOOL="reader"
    echo "Using reader (Mozilla Readability)"
elif command -v trafilatura &> /dev/null; then
    TOOL="trafilatura"
    echo "Using trafilatura"
else
    TOOL="fallback"
    echo "Using fallback method (may be less accurate)"
fi

# Extract article
case $TOOL in
    reader)
        # Get content
        reader "$ARTICLE_URL" > temp_article.txt

        # Get title (first line after # in markdown)
        TITLE=$(head -n 1 temp_article.txt | sed 's/^# //')
        ;;

    trafilatura)
        # Get title from metadata
        METADATA=$(trafilatura --URL "$ARTICLE_URL" --json)
        TITLE=$(echo "$METADATA" | python3 -c "import json, sys; print(json.load(sys.stdin).get('title', 'Article'))")

        # Get clean content
        trafilatura --URL "$ARTICLE_URL" --output-format txt --no-comments > temp_article.txt
        ;;

    fallback)
        # Get title
        TITLE=$(curl -s "$ARTICLE_URL" | grep -oP '<title>\K[^<]+' | head -n 1)
        TITLE=${TITLE%% - *}  # Remove site name
        TITLE=${TITLE%% | *}  # Remove site name (alternate)

        # Get content (basic extraction)
        curl -s "$ARTICLE_URL" | python3 -c "
from html.parser import HTMLParser
import sys

class ArticleExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_content = False
        self.content = []
        self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside', 'form'}

    def handle_starttag(self, tag, attrs):
        if tag not in self.skip_tags:
            if tag in {'p', 'article', 'main'}:
                self.in_content = True
        if tag in {'h1', 'h2', 'h3'}:
            self.content.append('\n')

    def handle_data(self, data):
        if self.in_content and data.strip():
            self.content.append(data.strip())

    def get_content(self):
        return '\n\n'.join(self.content)

parser = ArticleExtractor()
parser.feed(sys.stdin.read())
print(parser.get_content())
" > temp_article.txt
        ;;
esac

# Clean filename
FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<>' '' | tr '|' '-' | cut -c 1-80 | sed 's/ *$//' | sed 's/^ *//')
FILENAME="${FILENAME}.txt"

# Move to final filename
mv temp_article.txt "$FILENAME"

# Show result
echo "✓ Extracted article: $TITLE"
echo "✓ Saved to: $FILENAME"
echo ""
echo "Preview (first 10 lines):"
head -n 10 "$FILENAME"

Error Handling

Common Issues

1. Tool not installed

  • Try alternate tool (reader → trafilatura → fallback)
  • Offer to install: "Install reader with: npm install -g reader-cli"

2. Paywall or login required

  • Extraction tools may fail
  • Inform user: "This article requires authentication. Cannot extract."

3. Invalid URL

  • Check URL format
  • Try with and without redirects

4. No content extracted

  • Site may use heavy JavaScript
  • Try fallback method
  • Inform user if extraction fails

5. Special characters in title

  • Clean title for filesystem
  • Remove: /, :, ?, ", <, >, |
  • Replace with - or remove

Output Format

Saved File Contains:
  • Article title (if available)
  • Author (if available from tool)
  • Main article text
  • Section headings
  • No navigation, ads, or clutter
What Gets Removed:
  • Navigation menus
  • Ads and promotional content
  • Newsletter signup forms
  • Related articles sidebars
  • Comment sections (optional)
  • Social media buttons
  • Cookie notices

Tips for Best Results

1. Use reader for most articles

  • Best all-around tool
  • Based on Firefox Reader View
  • Works on most news sites and blogs

2. Use trafilatura for:

  • Academic articles
  • News sites
  • Blogs with complex layouts
  • Non-English content

3. Fallback method limitations:

  • May include some noise
  • Less accurate paragraph detection
  • Better than nothing for simple sites

4. Check extraction quality:

  • Always show preview to user
  • Ask if it looks correct
  • Offer to try different tool if needed

Example Usage

Simple extraction:

# User: "Extract https://example.com/article"
reader "https://example.com/article" > temp.txt
TITLE=$(head -n 1 temp.txt | sed 's/^# //')
FILENAME="$(echo "$TITLE" | tr '/' '-').txt"
mv temp.txt "$FILENAME"
echo "✓ Saved to: $FILENAME"

With error handling:

if ! reader "$URL" > temp.txt 2>/dev/null; then
    if command -v trafilatura &> /dev/null; then
        trafilatura --URL "$URL" --output-format txt > temp.txt
    else
        echo "Error: Could not extract article. Install reader or trafilatura."
        exit 1
    fi
fi

Best Practices

  • ✅ Always show preview after extraction (first 10 lines)
  • ✅ Verify extraction succeeded before saving
  • ✅ Clean filename for filesystem compatibility
  • ✅ Try fallback method if primary fails
  • ✅ Inform user which tool was used
  • ✅ Keep filename length reasonable (< 100 chars)

After Extraction

Display to user:

  1. "✓ Extracted: [Article Title]"
  2. "✓ Saved to: [filename]"
  3. Show preview (first 10-15 lines)
  4. File size and location

Ask if needed:

  • "Would you like me to also create a Ship-Learn-Next plan from this?" (if using ship-learn-next skill)
  • "Should I extract another article?"
1---
2name: article-extractor
3description: Extract clean article content from URLs (blog posts, articles, tutorials) and save as readable text. Use when user wants to download, extract, or save an article/blog post from a URL without ads, navigation, or clutter.
4allowed-tools: Bash,Write
5---
6 
7# Article Extractor
8 
9This skill extracts the main content from web articles and blog posts, removing navigation, ads, newsletter signups, and other clutter. Saves clean, readable text.
10 
11## When to Use This Skill
12 
13Activate when the user:
14- Provides an article/blog URL and wants the text content
15- Asks to "download this article"
16- Wants to "extract the content from [URL]"
17- Asks to "save this blog post as text"
18- Needs clean article text without distractions
19 
20## How It Works
21 
22### Priority Order:
231. **Check if tools are installed** (reader or trafilatura)
242. **Download and extract article** using best available tool
253. **Clean up the content** (remove extra whitespace, format properly)
264. **Save to file** with article title as filename
275. **Confirm location** and show preview
28 
29## Installation Check
30 
31Check for article extraction tools in this order:
32 
33### Option 1: reader (Recommended - Mozilla's Readability)
34 
35```bash
36command -v reader
37```
38 
39If not installed:
40```bash
41npm install -g @mozilla/readability-cli
42# or
43npm install -g reader-cli
44```
45 
46### Option 2: trafilatura (Python-based, very good)
47 
48```bash
49command -v trafilatura
50```
51 
52If not installed:
53```bash
54pip3 install trafilatura
55```
56 
57### Option 3: Fallback (curl + simple parsing)
58 
59If no tools available, use basic curl + text extraction (less reliable but works)
60 
61## Extraction Methods
62 
63### Method 1: Using reader (Best for most articles)
64 
65```bash
66# Extract article
67reader "URL" > article.txt
68```
69 
70**Pros:**
71- Based on Mozilla's Readability algorithm
72- Excellent at removing clutter
73- Preserves article structure
74 
75### Method 2: Using trafilatura (Best for blogs/news)
76 
77```bash
78# Extract article
79trafilatura --URL "URL" --output-format txt > article.txt
80 
81# Or with more options
82trafilatura --URL "URL" --output-format txt --no-comments --no-tables > article.txt
83```
84 
85**Pros:**
86- Very accurate extraction
87- Good with various site structures
88- Handles multiple languages
89 
90**Options:**
91- `--no-comments`: Skip comment sections
92- `--no-tables`: Skip data tables
93- `--precision`: Favor precision over recall
94- `--recall`: Extract more content (may include some noise)
95 
96### Method 3: Fallback (curl + basic parsing)
97 
98```bash
99# Download and extract basic content
100curl -s "URL" | python3 -c "
101from html.parser import HTMLParser
102import sys
103 
104class ArticleExtractor(HTMLParser):
105 def __init__(self):
106 super().__init__()
107 self.in_content = False
108 self.content = []
109 self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside'}
110 self.current_tag = None
111 
112 def handle_starttag(self, tag, attrs):
113 if tag not in self.skip_tags:
114 if tag in {'p', 'article', 'main', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6'}:
115 self.in_content = True
116 self.current_tag = tag
117 
118 def handle_data(self, data):
119 if self.in_content and data.strip():
120 self.content.append(data.strip())
121 
122 def get_content(self):
123 return '\n\n'.join(self.content)
124 
125parser = ArticleExtractor()
126parser.feed(sys.stdin.read())
127print(parser.get_content())
128" > article.txt
129```
130 
131**Note:** This is less reliable but works without dependencies.
132 
133## Getting Article Title
134 
135Extract title for filename:
136 
137### Using reader:
138```bash
139# reader outputs markdown with title at top
140TITLE=$(reader "URL" | head -n 1 | sed 's/^# //')
141```
142 
143### Using trafilatura:
144```bash
145# Get metadata including title
146TITLE=$(trafilatura --URL "URL" --json | python3 -c "import json, sys; print(json.load(sys.stdin)['title'])")
147```
148 
149### Using curl (fallback):
150```bash
151TITLE=$(curl -s "URL" | grep -oP '<title>\K[^<]+' | sed 's/ - .*//' | sed 's/ | .*//')
152```
153 
154## Filename Creation
155 
156Clean title for filesystem:
157 
158```bash
159# Get title
160TITLE="Article Title from Website"
161 
162# Clean for filesystem (remove special chars, limit length)
163FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<' '' | tr '>' '' | tr '|' '-' | cut -c 1-100 | sed 's/ *$//')
164 
165# Add extension
166FILENAME="${FILENAME}.txt"
167```
168 
169## Complete Workflow
170 
171```bash
172ARTICLE_URL="https://example.com/article"
173 
174# Check for tools
175if command -v reader &> /dev/null; then
176 TOOL="reader"
177 echo "Using reader (Mozilla Readability)"
178elif command -v trafilatura &> /dev/null; then
179 TOOL="trafilatura"
180 echo "Using trafilatura"
181else
182 TOOL="fallback"
183 echo "Using fallback method (may be less accurate)"
184fi
185 
186# Extract article
187case $TOOL in
188 reader)
189 # Get content
190 reader "$ARTICLE_URL" > temp_article.txt
191 
192 # Get title (first line after # in markdown)
193 TITLE=$(head -n 1 temp_article.txt | sed 's/^# //')
194 ;;
195 
196 trafilatura)
197 # Get title from metadata
198 METADATA=$(trafilatura --URL "$ARTICLE_URL" --json)
199 TITLE=$(echo "$METADATA" | python3 -c "import json, sys; print(json.load(sys.stdin).get('title', 'Article'))")
200 
201 # Get clean content
202 trafilatura --URL "$ARTICLE_URL" --output-format txt --no-comments > temp_article.txt
203 ;;
204 
205 fallback)
206 # Get title
207 TITLE=$(curl -s "$ARTICLE_URL" | grep -oP '<title>\K[^<]+' | head -n 1)
208 TITLE=${TITLE%% - *} # Remove site name
209 TITLE=${TITLE%% | *} # Remove site name (alternate)
210 
211 # Get content (basic extraction)
212 curl -s "$ARTICLE_URL" | python3 -c "
213from html.parser import HTMLParser
214import sys
215 
216class ArticleExtractor(HTMLParser):
217 def __init__(self):
218 super().__init__()
219 self.in_content = False
220 self.content = []
221 self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside', 'form'}
222 
223 def handle_starttag(self, tag, attrs):
224 if tag not in self.skip_tags:
225 if tag in {'p', 'article', 'main'}:
226 self.in_content = True
227 if tag in {'h1', 'h2', 'h3'}:
228 self.content.append('\n')
229 
230 def handle_data(self, data):
231 if self.in_content and data.strip():
232 self.content.append(data.strip())
233 
234 def get_content(self):
235 return '\n\n'.join(self.content)
236 
237parser = ArticleExtractor()
238parser.feed(sys.stdin.read())
239print(parser.get_content())
240" > temp_article.txt
241 ;;
242esac
243 
244# Clean filename
245FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<>' '' | tr '|' '-' | cut -c 1-80 | sed 's/ *$//' | sed 's/^ *//')
246FILENAME="${FILENAME}.txt"
247 
248# Move to final filename
249mv temp_article.txt "$FILENAME"
250 
251# Show result
252echo "✓ Extracted article: $TITLE"
253echo "✓ Saved to: $FILENAME"
254echo ""
255echo "Preview (first 10 lines):"
256head -n 10 "$FILENAME"
257```
258 
259## Error Handling
260 
261### Common Issues
262 
263**1. Tool not installed**
264- Try alternate tool (reader → trafilatura → fallback)
265- Offer to install: "Install reader with: npm install -g reader-cli"
266 
267**2. Paywall or login required**
268- Extraction tools may fail
269- Inform user: "This article requires authentication. Cannot extract."
270 
271**3. Invalid URL**
272- Check URL format
273- Try with and without redirects
274 
275**4. No content extracted**
276- Site may use heavy JavaScript
277- Try fallback method
278- Inform user if extraction fails
279 
280**5. Special characters in title**
281- Clean title for filesystem
282- Remove: `/`, `:`, `?`, `"`, `<`, `>`, `|`
283- Replace with `-` or remove
284 
285## Output Format
286 
287### Saved File Contains:
288- Article title (if available)
289- Author (if available from tool)
290- Main article text
291- Section headings
292- No navigation, ads, or clutter
293 
294### What Gets Removed:
295- Navigation menus
296- Ads and promotional content
297- Newsletter signup forms
298- Related articles sidebars
299- Comment sections (optional)
300- Social media buttons
301- Cookie notices
302 
303## Tips for Best Results
304 
305**1. Use reader for most articles**
306- Best all-around tool
307- Based on Firefox Reader View
308- Works on most news sites and blogs
309 
310**2. Use trafilatura for:**
311- Academic articles
312- News sites
313- Blogs with complex layouts
314- Non-English content
315 
316**3. Fallback method limitations:**
317- May include some noise
318- Less accurate paragraph detection
319- Better than nothing for simple sites
320 
321**4. Check extraction quality:**
322- Always show preview to user
323- Ask if it looks correct
324- Offer to try different tool if needed
325 
326## Example Usage
327 
328**Simple extraction:**
329```bash
330# User: "Extract https://example.com/article"
331reader "https://example.com/article" > temp.txt
332TITLE=$(head -n 1 temp.txt | sed 's/^# //')
333FILENAME="$(echo "$TITLE" | tr '/' '-').txt"
334mv temp.txt "$FILENAME"
335echo "✓ Saved to: $FILENAME"
336```
337 
338**With error handling:**
339```bash
340if ! reader "$URL" > temp.txt 2>/dev/null; then
341 if command -v trafilatura &> /dev/null; then
342 trafilatura --URL "$URL" --output-format txt > temp.txt
343 else
344 echo "Error: Could not extract article. Install reader or trafilatura."
345 exit 1
346 fi
347fi
348```
349 
350## Best Practices
351 
352- ✅ Always show preview after extraction (first 10 lines)
353- ✅ Verify extraction succeeded before saving
354- ✅ Clean filename for filesystem compatibility
355- ✅ Try fallback method if primary fails
356- ✅ Inform user which tool was used
357- ✅ Keep filename length reasonable (< 100 chars)
358 
359## After Extraction
360 
361Display to user:
3621. "✓ Extracted: [Article Title]"
3632. "✓ Saved to: [filename]"
3643. Show preview (first 10-15 lines)
3654. File size and location
366 
367Ask if needed:
368- "Would you like me to also create a Ship-Learn-Next plan from this?" (if using ship-learn-next skill)
369- "Should I extract another article?"
370 

Discussion

Alternatives