Apify skill
Scrapes social platforms, business data, and e-commerce via Apify actors — Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps, Amazon, and web crawls — filtering in code.
by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗
npx degit danielmiessler/LifeOS/LifeOS/install/skills/Apify#main ~/.claude/skills/apifyChecked ·commit main
Files of Apify
Show the full text520 lines
Customization
Before executing, check for user customizations at:
~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Apify/
If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
You MUST send this notification BEFORE doing anything else when this skill is invoked.
Send voice notification:
curl -s -X POST http://localhost:31337/notify \ -H "Content-Type: application/json" \ -d '{"message": "Running the WORKFLOWNAME workflow in the Apify skill to ACTION"}' \ > /dev/null 2>&1 &Output text notification:
Running the **WorkflowName** workflow in the **Apify** skill to ACTION...
This is not optional. Execute this curl command immediately upon skill invocation.
Apify - Social Media & Web Scraping
What It Does
Scrapes social platforms, business data, and e-commerce through Apify actors: Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps business search, Amazon, and general-purpose web crawling. TypeScript wrappers filter and transform the data in code before any of it reaches the model, so a 100-post scrape costs roughly what 10 posts would. Runs platforms in parallel for social-listening dashboards and chains Google Maps into LinkedIn for lead enrichment.
The Problem
Scraping through a raw MCP dumps every unfiltered result straight into model context — a single Instagram profile with 100 posts burns ~52,000 tokens, most of it noise you'll throw away. You usually want the top 10 posts, the negative reviews from the last week, the qualified leads with an email. Doing that filtering after the data hits the model is too late; the tokens are already spent. Filtering in code first cuts that 52,000 down to ~500.
How It Works
This skill is a file-based MCP — a code-first API wrapper that replaces token-heavy MCP protocol calls. You call an actor wrapper, filter and sort the result in TypeScript, and only the filtered slice reaches model context. That code-before-context step is where the 95-99% token savings come from.
Workflow Routing
| Workflow | Trigger | File |
|---|---|---|
| Update | update Apify skill, refresh actors, actor calls failing unexpectedly, monthly capability check | Workflows/Update.md |
| (inline) | all scrape/lead/crawl requests — scrape Instagram/LinkedIn/TikTok/YouTube/Facebook, Google Maps leads, Amazon reviews, web crawl | Actor wrappers under actors/ (see Actor Reference below) |
📊 Available Actors
Social Media (5 platforms)
- Instagram (145k users, 4.60★) - Profiles, posts, hashtags, comments
- LinkedIn (26k users, 4.10★) - Profiles, jobs, posts
- TikTok (90k users, 4.61★) - Profiles, videos, hashtags, comments
- YouTube (40k users, 4.40★) - Channels, videos, comments, search
- Facebook (35k users, 4.56★) - Posts, groups, comments
Business & Lead Generation
- Google Maps (198k users, 4.76★) - HIGHEST VALUE!
- Search businesses, extract contacts, reviews, images
- Perfect for lead generation
E-commerce
- Amazon (8k users, 4.97★) - Products, reviews, pricing
Web Scraping
- Web Scraper (94k users, 4.39★) - General-purpose, works with ANY website
🚀 Quick Start
Basic Usage Pattern
import { scrapeInstagramProfile, searchGoogleMaps } from 'actors'
// 1. Call the actor wrapper
const profile = await scrapeInstagramProfile({
username: 'target_username',
maxPosts: 50
})
// 2. Filter in code - BEFORE data reaches model!
const viral = profile.latestPosts?.filter(p => p.likesCount > 10000)
// 3. Only filtered results reach model context
console.log(viral) // ~10 posts instead of 50
📚 Examples by Use Case
Social Media Monitoring
Instagram - Track engagement:
import { scrapeInstagramProfile, scrapeInstagramPosts } from 'actors'
// Get profile with recent posts
const profile = await scrapeInstagramProfile({
username: 'competitor',
maxPosts: 100
})
// Filter in code - only high-performing posts from last 30 days
const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
const topRecent = profile.latestPosts
?.filter(p =>
new Date(p.timestamp).getTime() > thirtyDaysAgo &&
p.likesCount > 5000
)
.sort((a, b) => b.likesCount - a.likesCount)
.slice(0, 10)
// Only 10 posts reach model instead of 100!
LinkedIn - Job search:
import { searchLinkedInJobs } from 'actors'
const jobs = await searchLinkedInJobs({
keywords: 'AI engineer',
location: 'San Francisco',
remote: true,
maxResults: 200
})
// Filter in code - only senior roles at well-funded startups
const topJobs = jobs.filter(j =>
j.seniority?.includes('Senior') &&
parseInt(j.applicants || '0') > 50
)
TikTok - Trend analysis:
import { scrapeTikTokHashtag } from 'actors'
const videos = await scrapeTikTokHashtag({
hashtag: 'ai',
maxResults: 500
})
// Filter in code - only viral content
const viral = videos
.filter(v => v.playCount > 1000000)
.sort((a, b) => b.playCount - a.playCount)
.slice(0, 20)
Lead Generation (Business Intelligence)
Google Maps - Local business leads:
import { searchGoogleMaps } from 'actors'
// Search with contact info extraction
const places = await searchGoogleMaps({
query: 'restaurants in Austin',
maxResults: 500,
includeReviews: true,
maxReviewsPerPlace: 20,
scrapeContactInfo: true // Extracts emails from websites!
})
// Filter in code - only highly-rated with email/phone
const qualifiedLeads = places
.filter(p =>
p.rating >= 4.5 &&
p.reviewsCount >= 100 &&
(p.email || p.phone)
)
.map(p => ({
name: p.name,
rating: p.rating,
reviews: p.reviewsCount,
email: p.email,
phone: p.phone,
website: p.website,
address: p.address
}))
// Export leads - only qualified results!
console.log(`Found ${qualifiedLeads.length} qualified leads`)
Google Maps - Review sentiment analysis:
import { scrapeGoogleMapsReviews } from 'actors'
const reviews = await scrapeGoogleMapsReviews({
placeUrl: 'https://maps.google.com/maps?cid=12345',
maxResults: 1000
})
// Filter in code - analyze sentiment by rating
const recentNegative = reviews
.filter(r => {
const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
return (
r.rating <= 2 &&
new Date(r.publishedAtDate).getTime() > thirtyDaysAgo &&
r.text.length > 50
)
})
// Identify common complaints
const complaints = recentNegative.map(r => r.text)
E-commerce & Competitive Intelligence
Amazon - Price monitoring:
import { scrapeAmazonProduct } from 'actors'
const product = await scrapeAmazonProduct({
productUrl: 'https://www.amazon.com/dp/B08L5VT894',
includeReviews: true,
maxReviews: 200
})
// Filter in code - only recent negative reviews
const recentNegative = product.reviews
?.filter(r => {
const weekAgo = Date.now() - (7 * 24 * 60 * 60 * 1000)
return (
r.rating <= 2 &&
new Date(r.date).getTime() > weekAgo
)
})
console.log(`Price: $${product.price}`)
console.log(`Rating: ${product.rating}/5`)
console.log(`Recent issues: ${recentNegative?.length} complaints`)
Custom Web Scraping
Any Website - Custom extraction:
import { scrapeWebsite } from 'actors'
const products = await scrapeWebsite({
startUrls: ['https://example.com/products'],
linkSelector: 'a.product-link',
maxPagesPerCrawl: 100,
pageFunction: `
async function pageFunction(context) {
const { request, $, log } = context
return {
url: request.url,
title: $('h1.product-title').text(),
price: $('span.price').text(),
inStock: $('.in-stock').length > 0,
description: $('.description').text()
}
}
`
})
// Filter in code - only available products under $100
const affordable = products.filter(p =>
p.inStock &&
parseFloat(p.price.replace('$', '')) < 100
)
🎨 Advanced Patterns
Pattern 1: Multi-Platform Social Listening
import {
scrapeInstagramHashtag,
scrapeTikTokHashtag,
searchYouTube
} from 'actors'
// Run all platforms in parallel
const [instagramPosts, tiktokVideos, youtubeVideos] = await Promise.all([
scrapeInstagramHashtag({ hashtag: 'ai', maxResults: 100 }),
scrapeTikTokHashtag({ hashtag: 'ai', maxResults: 100 }),
searchYouTube({ query: '#ai', maxResults: 100 })
])
// Combine and filter - only viral content across all platforms
const allViral = [
...instagramPosts.filter(p => p.likesCount > 10000),
...tiktokVideos.filter(v => v.playCount > 100000),
...youtubeVideos.filter(v => v.viewsCount > 50000)
]
console.log(`Found ${allViral.length} viral posts across 3 platforms`)
Pattern 2: Lead Enrichment Pipeline
import { searchGoogleMaps, scrapeLinkedInProfile } from 'actors'
// 1. Find businesses on Google Maps
const restaurants = await searchGoogleMaps({
query: 'restaurants in SF',
maxResults: 100,
scrapeContactInfo: true
})
// 2. Filter for qualified leads
const qualified = restaurants.filter(r =>
r.rating >= 4.5 &&
r.email &&
r.reviewsCount >= 50
)
// 3. Enrich with LinkedIn data (if available)
const enriched = await Promise.all(
qualified.map(async (restaurant) => {
// Try to find LinkedIn company page
// ... additional enrichment logic
return restaurant
})
)
Pattern 3: Competitive Analysis Dashboard
import {
scrapeInstagramProfile,
scrapeYouTubeChannel,
scrapeTikTokProfile
} from 'actors'
async function analyzeCompetitor(username: string) {
// Gather data from all platforms
const [instagram, youtube, tiktok] = await Promise.all([
scrapeInstagramProfile({ username, maxPosts: 30 }),
scrapeYouTubeChannel({ channelUrl: `https://youtube.com/@${username}`, maxVideos: 30 }),
scrapeTikTokProfile({ username, maxVideos: 30 })
])
// Calculate engagement metrics in code
return {
username,
instagram: {
followers: instagram.followersCount,
avgLikes: average(instagram.latestPosts?.map(p => p.likesCount) || []),
engagementRate: calculateEngagement(instagram)
},
youtube: {
subscribers: youtube.subscribersCount,
avgViews: average(youtube.videos?.map(v => v.viewsCount) || [])
},
tiktok: {
followers: tiktok.followersCount,
avgPlays: average(tiktok.videos?.map(v => v.playCount) || [])
}
}
}
💰 Token Savings Calculator
Example: Instagram profile with 100 posts
MCP Approach:
1. search-actors → 1,000 tokens
2. call-actor → 1,000 tokens
3. get-actor-output → 50,000 tokens (100 unfiltered posts)
TOTAL: ~52,000 tokens
File-Based Approach:
const profile = await scrapeInstagramProfile({
username: 'user',
maxPosts: 100
})
// Filter in code - only top 10 posts
const top = profile.latestPosts
?.sort((a, b) => b.likesCount - a.likesCount)
.slice(0, 10)
// TOTAL: ~500 tokens (only 10 filtered posts reach model)
Savings: 99% reduction (52,000 → 500 tokens)
🔧 Actor Reference
Social Media
scrapeInstagramProfile(input)- Profile + postsscrapeInstagramPosts(input)- Posts from userscrapeInstagramHashtag(input)- Posts by hashtagscrapeInstagramComments(input)- Comments on post
scrapeLinkedInProfile(input)- Profile + experience + emailsearchLinkedInJobs(input)- Job listingsscrapeLinkedInPosts(input)- Posts from profile/company
TikTok
scrapeTikTokProfile(input)- Profile + videosscrapeTikTokHashtag(input)- Videos by hashtagscrapeTikTokComments(input)- Comments on video
YouTube
scrapeYouTubeChannel(input)- Channel + videossearchYouTube(input)- Search videosscrapeYouTubeComments(input)- Comments on video
scrapeFacebookPosts(input)- Posts from pagesscrapeFacebookGroups(input)- Group postsscrapeFacebookComments(input)- Post comments
Business & Lead Generation
Google Maps
searchGoogleMaps(input)- Search places (with contact extraction!)scrapeGoogleMapsPlace(input)- Single place detailsscrapeGoogleMapsReviews(input)- Place reviews
E-commerce
Amazon
scrapeAmazonProduct(input)- Product details + reviewsscrapeAmazonReviews(input)- Product reviews only
Web Scraping
General Web
scrapeWebsite(input)- Custom multi-page crawlingscrapePage(url, pageFunction)- Single page extraction
⚙️ Configuration
Environment Variables:
# Required - Get from https://console.apify.com/account/integrations
APIFY_TOKEN=apify_api_xxxxx...
Actor Run Options:
{
memory: 2048, // MB: 128, 256, 512, 1024, 2048, 4096, 8192
timeout: 300, // seconds
build: 'latest' // or specific build number
}
🎯 When to Use This vs MCP
Use File-Based (this skill):
- ✅ Need to filter large datasets (>100 results)
- ✅ Want to transform/aggregate data in code
- ✅ Multiple sequential operations
- ✅ Control flow (loops, conditionals)
- ✅ Maximum token efficiency
Use MCP:
- ❌ Simple single operations with small results (<10 items)
- ❌ One-off exploratory queries
- ❌ Don't want to write code
🔗 Links
- Apify Platform: https://apify.com
- Actor Store: https://apify.com/store
- API Docs: https://docs.apify.com/api/v2
Remember: Filter data in code BEFORE returning to model context. This is where the 99% token savings happen!
Gotchas
- Actor selection matters. Each social platform has specific actors — don't use a generic scraper for Instagram when a dedicated Instagram actor exists.
- Rate limits vary by platform and plan. Check actor documentation for limits before running large scrapes.
- Scraped data format varies by actor. Read the actor's output schema before processing results.
Examples
Example 1: Scrape Instagram profile
User: "get the recent posts from this Instagram account"
→ Selects Instagram Profile actor
→ Runs with target profile URL
→ Returns structured post data (text, engagement, dates)
Example 2: LinkedIn company scrape
User: "scrape this company's LinkedIn page"
→ Selects LinkedIn Company actor
→ Returns company info, employee count, recent posts
Execution Log
After completing any workflow, append a single JSONL entry:
echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Apify","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
Replace WORKFLOW_USED with the workflow executed, 8_WORD_SUMMARY with a brief input description, and SECONDS with approximate wall-clock time. Log status: "error" if the workflow failed.
| 1 | |
| 2 | name Apify |
| 3 | version 1.1.22 |
| 4 | description "Scrapes social platforms, business data, and e-commerce via Apify actors — Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps, Amazon, and web crawls — filtering in code. USE WHEN scrape Instagram, scrape LinkedIn, scrape TikTok, scrape YouTube, scrape Facebook, Google Maps leads, Amazon reviews, business intelligence, multi-platform social listening, competitive analysis, lead generation, social monitoring, Apify actors, web crawl, extract contacts. NOT FOR X/Twitter account operations like posting, threads, or bookmarks (those need a dedicated X API client), 4-tier progressive scraping with proxy escalation (use BrightData), or real-Chrome bot bypass and computer use (use Interceptor)." |
| 5 | |
| 6 | |
| 7 | ## Customization |
| 8 | |
| 9 | **Before executing, check for user customizations at:** |
| 10 | `~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Apify/` |
| 11 | |
| 12 | If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults. |
| 13 | |
| 14 | |
| 15 | ## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION) |
| 16 | |
| 17 | **You MUST send this notification BEFORE doing anything else when this skill is invoked.** |
| 18 | |
| 19 | **Send voice notification**: |
| 20 | |
| 21 | curl -s -X POST http://localhost:31337/notify \ |
| 22 | -H "Content-Type: application/json" \ |
| 23 | -d '{"message": "Running the WORKFLOWNAME workflow in the Apify skill to ACTION"}' \ |
| 24 | > /dev/null 2>&1 & |
| 25 | |
| 26 | |
| 27 | **Output text notification**: |
| 28 | |
| 29 | Running the **WorkflowName** workflow in the **Apify** skill to ACTION... |
| 30 | |
| 31 | |
| 32 | **This is not optional. Execute this curl command immediately upon skill invocation.** |
| 33 | |
| 34 | # Apify - Social Media & Web Scraping |
| 35 | |
| 36 | ## What It Does |
| 37 | |
| 38 | Scrapes social platforms, business data, and e-commerce through Apify actors: Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps business search, Amazon, and general-purpose web crawling. TypeScript wrappers filter and transform the data in code before any of it reaches the model, so a 100-post scrape costs roughly what 10 posts would. Runs platforms in parallel for social-listening dashboards and chains Google Maps into LinkedIn for lead enrichment. |
| 39 | |
| 40 | ## The Problem |
| 41 | |
| 42 | Scraping through a raw MCP dumps every unfiltered result straight into model context — a single Instagram profile with 100 posts burns ~52,000 tokens, most of it noise you'll throw away. You usually want the top 10 posts, the negative reviews from the last week, the qualified leads with an email. Doing that filtering after the data hits the model is too late; the tokens are already spent. Filtering in code first cuts that 52,000 down to ~500. |
| 43 | |
| 44 | ## How It Works |
| 45 | |
| 46 | This skill is a **file-based MCP** — a code-first API wrapper that replaces token-heavy MCP protocol calls. You call an actor wrapper, filter and sort the result in TypeScript, and only the filtered slice reaches model context. That code-before-context step is where the 95-99% token savings come from. |
| 47 | |
| 48 | ## Workflow Routing |
| 49 | |
| 50 | | Workflow | Trigger | File | |
| 51 | |----------|---------|------| |
| 52 | | Update | update Apify skill, refresh actors, actor calls failing unexpectedly, monthly capability check | `Workflows/Update.md` | |
| 53 | | (inline) | all scrape/lead/crawl requests — scrape Instagram/LinkedIn/TikTok/YouTube/Facebook, Google Maps leads, Amazon reviews, web crawl | Actor wrappers under `actors/` (see Actor Reference below) | |
| 54 | |
| 55 | ## 📊 Available Actors |
| 56 | |
| 57 | ### Social Media (5 platforms) |
| 58 | **Instagram** (145k users, 4.60★) - Profiles, posts, hashtags, comments |
| 59 | **LinkedIn** (26k users, 4.10★) - Profiles, jobs, posts |
| 60 | **TikTok** (90k users, 4.61★) - Profiles, videos, hashtags, comments |
| 61 | **YouTube** (40k users, 4.40★) - Channels, videos, comments, search |
| 62 | **Facebook** (35k users, 4.56★) - Posts, groups, comments |
| 63 | |
| 64 | ### Business & Lead Generation |
| 65 | **Google Maps** (198k users, 4.76★) - **HIGHEST VALUE!** |
| 66 | Search businesses, extract contacts, reviews, images |
| 67 | Perfect for lead generation |
| 68 | |
| 69 | ### E-commerce |
| 70 | **Amazon** (8k users, 4.97★) - Products, reviews, pricing |
| 71 | |
| 72 | ### Web Scraping |
| 73 | **Web Scraper** (94k users, 4.39★) - General-purpose, works with ANY website |
| 74 | |
| 75 | ## 🚀 Quick Start |
| 76 | |
| 77 | ### Basic Usage Pattern |
| 78 | |
| 79 | |
| 80 | import { scrapeInstagramProfile, searchGoogleMaps } from 'actors' |
| 81 | |
| 82 | // 1. Call the actor wrapper |
| 83 | const profile = await scrapeInstagramProfile({ |
| 84 | username: 'target_username', |
| 85 | maxPosts: 50 |
| 86 | }) |
| 87 | |
| 88 | // 2. Filter in code - BEFORE data reaches model! |
| 89 | const viral = profile.latestPosts?.filter(p => p.likesCount > 10000) |
| 90 | |
| 91 | // 3. Only filtered results reach model context |
| 92 | console.log(viral) // ~10 posts instead of 50 |
| 93 | |
| 94 | |
| 95 | ## 📚 Examples by Use Case |
| 96 | |
| 97 | ### Social Media Monitoring |
| 98 | |
| 99 | **Instagram - Track engagement:** |
| 100 | |
| 101 | import { scrapeInstagramProfile, scrapeInstagramPosts } from 'actors' |
| 102 | |
| 103 | // Get profile with recent posts |
| 104 | const profile = await scrapeInstagramProfile({ |
| 105 | username: 'competitor', |
| 106 | maxPosts: 100 |
| 107 | }) |
| 108 | |
| 109 | // Filter in code - only high-performing posts from last 30 days |
| 110 | const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000) |
| 111 | const topRecent = profile.latestPosts |
| 112 | ?.filter(p => |
| 113 | new Date(p.timestamp).getTime() > thirtyDaysAgo && |
| 114 | p.likesCount > 5000 |
| 115 | ) |
| 116 | .sort((a, b) => b.likesCount - a.likesCount) |
| 117 | .slice(0, 10) |
| 118 | |
| 119 | // Only 10 posts reach model instead of 100! |
| 120 | |
| 121 | |
| 122 | **LinkedIn - Job search:** |
| 123 | |
| 124 | import { searchLinkedInJobs } from 'actors' |
| 125 | |
| 126 | const jobs = await searchLinkedInJobs({ |
| 127 | keywords: 'AI engineer', |
| 128 | location: 'San Francisco', |
| 129 | remote: true, |
| 130 | maxResults: 200 |
| 131 | }) |
| 132 | |
| 133 | // Filter in code - only senior roles at well-funded startups |
| 134 | const topJobs = jobs.filter(j => |
| 135 | j.seniority?.includes('Senior') && |
| 136 | parseInt(j.applicants || '0') > 50 |
| 137 | ) |
| 138 | |
| 139 | |
| 140 | **TikTok - Trend analysis:** |
| 141 | |
| 142 | import { scrapeTikTokHashtag } from 'actors' |
| 143 | |
| 144 | const videos = await scrapeTikTokHashtag({ |
| 145 | hashtag: 'ai', |
| 146 | maxResults: 500 |
| 147 | }) |
| 148 | |
| 149 | // Filter in code - only viral content |
| 150 | const viral = videos |
| 151 | .filter(v => v.playCount > 1000000) |
| 152 | .sort((a, b) => b.playCount - a.playCount) |
| 153 | .slice(0, 20) |
| 154 | |
| 155 | |
| 156 | ### Lead Generation (Business Intelligence) |
| 157 | |
| 158 | **Google Maps - Local business leads:** |
| 159 | |
| 160 | import { searchGoogleMaps } from 'actors' |
| 161 | |
| 162 | // Search with contact info extraction |
| 163 | const places = await searchGoogleMaps({ |
| 164 | query: 'restaurants in Austin', |
| 165 | maxResults: 500, |
| 166 | includeReviews: true, |
| 167 | maxReviewsPerPlace: 20, |
| 168 | scrapeContactInfo: true // Extracts emails from websites! |
| 169 | }) |
| 170 | |
| 171 | // Filter in code - only highly-rated with email/phone |
| 172 | const qualifiedLeads = places |
| 173 | .filter(p => |
| 174 | p.rating >= 4.5 && |
| 175 | p.reviewsCount >= 100 && |
| 176 | (p.email || p.phone) |
| 177 | ) |
| 178 | .map(p => ({ |
| 179 | name: p.name, |
| 180 | rating: p.rating, |
| 181 | reviews: p.reviewsCount, |
| 182 | email: p.email, |
| 183 | phone: p.phone, |
| 184 | website: p.website, |
| 185 | address: p.address |
| 186 | })) |
| 187 | |
| 188 | // Export leads - only qualified results! |
| 189 | console.log(`Found ${qualifiedLeads.length} qualified leads`) |
| 190 | |
| 191 | |
| 192 | **Google Maps - Review sentiment analysis:** |
| 193 | |
| 194 | import { scrapeGoogleMapsReviews } from 'actors' |
| 195 | |
| 196 | const reviews = await scrapeGoogleMapsReviews({ |
| 197 | placeUrl: 'https://maps.google.com/maps?cid=12345', |
| 198 | maxResults: 1000 |
| 199 | }) |
| 200 | |
| 201 | // Filter in code - analyze sentiment by rating |
| 202 | const recentNegative = reviews |
| 203 | .filter(r => { |
| 204 | const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000) |
| 205 | return ( |
| 206 | r.rating <= 2 && |
| 207 | new Date(r.publishedAtDate).getTime() > thirtyDaysAgo && |
| 208 | r.text.length > 50 |
| 209 | ) |
| 210 | }) |
| 211 | |
| 212 | // Identify common complaints |
| 213 | const complaints = recentNegative.map(r => r.text) |
| 214 | |
| 215 | |
| 216 | ### E-commerce & Competitive Intelligence |
| 217 | |
| 218 | **Amazon - Price monitoring:** |
| 219 | |
| 220 | import { scrapeAmazonProduct } from 'actors' |
| 221 | |
| 222 | const product = await scrapeAmazonProduct({ |
| 223 | productUrl: 'https://www.amazon.com/dp/B08L5VT894', |
| 224 | includeReviews: true, |
| 225 | maxReviews: 200 |
| 226 | }) |
| 227 | |
| 228 | // Filter in code - only recent negative reviews |
| 229 | const recentNegative = product.reviews |
| 230 | ?.filter(r => { |
| 231 | const weekAgo = Date.now() - (7 * 24 * 60 * 60 * 1000) |
| 232 | return ( |
| 233 | r.rating <= 2 && |
| 234 | new Date(r.date).getTime() > weekAgo |
| 235 | ) |
| 236 | }) |
| 237 | |
| 238 | console.log(`Price: $${product.price}`) |
| 239 | console.log(`Rating: ${product.rating}/5`) |
| 240 | console.log(`Recent issues: ${recentNegative?.length} complaints`) |
| 241 | |
| 242 | |
| 243 | ### Custom Web Scraping |
| 244 | |
| 245 | **Any Website - Custom extraction:** |
| 246 | |
| 247 | import { scrapeWebsite } from 'actors' |
| 248 | |
| 249 | const products = await scrapeWebsite({ |
| 250 | startUrls: ['https://example.com/products'], |
| 251 | linkSelector: 'a.product-link', |
| 252 | maxPagesPerCrawl: 100, |
| 253 | pageFunction: ` |
| 254 | async function pageFunction(context) { |
| 255 | const { request, $, log } = context |
| 256 | |
| 257 | return { |
| 258 | url: request.url, |
| 259 | title: $('h1.product-title').text(), |
| 260 | price: $('span.price').text(), |
| 261 | inStock: $('.in-stock').length > 0, |
| 262 | description: $('.description').text() |
| 263 | } |
| 264 | } |
| 265 | ` |
| 266 | }) |
| 267 | |
| 268 | // Filter in code - only available products under $100 |
| 269 | const affordable = products.filter(p => |
| 270 | p.inStock && |
| 271 | parseFloat(p.price.replace('$', '')) < 100 |
| 272 | ) |
| 273 | |
| 274 | |
| 275 | ## 🎨 Advanced Patterns |
| 276 | |
| 277 | ### Pattern 1: Multi-Platform Social Listening |
| 278 | |
| 279 | |
| 280 | import { |
| 281 | scrapeInstagramHashtag, |
| 282 | scrapeTikTokHashtag, |
| 283 | searchYouTube |
| 284 | } from 'actors' |
| 285 | |
| 286 | // Run all platforms in parallel |
| 287 | const [instagramPosts, tiktokVideos, youtubeVideos] = await Promise.all([ |
| 288 | scrapeInstagramHashtag({ hashtag: 'ai', maxResults: 100 }), |
| 289 | scrapeTikTokHashtag({ hashtag: 'ai', maxResults: 100 }), |
| 290 | searchYouTube({ query: '#ai', maxResults: 100 }) |
| 291 | ]) |
| 292 | |
| 293 | // Combine and filter - only viral content across all platforms |
| 294 | const allViral = [ |
| 295 | ...instagramPosts.filter(p => p.likesCount > 10000), |
| 296 | ...tiktokVideos.filter(v => v.playCount > 100000), |
| 297 | ...youtubeVideos.filter(v => v.viewsCount > 50000) |
| 298 | ] |
| 299 | |
| 300 | console.log(`Found ${allViral.length} viral posts across 3 platforms`) |
| 301 | |
| 302 | |
| 303 | ### Pattern 2: Lead Enrichment Pipeline |
| 304 | |
| 305 | |
| 306 | import { searchGoogleMaps, scrapeLinkedInProfile } from 'actors' |
| 307 | |
| 308 | // 1. Find businesses on Google Maps |
| 309 | const restaurants = await searchGoogleMaps({ |
| 310 | query: 'restaurants in SF', |
| 311 | maxResults: 100, |
| 312 | scrapeContactInfo: true |
| 313 | }) |
| 314 | |
| 315 | // 2. Filter for qualified leads |
| 316 | const qualified = restaurants.filter(r => |
| 317 | r.rating >= 4.5 && |
| 318 | r.email && |
| 319 | r.reviewsCount >= 50 |
| 320 | ) |
| 321 | |
| 322 | // 3. Enrich with LinkedIn data (if available) |
| 323 | const enriched = await Promise.all( |
| 324 | qualified.map(async (restaurant) => { |
| 325 | // Try to find LinkedIn company page |
| 326 | // ... additional enrichment logic |
| 327 | return restaurant |
| 328 | }) |
| 329 | ) |
| 330 | |
| 331 | |
| 332 | ### Pattern 3: Competitive Analysis Dashboard |
| 333 | |
| 334 | |
| 335 | import { |
| 336 | scrapeInstagramProfile, |
| 337 | scrapeYouTubeChannel, |
| 338 | scrapeTikTokProfile |
| 339 | } from 'actors' |
| 340 | |
| 341 | async function analyzeCompetitor(username: string) { |
| 342 | // Gather data from all platforms |
| 343 | const [instagram, youtube, tiktok] = await Promise.all([ |
| 344 | scrapeInstagramProfile({ username, maxPosts: 30 }), |
| 345 | scrapeYouTubeChannel({ channelUrl: `https://youtube.com/@${username}`, maxVideos: 30 }), |
| 346 | scrapeTikTokProfile({ username, maxVideos: 30 }) |
| 347 | ]) |
| 348 | |
| 349 | // Calculate engagement metrics in code |
| 350 | return { |
| 351 | username, |
| 352 | instagram: { |
| 353 | followers: instagram.followersCount, |
| 354 | avgLikes: average(instagram.latestPosts?.map(p => p.likesCount) || []), |
| 355 | engagementRate: calculateEngagement(instagram) |
| 356 | }, |
| 357 | youtube: { |
| 358 | subscribers: youtube.subscribersCount, |
| 359 | avgViews: average(youtube.videos?.map(v => v.viewsCount) || []) |
| 360 | }, |
| 361 | tiktok: { |
| 362 | followers: tiktok.followersCount, |
| 363 | avgPlays: average(tiktok.videos?.map(v => v.playCount) || []) |
| 364 | } |
| 365 | } |
| 366 | } |
| 367 | |
| 368 | |
| 369 | ## 💰 Token Savings Calculator |
| 370 | |
| 371 | **Example: Instagram profile with 100 posts** |
| 372 | |
| 373 | **MCP Approach:** |
| 374 | |
| 375 | 1. search-actors → 1,000 tokens |
| 376 | 2. call-actor → 1,000 tokens |
| 377 | 3. get-actor-output → 50,000 tokens (100 unfiltered posts) |
| 378 | TOTAL: ~52,000 tokens |
| 379 | |
| 380 | |
| 381 | **File-Based Approach:** |
| 382 | |
| 383 | const profile = await scrapeInstagramProfile({ |
| 384 | username: 'user', |
| 385 | maxPosts: 100 |
| 386 | }) |
| 387 | |
| 388 | // Filter in code - only top 10 posts |
| 389 | const top = profile.latestPosts |
| 390 | ?.sort((a, b) => b.likesCount - a.likesCount) |
| 391 | .slice(0, 10) |
| 392 | |
| 393 | // TOTAL: ~500 tokens (only 10 filtered posts reach model) |
| 394 | |
| 395 | |
| 396 | **Savings: 99% reduction (52,000 → 500 tokens)** |
| 397 | |
| 398 | ## 🔧 Actor Reference |
| 399 | |
| 400 | ### Social Media |
| 401 | |
| 402 | |
| 403 | `scrapeInstagramProfile(input)` - Profile + posts |
| 404 | `scrapeInstagramPosts(input)` - Posts from user |
| 405 | `scrapeInstagramHashtag(input)` - Posts by hashtag |
| 406 | `scrapeInstagramComments(input)` - Comments on post |
| 407 | |
| 408 | |
| 409 | `scrapeLinkedInProfile(input)` - Profile + experience + email |
| 410 | `searchLinkedInJobs(input)` - Job listings |
| 411 | `scrapeLinkedInPosts(input)` - Posts from profile/company |
| 412 | |
| 413 | #### TikTok |
| 414 | `scrapeTikTokProfile(input)` - Profile + videos |
| 415 | `scrapeTikTokHashtag(input)` - Videos by hashtag |
| 416 | `scrapeTikTokComments(input)` - Comments on video |
| 417 | |
| 418 | #### YouTube |
| 419 | `scrapeYouTubeChannel(input)` - Channel + videos |
| 420 | `searchYouTube(input)` - Search videos |
| 421 | `scrapeYouTubeComments(input)` - Comments on video |
| 422 | |
| 423 | |
| 424 | `scrapeFacebookPosts(input)` - Posts from pages |
| 425 | `scrapeFacebookGroups(input)` - Group posts |
| 426 | `scrapeFacebookComments(input)` - Post comments |
| 427 | |
| 428 | ### Business & Lead Generation |
| 429 | |
| 430 | #### Google Maps |
| 431 | `searchGoogleMaps(input)` - Search places (with contact extraction!) |
| 432 | `scrapeGoogleMapsPlace(input)` - Single place details |
| 433 | `scrapeGoogleMapsReviews(input)` - Place reviews |
| 434 | |
| 435 | ### E-commerce |
| 436 | |
| 437 | #### Amazon |
| 438 | `scrapeAmazonProduct(input)` - Product details + reviews |
| 439 | `scrapeAmazonReviews(input)` - Product reviews only |
| 440 | |
| 441 | ### Web Scraping |
| 442 | |
| 443 | #### General Web |
| 444 | `scrapeWebsite(input)` - Custom multi-page crawling |
| 445 | `scrapePage(url, pageFunction)` - Single page extraction |
| 446 | |
| 447 | ## ⚙️ Configuration |
| 448 | |
| 449 | **Environment Variables:** |
| 450 | |
| 451 | # Required - Get from https://console.apify.com/account/integrations |
| 452 | APIFY_TOKEN=apify_api_xxxxx... |
| 453 | |
| 454 | |
| 455 | **Actor Run Options:** |
| 456 | |
| 457 | { |
| 458 | memory: 2048, // MB: 128, 256, 512, 1024, 2048, 4096, 8192 |
| 459 | timeout: 300, // seconds |
| 460 | build: 'latest' // or specific build number |
| 461 | } |
| 462 | |
| 463 | |
| 464 | ## 🎯 When to Use This vs MCP |
| 465 | |
| 466 | **Use File-Based (this skill):** |
| 467 | ✅ Need to filter large datasets (>100 results) |
| 468 | ✅ Want to transform/aggregate data in code |
| 469 | ✅ Multiple sequential operations |
| 470 | ✅ Control flow (loops, conditionals) |
| 471 | ✅ Maximum token efficiency |
| 472 | |
| 473 | **Use MCP:** |
| 474 | ❌ Simple single operations with small results (<10 items) |
| 475 | ❌ One-off exploratory queries |
| 476 | ❌ Don't want to write code |
| 477 | |
| 478 | ## 🔗 Links |
| 479 | |
| 480 | Apify Platform: https://apify.com |
| 481 | Actor Store: https://apify.com/store |
| 482 | API Docs: https://docs.apify.com/api/v2 |
| 483 | |
| 484 | |
| 485 | |
| 486 | **Remember: Filter data in code BEFORE returning to model context. This is where the 99% token savings happen!** |
| 487 | |
| 488 | ## Gotchas |
| 489 | |
| 490 | **Actor selection matters.** Each social platform has specific actors — don't use a generic scraper for Instagram when a dedicated Instagram actor exists. |
| 491 | **Rate limits vary by platform and plan.** Check actor documentation for limits before running large scrapes. |
| 492 | **Scraped data format varies by actor.** Read the actor's output schema before processing results. |
| 493 | |
| 494 | ## Examples |
| 495 | |
| 496 | **Example 1: Scrape Instagram profile** |
| 497 | |
| 498 | User: "get the recent posts from this Instagram account" |
| 499 | → Selects Instagram Profile actor |
| 500 | → Runs with target profile URL |
| 501 | → Returns structured post data (text, engagement, dates) |
| 502 | |
| 503 | |
| 504 | **Example 2: LinkedIn company scrape** |
| 505 | |
| 506 | User: "scrape this company's LinkedIn page" |
| 507 | → Selects LinkedIn Company actor |
| 508 | → Returns company info, employee count, recent posts |
| 509 | |
| 510 | |
| 511 | ## Execution Log |
| 512 | |
| 513 | After completing any workflow, append a single JSONL entry: |
| 514 | |
| 515 | |
| 516 | echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Apify","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl |
| 517 | |
| 518 | |
| 519 | Replace `WORKFLOW_USED` with the workflow executed, `8_WORD_SUMMARY` with a brief input description, and `SECONDS` with approximate wall-clock time. Log `status: "error"` if the workflow failed. |
| 520 |
Discussion
Alternatives
Browse more free Claude skills or everything in Marketing.