Apify skill

Scrapes social platforms, business data, and e-commerce via Apify actors — Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps, Amazon, and web crawls — filtering in code.

by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗

Use now

Files of Apify

danielmiessler/main1 file shown
SKILL.md
Show the full text520 lines

Customization

Before executing, check for user customizations at: ~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Apify/

If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.

🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)

You MUST send this notification BEFORE doing anything else when this skill is invoked.

  1. Send voice notification:

    curl -s -X POST http://localhost:31337/notify \
      -H "Content-Type: application/json" \
      -d '{"message": "Running the WORKFLOWNAME workflow in the Apify skill to ACTION"}' \
      > /dev/null 2>&1 &
    
  2. Output text notification:

    Running the **WorkflowName** workflow in the **Apify** skill to ACTION...
    

This is not optional. Execute this curl command immediately upon skill invocation.

Apify - Social Media & Web Scraping

What It Does

Scrapes social platforms, business data, and e-commerce through Apify actors: Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps business search, Amazon, and general-purpose web crawling. TypeScript wrappers filter and transform the data in code before any of it reaches the model, so a 100-post scrape costs roughly what 10 posts would. Runs platforms in parallel for social-listening dashboards and chains Google Maps into LinkedIn for lead enrichment.

The Problem

Scraping through a raw MCP dumps every unfiltered result straight into model context — a single Instagram profile with 100 posts burns ~52,000 tokens, most of it noise you'll throw away. You usually want the top 10 posts, the negative reviews from the last week, the qualified leads with an email. Doing that filtering after the data hits the model is too late; the tokens are already spent. Filtering in code first cuts that 52,000 down to ~500.

How It Works

This skill is a file-based MCP — a code-first API wrapper that replaces token-heavy MCP protocol calls. You call an actor wrapper, filter and sort the result in TypeScript, and only the filtered slice reaches model context. That code-before-context step is where the 95-99% token savings come from.

Workflow Routing

Workflow Trigger File
Update update Apify skill, refresh actors, actor calls failing unexpectedly, monthly capability check Workflows/Update.md
(inline) all scrape/lead/crawl requests — scrape Instagram/LinkedIn/TikTok/YouTube/Facebook, Google Maps leads, Amazon reviews, web crawl Actor wrappers under actors/ (see Actor Reference below)

📊 Available Actors

Social Media (5 platforms)
  • Instagram (145k users, 4.60★) - Profiles, posts, hashtags, comments
  • LinkedIn (26k users, 4.10★) - Profiles, jobs, posts
  • TikTok (90k users, 4.61★) - Profiles, videos, hashtags, comments
  • YouTube (40k users, 4.40★) - Channels, videos, comments, search
  • Facebook (35k users, 4.56★) - Posts, groups, comments
Business & Lead Generation
  • Google Maps (198k users, 4.76★) - HIGHEST VALUE!
    • Search businesses, extract contacts, reviews, images
    • Perfect for lead generation
E-commerce
  • Amazon (8k users, 4.97★) - Products, reviews, pricing
Web Scraping
  • Web Scraper (94k users, 4.39★) - General-purpose, works with ANY website

🚀 Quick Start

Basic Usage Pattern
import { scrapeInstagramProfile, searchGoogleMaps } from 'actors'

// 1. Call the actor wrapper
const profile = await scrapeInstagramProfile({
  username: 'target_username',
  maxPosts: 50
})

// 2. Filter in code - BEFORE data reaches model!
const viral = profile.latestPosts?.filter(p => p.likesCount > 10000)

// 3. Only filtered results reach model context
console.log(viral) // ~10 posts instead of 50

📚 Examples by Use Case

Social Media Monitoring

Instagram - Track engagement:

import { scrapeInstagramProfile, scrapeInstagramPosts } from 'actors'

// Get profile with recent posts
const profile = await scrapeInstagramProfile({
  username: 'competitor',
  maxPosts: 100
})

// Filter in code - only high-performing posts from last 30 days
const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
const topRecent = profile.latestPosts
  ?.filter(p =>
    new Date(p.timestamp).getTime() > thirtyDaysAgo &&
    p.likesCount > 5000
  )
  .sort((a, b) => b.likesCount - a.likesCount)
  .slice(0, 10)

// Only 10 posts reach model instead of 100!

LinkedIn - Job search:

import { searchLinkedInJobs } from 'actors'

const jobs = await searchLinkedInJobs({
  keywords: 'AI engineer',
  location: 'San Francisco',
  remote: true,
  maxResults: 200
})

// Filter in code - only senior roles at well-funded startups
const topJobs = jobs.filter(j =>
  j.seniority?.includes('Senior') &&
  parseInt(j.applicants || '0') > 50
)

TikTok - Trend analysis:

import { scrapeTikTokHashtag } from 'actors'

const videos = await scrapeTikTokHashtag({
  hashtag: 'ai',
  maxResults: 500
})

// Filter in code - only viral content
const viral = videos
  .filter(v => v.playCount > 1000000)
  .sort((a, b) => b.playCount - a.playCount)
  .slice(0, 20)
Lead Generation (Business Intelligence)

Google Maps - Local business leads:

import { searchGoogleMaps } from 'actors'

// Search with contact info extraction
const places = await searchGoogleMaps({
  query: 'restaurants in Austin',
  maxResults: 500,
  includeReviews: true,
  maxReviewsPerPlace: 20,
  scrapeContactInfo: true // Extracts emails from websites!
})

// Filter in code - only highly-rated with email/phone
const qualifiedLeads = places
  .filter(p =>
    p.rating >= 4.5 &&
    p.reviewsCount >= 100 &&
    (p.email || p.phone)
  )
  .map(p => ({
    name: p.name,
    rating: p.rating,
    reviews: p.reviewsCount,
    email: p.email,
    phone: p.phone,
    website: p.website,
    address: p.address
  }))

// Export leads - only qualified results!
console.log(`Found ${qualifiedLeads.length} qualified leads`)

Google Maps - Review sentiment analysis:

import { scrapeGoogleMapsReviews } from 'actors'

const reviews = await scrapeGoogleMapsReviews({
  placeUrl: 'https://maps.google.com/maps?cid=12345',
  maxResults: 1000
})

// Filter in code - analyze sentiment by rating
const recentNegative = reviews
  .filter(r => {
    const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
    return (
      r.rating <= 2 &&
      new Date(r.publishedAtDate).getTime() > thirtyDaysAgo &&
      r.text.length > 50
    )
  })

// Identify common complaints
const complaints = recentNegative.map(r => r.text)
E-commerce & Competitive Intelligence

Amazon - Price monitoring:

import { scrapeAmazonProduct } from 'actors'

const product = await scrapeAmazonProduct({
  productUrl: 'https://www.amazon.com/dp/B08L5VT894',
  includeReviews: true,
  maxReviews: 200
})

// Filter in code - only recent negative reviews
const recentNegative = product.reviews
  ?.filter(r => {
    const weekAgo = Date.now() - (7 * 24 * 60 * 60 * 1000)
    return (
      r.rating <= 2 &&
      new Date(r.date).getTime() > weekAgo
    )
  })

console.log(`Price: $${product.price}`)
console.log(`Rating: ${product.rating}/5`)
console.log(`Recent issues: ${recentNegative?.length} complaints`)
Custom Web Scraping

Any Website - Custom extraction:

import { scrapeWebsite } from 'actors'

const products = await scrapeWebsite({
  startUrls: ['https://example.com/products'],
  linkSelector: 'a.product-link',
  maxPagesPerCrawl: 100,
  pageFunction: `
    async function pageFunction(context) {
      const { request, $, log } = context

      return {
        url: request.url,
        title: $('h1.product-title').text(),
        price: $('span.price').text(),
        inStock: $('.in-stock').length > 0,
        description: $('.description').text()
      }
    }
  `
})

// Filter in code - only available products under $100
const affordable = products.filter(p =>
  p.inStock &&
  parseFloat(p.price.replace('$', '')) < 100
)

🎨 Advanced Patterns

Pattern 1: Multi-Platform Social Listening
import {
  scrapeInstagramHashtag,
  scrapeTikTokHashtag,
  searchYouTube
} from 'actors'

// Run all platforms in parallel
const [instagramPosts, tiktokVideos, youtubeVideos] = await Promise.all([
  scrapeInstagramHashtag({ hashtag: 'ai', maxResults: 100 }),
  scrapeTikTokHashtag({ hashtag: 'ai', maxResults: 100 }),
  searchYouTube({ query: '#ai', maxResults: 100 })
])

// Combine and filter - only viral content across all platforms
const allViral = [
  ...instagramPosts.filter(p => p.likesCount > 10000),
  ...tiktokVideos.filter(v => v.playCount > 100000),
  ...youtubeVideos.filter(v => v.viewsCount > 50000)
]

console.log(`Found ${allViral.length} viral posts across 3 platforms`)
Pattern 2: Lead Enrichment Pipeline
import { searchGoogleMaps, scrapeLinkedInProfile } from 'actors'

// 1. Find businesses on Google Maps
const restaurants = await searchGoogleMaps({
  query: 'restaurants in SF',
  maxResults: 100,
  scrapeContactInfo: true
})

// 2. Filter for qualified leads
const qualified = restaurants.filter(r =>
  r.rating >= 4.5 &&
  r.email &&
  r.reviewsCount >= 50
)

// 3. Enrich with LinkedIn data (if available)
const enriched = await Promise.all(
  qualified.map(async (restaurant) => {
    // Try to find LinkedIn company page
    // ... additional enrichment logic
    return restaurant
  })
)
Pattern 3: Competitive Analysis Dashboard
import {
  scrapeInstagramProfile,
  scrapeYouTubeChannel,
  scrapeTikTokProfile
} from 'actors'

async function analyzeCompetitor(username: string) {
  // Gather data from all platforms
  const [instagram, youtube, tiktok] = await Promise.all([
    scrapeInstagramProfile({ username, maxPosts: 30 }),
    scrapeYouTubeChannel({ channelUrl: `https://youtube.com/@${username}`, maxVideos: 30 }),
    scrapeTikTokProfile({ username, maxVideos: 30 })
  ])

  // Calculate engagement metrics in code
  return {
    username,
    instagram: {
      followers: instagram.followersCount,
      avgLikes: average(instagram.latestPosts?.map(p => p.likesCount) || []),
      engagementRate: calculateEngagement(instagram)
    },
    youtube: {
      subscribers: youtube.subscribersCount,
      avgViews: average(youtube.videos?.map(v => v.viewsCount) || [])
    },
    tiktok: {
      followers: tiktok.followersCount,
      avgPlays: average(tiktok.videos?.map(v => v.playCount) || [])
    }
  }
}

💰 Token Savings Calculator

Example: Instagram profile with 100 posts

MCP Approach:

1. search-actors → 1,000 tokens
2. call-actor → 1,000 tokens
3. get-actor-output → 50,000 tokens (100 unfiltered posts)
TOTAL: ~52,000 tokens

File-Based Approach:

const profile = await scrapeInstagramProfile({
  username: 'user',
  maxPosts: 100
})

// Filter in code - only top 10 posts
const top = profile.latestPosts
  ?.sort((a, b) => b.likesCount - a.likesCount)
  .slice(0, 10)

// TOTAL: ~500 tokens (only 10 filtered posts reach model)

Savings: 99% reduction (52,000 → 500 tokens)

🔧 Actor Reference

Social Media
Instagram
  • scrapeInstagramProfile(input) - Profile + posts
  • scrapeInstagramPosts(input) - Posts from user
  • scrapeInstagramHashtag(input) - Posts by hashtag
  • scrapeInstagramComments(input) - Comments on post
LinkedIn
  • scrapeLinkedInProfile(input) - Profile + experience + email
  • searchLinkedInJobs(input) - Job listings
  • scrapeLinkedInPosts(input) - Posts from profile/company
TikTok
  • scrapeTikTokProfile(input) - Profile + videos
  • scrapeTikTokHashtag(input) - Videos by hashtag
  • scrapeTikTokComments(input) - Comments on video
YouTube
  • scrapeYouTubeChannel(input) - Channel + videos
  • searchYouTube(input) - Search videos
  • scrapeYouTubeComments(input) - Comments on video
Facebook
  • scrapeFacebookPosts(input) - Posts from pages
  • scrapeFacebookGroups(input) - Group posts
  • scrapeFacebookComments(input) - Post comments
Business & Lead Generation
Google Maps
  • searchGoogleMaps(input) - Search places (with contact extraction!)
  • scrapeGoogleMapsPlace(input) - Single place details
  • scrapeGoogleMapsReviews(input) - Place reviews
E-commerce
Amazon
  • scrapeAmazonProduct(input) - Product details + reviews
  • scrapeAmazonReviews(input) - Product reviews only
Web Scraping
General Web
  • scrapeWebsite(input) - Custom multi-page crawling
  • scrapePage(url, pageFunction) - Single page extraction

⚙️ Configuration

Environment Variables:

# Required - Get from https://console.apify.com/account/integrations
APIFY_TOKEN=apify_api_xxxxx...

Actor Run Options:

{
  memory: 2048,    // MB: 128, 256, 512, 1024, 2048, 4096, 8192
  timeout: 300,    // seconds
  build: 'latest'  // or specific build number
}

🎯 When to Use This vs MCP

Use File-Based (this skill):

  • ✅ Need to filter large datasets (>100 results)
  • ✅ Want to transform/aggregate data in code
  • ✅ Multiple sequential operations
  • ✅ Control flow (loops, conditionals)
  • ✅ Maximum token efficiency

Use MCP:

  • ❌ Simple single operations with small results (<10 items)
  • ❌ One-off exploratory queries
  • ❌ Don't want to write code

Remember: Filter data in code BEFORE returning to model context. This is where the 99% token savings happen!

Gotchas

  • Actor selection matters. Each social platform has specific actors — don't use a generic scraper for Instagram when a dedicated Instagram actor exists.
  • Rate limits vary by platform and plan. Check actor documentation for limits before running large scrapes.
  • Scraped data format varies by actor. Read the actor's output schema before processing results.

Examples

Example 1: Scrape Instagram profile

User: "get the recent posts from this Instagram account"
→ Selects Instagram Profile actor
→ Runs with target profile URL
→ Returns structured post data (text, engagement, dates)

Example 2: LinkedIn company scrape

User: "scrape this company's LinkedIn page"
→ Selects LinkedIn Company actor
→ Returns company info, employee count, recent posts

Execution Log

After completing any workflow, append a single JSONL entry:

echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Apify","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl

Replace WORKFLOW_USED with the workflow executed, 8_WORD_SUMMARY with a brief input description, and SECONDS with approximate wall-clock time. Log status: "error" if the workflow failed.

1---
2name: Apify
3version: 1.1.22
4description: "Scrapes social platforms, business data, and e-commerce via Apify actors — Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps, Amazon, and web crawls — filtering in code. USE WHEN scrape Instagram, scrape LinkedIn, scrape TikTok, scrape YouTube, scrape Facebook, Google Maps leads, Amazon reviews, business intelligence, multi-platform social listening, competitive analysis, lead generation, social monitoring, Apify actors, web crawl, extract contacts. NOT FOR X/Twitter account operations like posting, threads, or bookmarks (those need a dedicated X API client), 4-tier progressive scraping with proxy escalation (use BrightData), or real-Chrome bot bypass and computer use (use Interceptor)."
5---
6 
7## Customization
8 
9**Before executing, check for user customizations at:**
10`~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Apify/`
11 
12If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
13 
14 
15## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
16 
17**You MUST send this notification BEFORE doing anything else when this skill is invoked.**
18 
191. **Send voice notification**:
20 ```bash
21 curl -s -X POST http://localhost:31337/notify \
22 -H "Content-Type: application/json" \
23 -d '{"message": "Running the WORKFLOWNAME workflow in the Apify skill to ACTION"}' \
24 > /dev/null 2>&1 &
25 ```
26 
272. **Output text notification**:
28 ```
29 Running the **WorkflowName** workflow in the **Apify** skill to ACTION...
30 ```
31 
32**This is not optional. Execute this curl command immediately upon skill invocation.**
33 
34# Apify - Social Media & Web Scraping
35 
36## What It Does
37 
38Scrapes social platforms, business data, and e-commerce through Apify actors: Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps business search, Amazon, and general-purpose web crawling. TypeScript wrappers filter and transform the data in code before any of it reaches the model, so a 100-post scrape costs roughly what 10 posts would. Runs platforms in parallel for social-listening dashboards and chains Google Maps into LinkedIn for lead enrichment.
39 
40## The Problem
41 
42Scraping through a raw MCP dumps every unfiltered result straight into model context — a single Instagram profile with 100 posts burns ~52,000 tokens, most of it noise you'll throw away. You usually want the top 10 posts, the negative reviews from the last week, the qualified leads with an email. Doing that filtering after the data hits the model is too late; the tokens are already spent. Filtering in code first cuts that 52,000 down to ~500.
43 
44## How It Works
45 
46This skill is a **file-based MCP** — a code-first API wrapper that replaces token-heavy MCP protocol calls. You call an actor wrapper, filter and sort the result in TypeScript, and only the filtered slice reaches model context. That code-before-context step is where the 95-99% token savings come from.
47 
48## Workflow Routing
49 
50| Workflow | Trigger | File |
51|----------|---------|------|
52| Update | update Apify skill, refresh actors, actor calls failing unexpectedly, monthly capability check | `Workflows/Update.md` |
53| (inline) | all scrape/lead/crawl requests — scrape Instagram/LinkedIn/TikTok/YouTube/Facebook, Google Maps leads, Amazon reviews, web crawl | Actor wrappers under `actors/` (see Actor Reference below) |
54 
55## 📊 Available Actors
56 
57### Social Media (5 platforms)
58- **Instagram** (145k users, 4.60★) - Profiles, posts, hashtags, comments
59- **LinkedIn** (26k users, 4.10★) - Profiles, jobs, posts
60- **TikTok** (90k users, 4.61★) - Profiles, videos, hashtags, comments
61- **YouTube** (40k users, 4.40★) - Channels, videos, comments, search
62- **Facebook** (35k users, 4.56★) - Posts, groups, comments
63 
64### Business & Lead Generation
65- **Google Maps** (198k users, 4.76★) - **HIGHEST VALUE!**
66 - Search businesses, extract contacts, reviews, images
67 - Perfect for lead generation
68 
69### E-commerce
70- **Amazon** (8k users, 4.97★) - Products, reviews, pricing
71 
72### Web Scraping
73- **Web Scraper** (94k users, 4.39★) - General-purpose, works with ANY website
74 
75## 🚀 Quick Start
76 
77### Basic Usage Pattern
78 
79```typescript
80import { scrapeInstagramProfile, searchGoogleMaps } from 'actors'
81 
82// 1. Call the actor wrapper
83const profile = await scrapeInstagramProfile({
84 username: 'target_username',
85 maxPosts: 50
86})
87 
88// 2. Filter in code - BEFORE data reaches model!
89const viral = profile.latestPosts?.filter(p => p.likesCount > 10000)
90 
91// 3. Only filtered results reach model context
92console.log(viral) // ~10 posts instead of 50
93```
94 
95## 📚 Examples by Use Case
96 
97### Social Media Monitoring
98 
99**Instagram - Track engagement:**
100```typescript
101import { scrapeInstagramProfile, scrapeInstagramPosts } from 'actors'
102 
103// Get profile with recent posts
104const profile = await scrapeInstagramProfile({
105 username: 'competitor',
106 maxPosts: 100
107})
108 
109// Filter in code - only high-performing posts from last 30 days
110const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
111const topRecent = profile.latestPosts
112 ?.filter(p =>
113 new Date(p.timestamp).getTime() > thirtyDaysAgo &&
114 p.likesCount > 5000
115 )
116 .sort((a, b) => b.likesCount - a.likesCount)
117 .slice(0, 10)
118 
119// Only 10 posts reach model instead of 100!
120```
121 
122**LinkedIn - Job search:**
123```typescript
124import { searchLinkedInJobs } from 'actors'
125 
126const jobs = await searchLinkedInJobs({
127 keywords: 'AI engineer',
128 location: 'San Francisco',
129 remote: true,
130 maxResults: 200
131})
132 
133// Filter in code - only senior roles at well-funded startups
134const topJobs = jobs.filter(j =>
135 j.seniority?.includes('Senior') &&
136 parseInt(j.applicants || '0') > 50
137)
138```
139 
140**TikTok - Trend analysis:**
141```typescript
142import { scrapeTikTokHashtag } from 'actors'
143 
144const videos = await scrapeTikTokHashtag({
145 hashtag: 'ai',
146 maxResults: 500
147})
148 
149// Filter in code - only viral content
150const viral = videos
151 .filter(v => v.playCount > 1000000)
152 .sort((a, b) => b.playCount - a.playCount)
153 .slice(0, 20)
154```
155 
156### Lead Generation (Business Intelligence)
157 
158**Google Maps - Local business leads:**
159```typescript
160import { searchGoogleMaps } from 'actors'
161 
162// Search with contact info extraction
163const places = await searchGoogleMaps({
164 query: 'restaurants in Austin',
165 maxResults: 500,
166 includeReviews: true,
167 maxReviewsPerPlace: 20,
168 scrapeContactInfo: true // Extracts emails from websites!
169})
170 
171// Filter in code - only highly-rated with email/phone
172const qualifiedLeads = places
173 .filter(p =>
174 p.rating >= 4.5 &&
175 p.reviewsCount >= 100 &&
176 (p.email || p.phone)
177 )
178 .map(p => ({
179 name: p.name,
180 rating: p.rating,
181 reviews: p.reviewsCount,
182 email: p.email,
183 phone: p.phone,
184 website: p.website,
185 address: p.address
186 }))
187 
188// Export leads - only qualified results!
189console.log(`Found ${qualifiedLeads.length} qualified leads`)
190```
191 
192**Google Maps - Review sentiment analysis:**
193```typescript
194import { scrapeGoogleMapsReviews } from 'actors'
195 
196const reviews = await scrapeGoogleMapsReviews({
197 placeUrl: 'https://maps.google.com/maps?cid=12345',
198 maxResults: 1000
199})
200 
201// Filter in code - analyze sentiment by rating
202const recentNegative = reviews
203 .filter(r => {
204 const thirtyDaysAgo = Date.now() - (30 * 24 * 60 * 60 * 1000)
205 return (
206 r.rating <= 2 &&
207 new Date(r.publishedAtDate).getTime() > thirtyDaysAgo &&
208 r.text.length > 50
209 )
210 })
211 
212// Identify common complaints
213const complaints = recentNegative.map(r => r.text)
214```
215 
216### E-commerce & Competitive Intelligence
217 
218**Amazon - Price monitoring:**
219```typescript
220import { scrapeAmazonProduct } from 'actors'
221 
222const product = await scrapeAmazonProduct({
223 productUrl: 'https://www.amazon.com/dp/B08L5VT894',
224 includeReviews: true,
225 maxReviews: 200
226})
227 
228// Filter in code - only recent negative reviews
229const recentNegative = product.reviews
230 ?.filter(r => {
231 const weekAgo = Date.now() - (7 * 24 * 60 * 60 * 1000)
232 return (
233 r.rating <= 2 &&
234 new Date(r.date).getTime() > weekAgo
235 )
236 })
237 
238console.log(`Price: $${product.price}`)
239console.log(`Rating: ${product.rating}/5`)
240console.log(`Recent issues: ${recentNegative?.length} complaints`)
241```
242 
243### Custom Web Scraping
244 
245**Any Website - Custom extraction:**
246```typescript
247import { scrapeWebsite } from 'actors'
248 
249const products = await scrapeWebsite({
250 startUrls: ['https://example.com/products'],
251 linkSelector: 'a.product-link',
252 maxPagesPerCrawl: 100,
253 pageFunction: `
254 async function pageFunction(context) {
255 const { request, $, log } = context
256 
257 return {
258 url: request.url,
259 title: $('h1.product-title').text(),
260 price: $('span.price').text(),
261 inStock: $('.in-stock').length > 0,
262 description: $('.description').text()
263 }
264 }
265 `
266})
267 
268// Filter in code - only available products under $100
269const affordable = products.filter(p =>
270 p.inStock &&
271 parseFloat(p.price.replace('$', '')) < 100
272)
273```
274 
275## 🎨 Advanced Patterns
276 
277### Pattern 1: Multi-Platform Social Listening
278 
279```typescript
280import {
281 scrapeInstagramHashtag,
282 scrapeTikTokHashtag,
283 searchYouTube
284} from 'actors'
285 
286// Run all platforms in parallel
287const [instagramPosts, tiktokVideos, youtubeVideos] = await Promise.all([
288 scrapeInstagramHashtag({ hashtag: 'ai', maxResults: 100 }),
289 scrapeTikTokHashtag({ hashtag: 'ai', maxResults: 100 }),
290 searchYouTube({ query: '#ai', maxResults: 100 })
291])
292 
293// Combine and filter - only viral content across all platforms
294const allViral = [
295 ...instagramPosts.filter(p => p.likesCount > 10000),
296 ...tiktokVideos.filter(v => v.playCount > 100000),
297 ...youtubeVideos.filter(v => v.viewsCount > 50000)
298]
299 
300console.log(`Found ${allViral.length} viral posts across 3 platforms`)
301```
302 
303### Pattern 2: Lead Enrichment Pipeline
304 
305```typescript
306import { searchGoogleMaps, scrapeLinkedInProfile } from 'actors'
307 
308// 1. Find businesses on Google Maps
309const restaurants = await searchGoogleMaps({
310 query: 'restaurants in SF',
311 maxResults: 100,
312 scrapeContactInfo: true
313})
314 
315// 2. Filter for qualified leads
316const qualified = restaurants.filter(r =>
317 r.rating >= 4.5 &&
318 r.email &&
319 r.reviewsCount >= 50
320)
321 
322// 3. Enrich with LinkedIn data (if available)
323const enriched = await Promise.all(
324 qualified.map(async (restaurant) => {
325 // Try to find LinkedIn company page
326 // ... additional enrichment logic
327 return restaurant
328 })
329)
330```
331 
332### Pattern 3: Competitive Analysis Dashboard
333 
334```typescript
335import {
336 scrapeInstagramProfile,
337 scrapeYouTubeChannel,
338 scrapeTikTokProfile
339} from 'actors'
340 
341async function analyzeCompetitor(username: string) {
342 // Gather data from all platforms
343 const [instagram, youtube, tiktok] = await Promise.all([
344 scrapeInstagramProfile({ username, maxPosts: 30 }),
345 scrapeYouTubeChannel({ channelUrl: `https://youtube.com/@${username}`, maxVideos: 30 }),
346 scrapeTikTokProfile({ username, maxVideos: 30 })
347 ])
348 
349 // Calculate engagement metrics in code
350 return {
351 username,
352 instagram: {
353 followers: instagram.followersCount,
354 avgLikes: average(instagram.latestPosts?.map(p => p.likesCount) || []),
355 engagementRate: calculateEngagement(instagram)
356 },
357 youtube: {
358 subscribers: youtube.subscribersCount,
359 avgViews: average(youtube.videos?.map(v => v.viewsCount) || [])
360 },
361 tiktok: {
362 followers: tiktok.followersCount,
363 avgPlays: average(tiktok.videos?.map(v => v.playCount) || [])
364 }
365 }
366}
367```
368 
369## 💰 Token Savings Calculator
370 
371**Example: Instagram profile with 100 posts**
372 
373**MCP Approach:**
374```
3751. search-actors → 1,000 tokens
3762. call-actor → 1,000 tokens
3773. get-actor-output → 50,000 tokens (100 unfiltered posts)
378TOTAL: ~52,000 tokens
379```
380 
381**File-Based Approach:**
382```typescript
383const profile = await scrapeInstagramProfile({
384 username: 'user',
385 maxPosts: 100
386})
387 
388// Filter in code - only top 10 posts
389const top = profile.latestPosts
390 ?.sort((a, b) => b.likesCount - a.likesCount)
391 .slice(0, 10)
392 
393// TOTAL: ~500 tokens (only 10 filtered posts reach model)
394```
395 
396**Savings: 99% reduction (52,000 → 500 tokens)**
397 
398## 🔧 Actor Reference
399 
400### Social Media
401 
402#### Instagram
403- `scrapeInstagramProfile(input)` - Profile + posts
404- `scrapeInstagramPosts(input)` - Posts from user
405- `scrapeInstagramHashtag(input)` - Posts by hashtag
406- `scrapeInstagramComments(input)` - Comments on post
407 
408#### LinkedIn
409- `scrapeLinkedInProfile(input)` - Profile + experience + email
410- `searchLinkedInJobs(input)` - Job listings
411- `scrapeLinkedInPosts(input)` - Posts from profile/company
412 
413#### TikTok
414- `scrapeTikTokProfile(input)` - Profile + videos
415- `scrapeTikTokHashtag(input)` - Videos by hashtag
416- `scrapeTikTokComments(input)` - Comments on video
417 
418#### YouTube
419- `scrapeYouTubeChannel(input)` - Channel + videos
420- `searchYouTube(input)` - Search videos
421- `scrapeYouTubeComments(input)` - Comments on video
422 
423#### Facebook
424- `scrapeFacebookPosts(input)` - Posts from pages
425- `scrapeFacebookGroups(input)` - Group posts
426- `scrapeFacebookComments(input)` - Post comments
427 
428### Business & Lead Generation
429 
430#### Google Maps
431- `searchGoogleMaps(input)` - Search places (with contact extraction!)
432- `scrapeGoogleMapsPlace(input)` - Single place details
433- `scrapeGoogleMapsReviews(input)` - Place reviews
434 
435### E-commerce
436 
437#### Amazon
438- `scrapeAmazonProduct(input)` - Product details + reviews
439- `scrapeAmazonReviews(input)` - Product reviews only
440 
441### Web Scraping
442 
443#### General Web
444- `scrapeWebsite(input)` - Custom multi-page crawling
445- `scrapePage(url, pageFunction)` - Single page extraction
446 
447## ⚙️ Configuration
448 
449**Environment Variables:**
450```bash
451# Required - Get from https://console.apify.com/account/integrations
452APIFY_TOKEN=apify_api_xxxxx...
453```
454 
455**Actor Run Options:**
456```typescript
457{
458 memory: 2048, // MB: 128, 256, 512, 1024, 2048, 4096, 8192
459 timeout: 300, // seconds
460 build: 'latest' // or specific build number
461}
462```
463 
464## 🎯 When to Use This vs MCP
465 
466**Use File-Based (this skill):**
467- ✅ Need to filter large datasets (>100 results)
468- ✅ Want to transform/aggregate data in code
469- ✅ Multiple sequential operations
470- ✅ Control flow (loops, conditionals)
471- ✅ Maximum token efficiency
472 
473**Use MCP:**
474- ❌ Simple single operations with small results (<10 items)
475- ❌ One-off exploratory queries
476- ❌ Don't want to write code
477 
478## 🔗 Links
479 
480- Apify Platform: https://apify.com
481- Actor Store: https://apify.com/store
482- API Docs: https://docs.apify.com/api/v2
483 
484---
485 
486**Remember: Filter data in code BEFORE returning to model context. This is where the 99% token savings happen!**
487 
488## Gotchas
489 
490- **Actor selection matters.** Each social platform has specific actors — don't use a generic scraper for Instagram when a dedicated Instagram actor exists.
491- **Rate limits vary by platform and plan.** Check actor documentation for limits before running large scrapes.
492- **Scraped data format varies by actor.** Read the actor's output schema before processing results.
493 
494## Examples
495 
496**Example 1: Scrape Instagram profile**
497```
498User: "get the recent posts from this Instagram account"
499→ Selects Instagram Profile actor
500→ Runs with target profile URL
501→ Returns structured post data (text, engagement, dates)
502```
503 
504**Example 2: LinkedIn company scrape**
505```
506User: "scrape this company's LinkedIn page"
507→ Selects LinkedIn Company actor
508→ Returns company info, employee count, recent posts
509```
510 
511## Execution Log
512 
513After completing any workflow, append a single JSONL entry:
514 
515```bash
516echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Apify","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
517```
518 
519Replace `WORKFLOW_USED` with the workflow executed, `8_WORD_SUMMARY` with a brief input description, and `SECONDS` with approximate wall-clock time. Log `status: "error"` if the workflow failed.
520 

Discussion

Alternatives

Twitter/X SkillSearch and retrieve content from Twitter/X. Get user info, tweets, replies, followers, communities, spaces, and trends via twitterapi.io. Use when user mentions Twitter, X, or tweets.Marketing · Apache-2.0🎬 YouTube Script: Writing retention-optimized video script...Use this skill when the user says 'YouTube script', 'video script', 'write script for YouTube', 'YouTube video outline', or is creating scripted content for a YouTube video with hooks, chapters, and CTAs. Do NOT use for TikTok/Reels short-form scripts or webinar presentations.Business & ops · MITInstagram Audience InsightsRead your Instagram niche and profile from real data via Apify, no login. Scan a hashtag for the posts traveling now (likes, comments, owner) to see the format and hook that works. Pull profile stats for any handle, yours or a competitor's: followers, posts, bio, category. Instagram hides who liked or commented on other accounts, so this is discovery plus profiles, not engagers. Triggers on "what works in my niche", "scan the hashtag", "competitor stats". Not for writing captions (use ig-caption-writer).Marketing · MITX (Twitter) Marketing SkillsPlan, draft, audit, and publish posts and threads for X (Twitter). Use when the user wants to write a single tweet or an auto-numbered thread, build a long-form tweetstorm, remove AI tells from a draft, reverse-engineer the hook from a viral tweet, draft a reply or quote tweet, or plan a week of X content. Tweets and threads publish via the Publora API, which auto-splits long content into a numbered thread. User provides notes or a tweet URL, the skill drafts, the user approves, then it publishes.Marketing · MIT