{"tier":2,"imagePrompt":false,"slug":"pdf-processor","kind":"skill","name":"PDF Processor - Extract Data from PDFs","tagline":"Process PDFs - extract text, tables, and structured data from documents","category":"Data & AI","tags":["Data & AI"],"body":"---\nname: pdf-processor\ndescription: Process PDFs - extract text, tables, and structured data from documents\nsource: orthogonal\n---\n\n\n# PDF Processor - Extract Data from PDFs\n\n## Setup\n\nRead your credentials from ~/.gooseworks/credentials.json:\n```bash\nexport GOOSEWORKS_API_KEY=$(python3 -c \"import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])\")\nexport GOOSEWORKS_API_BASE=$(python3 -c \"import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))\")\n```\n\nIf ~/.gooseworks/credentials.json does not exist, tell the user to run: `npx gooseworks login`\n\nAll endpoints use Bearer auth: `-H \"Authorization: Bearer $GOOSEWORKS_API_KEY\"`\n\n\nExtract text, tables, and structured data from PDF documents.\n\n## Workflow\n\n### Step 1: Fetch PDF Content\nUse Linkup to fetch PDF URLs:\n\n```bash\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"linkup\",\"path\":\"/fetch\",\"body\":{\"url\":\"https://example.com/document.pdf\"}}'\n```\n\n### Step 2: Extract with AI\nUse ScrapeGraph to extract specific content:\n\n```bash\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"scrapegraph\",\"path\":\"/v1/smartscraper\"}'\n  \"website_url\": \"https://example.com/report.pdf\",\n  \"user_prompt\": \"Extract all financial figures, tables, and key metrics from this document\"\n}'\n```\n\n### Step 3: Extract Tables\nGet structured table data:\n\n```bash\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"riveter\",\"path\":\"/v1/run\"}'\n  \"input\": {\n    \"urls\": [\"https://example.com/report.pdf\"]\n  },\n  \"output\": {\n    \"tables\": {\"prompt\": \"Extract all tables with titles, headers, and rows\", \"contexts\": [\"urls\"]}\n  }\n}'\n```\n\n### Step 4: Convert to Markdown\nGet readable markdown output:\n\n```bash\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"scrapegraph\",\"path\":\"/v1/markdownify\",\"body\":{\"website_url\":\"https://example.com/document.pdf\"}}'\n```\n\n## Example Usage\n\n```bash\n# Extract data from financial report\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"scrapegraph\",\"path\":\"/v1/smartscraper\"}'\n  \"website_url\": \"https://example.com/annual-report.pdf\",\n  \"user_prompt\": \"Extract revenue, profit, and key business metrics with their values\"\n}'\n\n# Extract invoice data\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"riveter\",\"path\":\"/v1/run\"}'\n  \"input\": {\"urls\": [\"https://example.com/invoice.pdf\"]},\n  \"output\": {\n    \"vendor\": {\"prompt\": \"Vendor name\", \"contexts\": [\"urls\"]},\n    \"amount\": {\"prompt\": \"Total amount\", \"contexts\": [\"urls\"]},\n    \"date\": {\"prompt\": \"Invoice date\", \"contexts\": [\"urls\"]}\n  }\n}'\n```\n\n## Tips\n\n- Specify exact data you need for better extraction\n- Use schemas for consistent structured output\n- Handle multi-page documents in chunks\n- Verify extracted numbers against source\n\n## Discover More\n\nList all endpoints, or add a path for parameter details:\n\n```bash\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"prompt\":\"linkup API endpoints\"}' api show riveter\ncurl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"prompt\":\"scrapegraph API endpoints\"}'\n\nExample: `curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \\\n  -H \"Authorization: Bearer $GOOSEWORKS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"api\":\"olostep\",\"path\":\"/v1/scrapes`\"}' for endpoint parameters.\n","bytes":4197,"lineCount":128,"installName":"pdf-processor","problem":"","solves":"","notFor":"","author":"gooseworks-ai","authorUrl":"https://github.com/gooseworks-ai/goose-skills","license":"MIT","repo":"gooseworks-ai/goose-skills","sourceUrl":"https://github.com/gooseworks-ai/goose-skills/blob/main/skills/research-tools/capabilities/pdf-processor/SKILL.md","vendoredVia":null,"basis":"repo","stars":1220,"hnPoints":null,"quetNgay":null,"forks":208,"contributors":4,"repoPushedAt":"2026-09-21T13:41:59Z","archived":false,"commits":3,"editedByCount":2,"lastTouchDays":107,"compat":[{"platform":"claude_code","label":"Claude Code","status":"partial","note":"Has SKILL.md but declares no allowed-tools — Claude Code will ask for permission each time","evidence":"auto"},{"platform":"cursor","label":"Cursor","status":"unknown","note":"We have not crawled the repo tree, so we will not guess","evidence":"auto"},{"platform":"codex","label":"Codex","status":"unknown","note":"We have not crawled the repo tree, so we will not guess","evidence":"auto"},{"platform":"gemini_cli","label":"Gemini CLI","status":"unknown","note":"The spec defines no detection rule for Gemini","evidence":"auto"},{"platform":"copilot","label":"Copilot","status":"unknown","note":"We have not crawled the repo tree, so we will not guess","evidence":"auto"}],"addedAt":"2026-09-21T08:34:24.006123+00:00","featured":false,"threads":[],"ghComments":[],"collectedAt":"2026-09-21T08:34:24.006123+00:00","iconUrl":"https://m.agentalley.io/logo/6ef0ad61519e4120b7e8.png","media":{"shots":[{"url":"https://m.agentalley.io/r2/bea4345ed06bffa34873.webp","alt":"PDF Processor - Extract Data from PDFs — CleanShot 2026-07-13 at 20 15 47@2x (from the gooseworks-ai/goose-skills README)","w":1600,"h":929,"source":"author","creditName":"gooseworks-ai","creditUrl":"https://github.com/gooseworks-ai/goose-skills/blob/main/README.md","from":"repo"},{"url":"https://m.agentalley.io/r2/2d15be91e8a6d3945693.webp","alt":"PDF Processor - Extract Data from PDFs — CleanShot 2026-07-13 at 20 16 54@2x (from the gooseworks-ai/goose-skills README)","w":1600,"h":861,"source":"author","creditName":"gooseworks-ai","creditUrl":"https://github.com/gooseworks-ai/goose-skills/blob/main/README.md","from":"repo"}]}}