Arxiv skill
Search and retrieve arXiv academic papers by topic, category, or paper ID — with AlphaXiv-enriched AI-generated overviews.
by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗
npx degit danielmiessler/LifeOS/LifeOS/install/skills/ArXiv#main ~/.claude/skills/arxivChecked ·commit main
Files of Arxiv
Show the full text120 lines
Customization
Before executing, check for user customizations at:
~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/ArXiv/
If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
You MUST send this notification BEFORE doing anything else when this skill is invoked.
Send voice notification:
curl -s -X POST http://localhost:31337/notify \ -H "Content-Type: application/json" \ -d '{"message": "Running the WORKFLOWNAME workflow in the ArXiv skill to ACTION"}' \ > /dev/null 2>&1 &Output text notification:
Running the **WorkflowName** workflow in the **ArXiv** skill to ACTION...
This is not optional. Execute this curl command immediately upon skill invocation.
ArXiv
What It Does
Searches and retrieves arXiv academic papers by topic, category, or paper ID, and pulls AlphaXiv's AI-generated overviews when a paper has one. Covers the cs.AI / cs.LG / cs.CL / cs.CR / cs.MA / cs.SE / cs.IR categories. Three workflows: Latest, Search, Paper. No API keys needed.
The Problem
arXiv ships thousands of papers a day and its native search is clunky — Atom XML, three-second rate limits, fields you have to know by name, and a lastUpdatedDate that quietly resurfaces old papers as if they were new. Reading a raw paper to decide whether it's worth your time is slow. This skill wraps the query mechanics, handles the XML, and layers AlphaXiv overviews on top so you can triage a paper in seconds instead of reading the whole PDF first.
How It Works
Uses arXiv's Atom API for search and discovery, and AlphaXiv's markdown endpoint for enriched paper overviews. Search fields, boolean operators, sort order, and pagination are all handled for you; overviews are fetched per paper ID when available (a 404 just means no overview exists yet).
Workflow Routing
| Trigger | Workflow |
|---|---|
| "latest papers in X", "new papers on X", "what's new in AI research" | Workflows/Latest.md |
| "search arxiv for X", "find papers about X", "arxiv papers on X" | Workflows/Search.md |
arxiv URL, paper ID like 2401.12345, "explain this paper" |
Workflows/Paper.md |
Quick Reference
arXiv API (no auth):
- Base:
https://export.arxiv.org/api/query - Search fields:
ti:(title),au:(author),abs:(abstract),cat:(category),all:(everything) - Booleans:
AND,OR,ANDNOT - Sort:
sortBy=lastUpdatedDate&sortOrder=descendingfor latest - Pagination:
start=0&max_results=10(max 2000 per call) - Rate limit: 3s between calls
AlphaXiv enrichment (no auth):
- Overview:
curl -s "https://alphaxiv.org/overview/{PAPER_ID}.md" - Full text:
curl -s "https://alphaxiv.org/abs/{PAPER_ID}.md"(fallback) - Not all papers have overviews — 404 means analysis not yet generated
Key categories for our work:
cs.AI— Artificial Intelligencecs.LG— Machine Learningcs.CL— Computation and Language (NLP/LLMs)cs.CR— Cryptography and Securitycs.SE— Software Engineeringcs.MA— Multi-Agent Systemscs.IR— Information Retrieval
Examples
Example 1: Latest papers in a category
User: "what's new in AI safety papers this week"
→ Latest workflow: queries cat:cs.AI sorted by lastUpdatedDate, filters by <published> date
→ Returns titles, authors, abstracts, links
Example 2: Topic search
User: "search arxiv for prompt injection defenses"
→ Search workflow: all:"prompt injection" query with boolean refinement
→ Returns ranked matches with abstracts
Example 3: Single paper lookup
User: "explain this paper: 2401.12345"
→ Paper workflow: fetches metadata, pulls AlphaXiv overview (falls back to abstract on 404)
→ Returns summary plus link to PDF
Gotchas
- arXiv API requires HTTPS and
-L(follows redirects). HTTP 301s to HTTPS silently. - arXiv API returns Atom XML, not JSON. Parse with text processing, not
jq. lastUpdatedDateincludes edits to old papers. For truly new submissions, check<published>dates.- AlphaXiv overviews are AI-generated summaries. Great for quick understanding, but verify claims against the actual paper for anything you'd cite.
- arXiv API rate limit is 3 seconds between calls. Batch your queries.
max_resultscaps at 2000. For broader sweeps, paginate withstart.- Category search (
cat:cs.AI) returns papers with that as primary OR cross-listed category.
Execution Log
After completing any workflow, append a single JSONL entry:
echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"ArXiv","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
| 1 | |
| 2 | name ArXiv |
| 3 | version 1.0.9 |
| 4 | description "Search and retrieve arXiv academic papers by topic, category, or paper ID — with AlphaXiv-enriched AI-generated overviews. Uses arXiv Atom API across cs.AI/cs.LG/cs.CL/cs.CR/cs.MA/cs.SE/cs.IR. Three workflows: Latest, Search, Paper. USE WHEN arxiv, papers, latest papers, research papers, recent ML papers, paper lookup, summarize paper, latest LLM papers, AI safety papers, cs.AI latest. NOT FOR general research (Research), URL parsing, or annual reports." |
| 5 | |
| 6 | |
| 7 | ## Customization |
| 8 | |
| 9 | **Before executing, check for user customizations at:** |
| 10 | `~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/ArXiv/` |
| 11 | |
| 12 | If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults. |
| 13 | |
| 14 | |
| 15 | ## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION) |
| 16 | |
| 17 | **You MUST send this notification BEFORE doing anything else when this skill is invoked.** |
| 18 | |
| 19 | **Send voice notification**: |
| 20 | |
| 21 | curl -s -X POST http://localhost:31337/notify \ |
| 22 | -H "Content-Type: application/json" \ |
| 23 | -d '{"message": "Running the WORKFLOWNAME workflow in the ArXiv skill to ACTION"}' \ |
| 24 | > /dev/null 2>&1 & |
| 25 | |
| 26 | |
| 27 | **Output text notification**: |
| 28 | |
| 29 | Running the **WorkflowName** workflow in the **ArXiv** skill to ACTION... |
| 30 | |
| 31 | |
| 32 | **This is not optional. Execute this curl command immediately upon skill invocation.** |
| 33 | |
| 34 | # ArXiv |
| 35 | |
| 36 | ## What It Does |
| 37 | |
| 38 | Searches and retrieves arXiv academic papers by topic, category, or paper ID, and pulls AlphaXiv's AI-generated overviews when a paper has one. Covers the cs.AI / cs.LG / cs.CL / cs.CR / cs.MA / cs.SE / cs.IR categories. Three workflows: Latest, Search, Paper. No API keys needed. |
| 39 | |
| 40 | ## The Problem |
| 41 | |
| 42 | arXiv ships thousands of papers a day and its native search is clunky — Atom XML, three-second rate limits, fields you have to know by name, and a `lastUpdatedDate` that quietly resurfaces old papers as if they were new. Reading a raw paper to decide whether it's worth your time is slow. This skill wraps the query mechanics, handles the XML, and layers AlphaXiv overviews on top so you can triage a paper in seconds instead of reading the whole PDF first. |
| 43 | |
| 44 | ## How It Works |
| 45 | |
| 46 | Uses arXiv's Atom API for search and discovery, and AlphaXiv's markdown endpoint for enriched paper overviews. Search fields, boolean operators, sort order, and pagination are all handled for you; overviews are fetched per paper ID when available (a 404 just means no overview exists yet). |
| 47 | |
| 48 | ## Workflow Routing |
| 49 | |
| 50 | | Trigger | Workflow | |
| 51 | |---------|----------| |
| 52 | | "latest papers in X", "new papers on X", "what's new in AI research" | `Workflows/Latest.md` | |
| 53 | | "search arxiv for X", "find papers about X", "arxiv papers on X" | `Workflows/Search.md` | |
| 54 | | arxiv URL, paper ID like `2401.12345`, "explain this paper" | `Workflows/Paper.md` | |
| 55 | |
| 56 | ## Quick Reference |
| 57 | |
| 58 | **arXiv API** (no auth): |
| 59 | Base: `https://export.arxiv.org/api/query` |
| 60 | Search fields: `ti:` (title), `au:` (author), `abs:` (abstract), `cat:` (category), `all:` (everything) |
| 61 | Booleans: `AND`, `OR`, `ANDNOT` |
| 62 | Sort: `sortBy=lastUpdatedDate&sortOrder=descending` for latest |
| 63 | Pagination: `start=0&max_results=10` (max 2000 per call) |
| 64 | Rate limit: 3s between calls |
| 65 | |
| 66 | **AlphaXiv enrichment** (no auth): |
| 67 | Overview: `curl -s "https://alphaxiv.org/overview/{PAPER_ID}.md"` |
| 68 | Full text: `curl -s "https://alphaxiv.org/abs/{PAPER_ID}.md"` (fallback) |
| 69 | Not all papers have overviews — 404 means analysis not yet generated |
| 70 | |
| 71 | **Key categories for our work:** |
| 72 | `cs.AI` — Artificial Intelligence |
| 73 | `cs.LG` — Machine Learning |
| 74 | `cs.CL` — Computation and Language (NLP/LLMs) |
| 75 | `cs.CR` — Cryptography and Security |
| 76 | `cs.SE` — Software Engineering |
| 77 | `cs.MA` — Multi-Agent Systems |
| 78 | `cs.IR` — Information Retrieval |
| 79 | |
| 80 | ## Examples |
| 81 | |
| 82 | **Example 1: Latest papers in a category** |
| 83 | |
| 84 | User: "what's new in AI safety papers this week" |
| 85 | → Latest workflow: queries cat:cs.AI sorted by lastUpdatedDate, filters by <published> date |
| 86 | → Returns titles, authors, abstracts, links |
| 87 | |
| 88 | |
| 89 | **Example 2: Topic search** |
| 90 | |
| 91 | User: "search arxiv for prompt injection defenses" |
| 92 | → Search workflow: all:"prompt injection" query with boolean refinement |
| 93 | → Returns ranked matches with abstracts |
| 94 | |
| 95 | |
| 96 | **Example 3: Single paper lookup** |
| 97 | |
| 98 | User: "explain this paper: 2401.12345" |
| 99 | → Paper workflow: fetches metadata, pulls AlphaXiv overview (falls back to abstract on 404) |
| 100 | → Returns summary plus link to PDF |
| 101 | |
| 102 | |
| 103 | ## Gotchas |
| 104 | |
| 105 | arXiv API **requires HTTPS** and `-L` (follows redirects). HTTP 301s to HTTPS silently. |
| 106 | arXiv API returns Atom XML, not JSON. Parse with text processing, not `jq`. |
| 107 | `lastUpdatedDate` includes edits to old papers. For truly new submissions, check `<published>` dates. |
| 108 | AlphaXiv overviews are AI-generated summaries. Great for quick understanding, but verify claims against the actual paper for anything you'd cite. |
| 109 | arXiv API rate limit is 3 seconds between calls. Batch your queries. |
| 110 | `max_results` caps at 2000. For broader sweeps, paginate with `start`. |
| 111 | Category search (`cat:cs.AI`) returns papers with that as primary OR cross-listed category. |
| 112 | |
| 113 | ## Execution Log |
| 114 | |
| 115 | After completing any workflow, append a single JSONL entry: |
| 116 | |
| 117 | |
| 118 | echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"ArXiv","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl |
| 119 | |
| 120 |
Discussion
Alternatives
Browse more free Claude skills or everything in Product.