Arxiv skill

Search and retrieve arXiv academic papers by topic, category, or paper ID — with AlphaXiv-enriched AI-generated overviews.

by danielmiessler·MIT license·★ 19,269 Stars on the repo·GitHub ↗

Use now

Files of Arxiv

danielmiessler/main1 file shown
SKILL.md
Show the full text120 lines

Customization

Before executing, check for user customizations at: ~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/ArXiv/

If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.

🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)

You MUST send this notification BEFORE doing anything else when this skill is invoked.

  1. Send voice notification:

    curl -s -X POST http://localhost:31337/notify \
      -H "Content-Type: application/json" \
      -d '{"message": "Running the WORKFLOWNAME workflow in the ArXiv skill to ACTION"}' \
      > /dev/null 2>&1 &
    
  2. Output text notification:

    Running the **WorkflowName** workflow in the **ArXiv** skill to ACTION...
    

This is not optional. Execute this curl command immediately upon skill invocation.

ArXiv

What It Does

Searches and retrieves arXiv academic papers by topic, category, or paper ID, and pulls AlphaXiv's AI-generated overviews when a paper has one. Covers the cs.AI / cs.LG / cs.CL / cs.CR / cs.MA / cs.SE / cs.IR categories. Three workflows: Latest, Search, Paper. No API keys needed.

The Problem

arXiv ships thousands of papers a day and its native search is clunky — Atom XML, three-second rate limits, fields you have to know by name, and a lastUpdatedDate that quietly resurfaces old papers as if they were new. Reading a raw paper to decide whether it's worth your time is slow. This skill wraps the query mechanics, handles the XML, and layers AlphaXiv overviews on top so you can triage a paper in seconds instead of reading the whole PDF first.

How It Works

Uses arXiv's Atom API for search and discovery, and AlphaXiv's markdown endpoint for enriched paper overviews. Search fields, boolean operators, sort order, and pagination are all handled for you; overviews are fetched per paper ID when available (a 404 just means no overview exists yet).

Workflow Routing

Trigger Workflow
"latest papers in X", "new papers on X", "what's new in AI research" Workflows/Latest.md
"search arxiv for X", "find papers about X", "arxiv papers on X" Workflows/Search.md
arxiv URL, paper ID like 2401.12345, "explain this paper" Workflows/Paper.md

Quick Reference

arXiv API (no auth):

  • Base: https://export.arxiv.org/api/query
  • Search fields: ti: (title), au: (author), abs: (abstract), cat: (category), all: (everything)
  • Booleans: AND, OR, ANDNOT
  • Sort: sortBy=lastUpdatedDate&sortOrder=descending for latest
  • Pagination: start=0&max_results=10 (max 2000 per call)
  • Rate limit: 3s between calls

AlphaXiv enrichment (no auth):

  • Overview: curl -s "https://alphaxiv.org/overview/{PAPER_ID}.md"
  • Full text: curl -s "https://alphaxiv.org/abs/{PAPER_ID}.md" (fallback)
  • Not all papers have overviews — 404 means analysis not yet generated

Key categories for our work:

  • cs.AI — Artificial Intelligence
  • cs.LG — Machine Learning
  • cs.CL — Computation and Language (NLP/LLMs)
  • cs.CR — Cryptography and Security
  • cs.SE — Software Engineering
  • cs.MA — Multi-Agent Systems
  • cs.IR — Information Retrieval

Examples

Example 1: Latest papers in a category

User: "what's new in AI safety papers this week"
→ Latest workflow: queries cat:cs.AI sorted by lastUpdatedDate, filters by <published> date
→ Returns titles, authors, abstracts, links

Example 2: Topic search

User: "search arxiv for prompt injection defenses"
→ Search workflow: all:"prompt injection" query with boolean refinement
→ Returns ranked matches with abstracts

Example 3: Single paper lookup

User: "explain this paper: 2401.12345"
→ Paper workflow: fetches metadata, pulls AlphaXiv overview (falls back to abstract on 404)
→ Returns summary plus link to PDF

Gotchas

  • arXiv API requires HTTPS and -L (follows redirects). HTTP 301s to HTTPS silently.
  • arXiv API returns Atom XML, not JSON. Parse with text processing, not jq.
  • lastUpdatedDate includes edits to old papers. For truly new submissions, check <published> dates.
  • AlphaXiv overviews are AI-generated summaries. Great for quick understanding, but verify claims against the actual paper for anything you'd cite.
  • arXiv API rate limit is 3 seconds between calls. Batch your queries.
  • max_results caps at 2000. For broader sweeps, paginate with start.
  • Category search (cat:cs.AI) returns papers with that as primary OR cross-listed category.

Execution Log

After completing any workflow, append a single JSONL entry:

echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"ArXiv","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
1---
2name: ArXiv
3version: 1.0.9
4description: "Search and retrieve arXiv academic papers by topic, category, or paper ID — with AlphaXiv-enriched AI-generated overviews. Uses arXiv Atom API across cs.AI/cs.LG/cs.CL/cs.CR/cs.MA/cs.SE/cs.IR. Three workflows: Latest, Search, Paper. USE WHEN arxiv, papers, latest papers, research papers, recent ML papers, paper lookup, summarize paper, latest LLM papers, AI safety papers, cs.AI latest. NOT FOR general research (Research), URL parsing, or annual reports."
5---
6 
7## Customization
8 
9**Before executing, check for user customizations at:**
10`~/.claude/LIFEOS/USER/CUSTOMIZATIONS/SKILLS/ArXiv/`
11 
12If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
13 
14 
15## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
16 
17**You MUST send this notification BEFORE doing anything else when this skill is invoked.**
18 
191. **Send voice notification**:
20 ```bash
21 curl -s -X POST http://localhost:31337/notify \
22 -H "Content-Type: application/json" \
23 -d '{"message": "Running the WORKFLOWNAME workflow in the ArXiv skill to ACTION"}' \
24 > /dev/null 2>&1 &
25 ```
26 
272. **Output text notification**:
28 ```
29 Running the **WorkflowName** workflow in the **ArXiv** skill to ACTION...
30 ```
31 
32**This is not optional. Execute this curl command immediately upon skill invocation.**
33 
34# ArXiv
35 
36## What It Does
37 
38Searches and retrieves arXiv academic papers by topic, category, or paper ID, and pulls AlphaXiv's AI-generated overviews when a paper has one. Covers the cs.AI / cs.LG / cs.CL / cs.CR / cs.MA / cs.SE / cs.IR categories. Three workflows: Latest, Search, Paper. No API keys needed.
39 
40## The Problem
41 
42arXiv ships thousands of papers a day and its native search is clunky — Atom XML, three-second rate limits, fields you have to know by name, and a `lastUpdatedDate` that quietly resurfaces old papers as if they were new. Reading a raw paper to decide whether it's worth your time is slow. This skill wraps the query mechanics, handles the XML, and layers AlphaXiv overviews on top so you can triage a paper in seconds instead of reading the whole PDF first.
43 
44## How It Works
45 
46Uses arXiv's Atom API for search and discovery, and AlphaXiv's markdown endpoint for enriched paper overviews. Search fields, boolean operators, sort order, and pagination are all handled for you; overviews are fetched per paper ID when available (a 404 just means no overview exists yet).
47 
48## Workflow Routing
49 
50| Trigger | Workflow |
51|---------|----------|
52| "latest papers in X", "new papers on X", "what's new in AI research" | `Workflows/Latest.md` |
53| "search arxiv for X", "find papers about X", "arxiv papers on X" | `Workflows/Search.md` |
54| arxiv URL, paper ID like `2401.12345`, "explain this paper" | `Workflows/Paper.md` |
55 
56## Quick Reference
57 
58**arXiv API** (no auth):
59- Base: `https://export.arxiv.org/api/query`
60- Search fields: `ti:` (title), `au:` (author), `abs:` (abstract), `cat:` (category), `all:` (everything)
61- Booleans: `AND`, `OR`, `ANDNOT`
62- Sort: `sortBy=lastUpdatedDate&sortOrder=descending` for latest
63- Pagination: `start=0&max_results=10` (max 2000 per call)
64- Rate limit: 3s between calls
65 
66**AlphaXiv enrichment** (no auth):
67- Overview: `curl -s "https://alphaxiv.org/overview/{PAPER_ID}.md"`
68- Full text: `curl -s "https://alphaxiv.org/abs/{PAPER_ID}.md"` (fallback)
69- Not all papers have overviews — 404 means analysis not yet generated
70 
71**Key categories for our work:**
72- `cs.AI` — Artificial Intelligence
73- `cs.LG` — Machine Learning
74- `cs.CL` — Computation and Language (NLP/LLMs)
75- `cs.CR` — Cryptography and Security
76- `cs.SE` — Software Engineering
77- `cs.MA` — Multi-Agent Systems
78- `cs.IR` — Information Retrieval
79 
80## Examples
81 
82**Example 1: Latest papers in a category**
83```
84User: "what's new in AI safety papers this week"
85→ Latest workflow: queries cat:cs.AI sorted by lastUpdatedDate, filters by <published> date
86→ Returns titles, authors, abstracts, links
87```
88 
89**Example 2: Topic search**
90```
91User: "search arxiv for prompt injection defenses"
92→ Search workflow: all:"prompt injection" query with boolean refinement
93→ Returns ranked matches with abstracts
94```
95 
96**Example 3: Single paper lookup**
97```
98User: "explain this paper: 2401.12345"
99→ Paper workflow: fetches metadata, pulls AlphaXiv overview (falls back to abstract on 404)
100→ Returns summary plus link to PDF
101```
102 
103## Gotchas
104 
105- arXiv API **requires HTTPS** and `-L` (follows redirects). HTTP 301s to HTTPS silently.
106- arXiv API returns Atom XML, not JSON. Parse with text processing, not `jq`.
107- `lastUpdatedDate` includes edits to old papers. For truly new submissions, check `<published>` dates.
108- AlphaXiv overviews are AI-generated summaries. Great for quick understanding, but verify claims against the actual paper for anything you'd cite.
109- arXiv API rate limit is 3 seconds between calls. Batch your queries.
110- `max_results` caps at 2000. For broader sweeps, paginate with `start`.
111- Category search (`cat:cs.AI`) returns papers with that as primary OR cross-listed category.
112 
113## Execution Log
114 
115After completing any workflow, append a single JSONL entry:
116 
117```bash
118echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"ArXiv","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
119```
120 

Discussion

Alternatives

Research methodology design for health literacy and medication adherence in aotearoa new zealandExplore the methodological design for researching health literacy and its impact on medication adherence among adults with chronic diseases in Aotearoa New Zealand.Business & ops · CC0-1.0Scientific critical thinkingEvaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.Science · MITAcademic research synthesizerAcademic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources. <example>Context: A podcast episode needs a segment grounded in peer-reviewed evidence with formal citations. user: "Research the current state of transformer efficiency techniques for the episode, with proper academic citations." assistant: "I'll use the academic-research-synthesizer agent to search arXiv and Semantic Scholar, extract full-text findings via WebFetch, and produce a cited literature synthesis with confidence levels." <commentary>Use academic-research-synthesizer (not comprehensive-researcher) when the episode segment needs peer-reviewed sourcing, formal citation format, and explicit confidence tagging rather than general-purpose multi-source coverage.</commentary></example> <example>Context: The episode-orchestrator has routed a "literature review" request for a technical deep-dive segment. user: "Summarize the research landscape on federated learning privacy guarantees." assistant: "I'll invoke academic-research-synthesizer to systematically search academic sources, note peer-review status per source, and synthesize consensus vs. open debates."</example>Business & ops · MITAcademic researcherAcademic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation. <example>Context: The research-orchestrator has kicked off Phase 4 parallel research on 'efficacy of intermittent fasting' and needs peer-reviewed evidence. user: "Find the academic evidence on intermittent fasting outcomes." assistant: "I'll use the academic-researcher agent to search Semantic Scholar, PubMed, and OpenAlex for peer-reviewed studies and write structured findings to academic-research.md." <commentary>The request is specifically for scholarly/peer-reviewed evidence rather than general web coverage or code, so academic-researcher (not web-researcher or technical-researcher) is the right specialist.</commentary></example> <example>Context: The user wants a literature review comparing methodologies across studies on a topic. user: "Can you review the literature on transformer model interpretability and identify research gaps?" assistant: "Let me invoke the academic-researcher agent to pull foundational and recent papers, extract methodologies, and surface open research gaps." <commentary>Literature review, methodology extraction, and research-gap identification are core academic-researcher capabilities, distinct from technical-researcher's focus on code repositories and implementations.</commentary></example>Business & ops · MIT