Derive an API client from a recorded session skill

Reverse-engineer a website's internal API by recording browser traffic into a HAR file, then generate a standalone client or CLI that calls the endpoints directly, with no browser needed after the first recording.

by fcakyon·Apache-2.0 license·★ 1,162 Stars on the repo·GitHub ↗

Use now

Files of Derive an API client from a recorded session

fcakyon/main1 file shown
SKILL.md
Show the full text88 lines

Derive an API client from a recorded session

Driving a browser is the right tool for the first visit and the wrong tool for the hundredth. This skill records a site's network traffic once while you use it, then turns the captured requests into a standalone client (script, CLI, or library) that talks to the site's internal API directly.

The recording alone contains everything needed: agent-browser embeds text response bodies (JSON/HTML/JS) in the HAR by default, so endpoint shapes can be studied offline after the browser is closed.

Workflow

1. Record     Start HAR capture, drive the flows you want in the client
2. Identify   Find the real API endpoints among the noise
3. Extract    Pull request shapes, response schemas, and auth material
4. Generate   Write the client, one function per flow
5. Verify     Call every endpoint for real before declaring done

1. Record

agent-browser network har start          # embeds text response bodies by default
# ... drive the site: search, open a detail page, paginate, etc. ...
agent-browser network har stop /tmp/site.har
  • Exercise every flow the client should support, and run each one at least twice with different inputs (two search terms, two detail pages). Diffing the recorded URLs reveals which parts are parameters.
  • If the site needs login, log in before starting the HAR so credentials don't land in the recording unnecessarily. The session cookies are exported separately in step 3.
  • --content all embeds binary bodies too (base64); --content none disables embedding. Per-body cap is 2 MB.

While the session is still open, agent-browser network requests and network request <id> give the same data interactively — but only the HAR survives navigation and browser close, so prefer it for anything multi-page.

2. Identify endpoints

Query the HAR with jq:

# All JSON API calls: method, URL, status
jq -r '.log.entries[]
  | select(.response.content.mimeType | test("json"))
  | "\(.request.method) \(.response.status) \(.request.url)"' /tmp/site.har

Ignore analytics and infrastructure noise: telemetry endpoints (/collect, /track, /beacon, /log), third-party domains (google-analytics, segment, sentry, datadog, intercom, hotjar), and static assets. The real API is usually first-party, JSON, and correlates with the actions you performed.

3. Extract shapes and auth

# Full detail for one endpoint: request headers, POST body, response body
jq '.log.entries[] | select(.request.url | test("api/search"))
  | {request: {method: .request.method, headers: .request.headers,
     postData: .request.postData.text},
     response: .response.content.text}' /tmp/site.har
  • Response schema: read .response.content.text — this is the real payload, use it to derive types.
  • Auth: compare request headers across endpoints. Look for authorization, cookie, x-csrf-token, x-api-key, and site-specific x-* headers. Replay only the ones that matter — test by omission in step 5.
  • Cookies: export the live session with agent-browser cookies get --json > cookies.json for the client to load at runtime. Never hardcode cookie values into generated source.

4. Generate the client

  • One function per recorded flow (search(query), getItem(id)), typed from the observed response bodies.
  • Auth material (cookies, bearer tokens) loads from a file or environment variable, with a clear error telling the user to re-run the browser login when it expires.
  • Reproduce the headers the API actually requires — some sites 403 without a matching user-agent, referer, or x-requested-with.
  • Keep pagination, sort, and filter parameters that appeared in the recorded query strings as function options.

5. Verify

Call every generated function against the live API and compare the response shape with the recording. Common failures:

Symptom Cause Fix
401/403 Expired or missing session Re-login via agent-browser, re-export cookies
403/419 on writes CSRF token is per-session or per-form Fetch the token endpoint first, or keep that flow browser-driven
Works then breaks Signed/expiring request params Fall back to the browser for that step; derive the rest
Different shape than HAR A/B tests or geo-dependent responses Re-record and treat the union as optional fields

Caveats

  • Internal APIs are unversioned and change without notice — keep the HAR so the client can be re-derived.
  • Respect the site's terms of service and rate limits; add delays for bulk fetching.
  • HAR files contain live session credentials (cookies, tokens, POST bodies). Treat them like secrets: keep them out of version control and delete them when done.
1---
2name: derive-client
3description: Reverse-engineer a website's internal API by recording browser traffic into a HAR file, then generate a standalone client or CLI that calls the endpoints directly, with no browser needed after the first recording. Use when asked to "derive a client", "build a CLI for <site>", "reverse engineer this site's API", "record network requests", "turn this site into an API", or when the same site will be automated repeatedly and direct HTTP calls would beat driving the browser every time.
4allowed-tools: Bash(agent-browser:*), Bash(npx agent-browser:*)
5license: Apache-2.0
6---
7 
8# Derive an API client from a recorded session
9 
10Driving a browser is the right tool for the first visit and the wrong tool for the hundredth. This skill records a site's network traffic once while you use it, then turns the captured requests into a standalone client (script, CLI, or library) that talks to the site's internal API directly.
11 
12The recording alone contains everything needed: agent-browser embeds text response bodies (JSON/HTML/JS) in the HAR by default, so endpoint shapes can be studied offline after the browser is closed.
13 
14## Workflow
15 
16```
171. Record Start HAR capture, drive the flows you want in the client
182. Identify Find the real API endpoints among the noise
193. Extract Pull request shapes, response schemas, and auth material
204. Generate Write the client, one function per flow
215. Verify Call every endpoint for real before declaring done
22```
23 
24## 1. Record
25 
26```bash
27agent-browser network har start # embeds text response bodies by default
28# ... drive the site: search, open a detail page, paginate, etc. ...
29agent-browser network har stop /tmp/site.har
30```
31 
32- Exercise **every flow the client should support**, and run each one at least twice with different inputs (two search terms, two detail pages). Diffing the recorded URLs reveals which parts are parameters.
33- If the site needs login, log in **before** starting the HAR so credentials don't land in the recording unnecessarily. The session cookies are exported separately in step 3.
34- `--content all` embeds binary bodies too (base64); `--content none` disables embedding. Per-body cap is 2 MB.
35 
36While the session is still open, `agent-browser network requests` and `network request <id>` give the same data interactively — but only the HAR survives navigation and browser close, so prefer it for anything multi-page.
37 
38## 2. Identify endpoints
39 
40Query the HAR with `jq`:
41 
42```bash
43# All JSON API calls: method, URL, status
44jq -r '.log.entries[]
45 | select(.response.content.mimeType | test("json"))
46 | "\(.request.method) \(.response.status) \(.request.url)"' /tmp/site.har
47```
48 
49Ignore analytics and infrastructure noise: telemetry endpoints (`/collect`, `/track`, `/beacon`, `/log`), third-party domains (google-analytics, segment, sentry, datadog, intercom, hotjar), and static assets. The real API is usually first-party, JSON, and correlates with the actions you performed.
50 
51## 3. Extract shapes and auth
52 
53```bash
54# Full detail for one endpoint: request headers, POST body, response body
55jq '.log.entries[] | select(.request.url | test("api/search"))
56 | {request: {method: .request.method, headers: .request.headers,
57 postData: .request.postData.text},
58 response: .response.content.text}' /tmp/site.har
59```
60 
61- **Response schema**: read `.response.content.text` — this is the real payload, use it to derive types.
62- **Auth**: compare request headers across endpoints. Look for `authorization`, `cookie`, `x-csrf-token`, `x-api-key`, and site-specific `x-*` headers. Replay only the ones that matter — test by omission in step 5.
63- **Cookies**: export the live session with `agent-browser cookies get --json > cookies.json` for the client to load at runtime. Never hardcode cookie values into generated source.
64 
65## 4. Generate the client
66 
67- One function per recorded flow (`search(query)`, `getItem(id)`), typed from the observed response bodies.
68- Auth material (cookies, bearer tokens) loads from a file or environment variable, with a clear error telling the user to re-run the browser login when it expires.
69- Reproduce the headers the API actually requires — some sites 403 without a matching `user-agent`, `referer`, or `x-requested-with`.
70- Keep pagination, sort, and filter parameters that appeared in the recorded query strings as function options.
71 
72## 5. Verify
73 
74Call every generated function against the live API and compare the response shape with the recording. Common failures:
75 
76| Symptom | Cause | Fix |
77|---------|-------|-----|
78| 401/403 | Expired or missing session | Re-login via agent-browser, re-export cookies |
79| 403/419 on writes | CSRF token is per-session or per-form | Fetch the token endpoint first, or keep that flow browser-driven |
80| Works then breaks | Signed/expiring request params | Fall back to the browser for that step; derive the rest |
81| Different shape than HAR | A/B tests or geo-dependent responses | Re-record and treat the union as optional fields |
82 
83## Caveats
84 
85- Internal APIs are unversioned and change without notice — keep the HAR so the client can be re-derived.
86- Respect the site's terms of service and rate limits; add delays for bulk fetching.
87- HAR files contain live session credentials (cookies, tokens, POST bodies). Treat them like secrets: keep them out of version control and delete them when done.
88 

Discussion

Alternatives

API and interface designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.Coding · MITContext7Pulls up-to-date, version-specific library docs and code examples into the prompt so the AI stops inventing old APIs.Coding · MITContext7 Documentation LookupFetch up-to-date documentation and code examples for any library, framework, SDK, CLI tool, or cloud service. Use whenever the user asks about a specific library — even well-known ones like React, Next.js, Prisma, Express, Tailwind, Django, or Spring Boot — because training data may not reflect recent API changes or version updates. Always use for: API syntax questions, configuration options, version migration issues, "how do I" questions mentioning a library name, debugging that involves library-specific behavior, setup instructions, and CLI tool usage. Use even when you think you know the answer. Do not rely on training data for API details, signatures, or configuration options — they are frequently out of date. Prefer this over web search for library documentation.Coding · MITAdaptyv Bio Foundry APIHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.Science · MIT