Agentic video understanding skill
Use when an agent must extract moments, quotes, objections, hooks, or evidence from long video or audio cheaper than full-frame ingest — sales calls, podcasts, YouTube episodes, Loom trials, discovery recordings.
by ericosiu·MIT license·★ 3,615 Stars on the repo·GitHub ↗
npx degit ericosiu/ai-marketing-skills/agentic-video-understanding#main ~/.claude/skills/agentic-video-understandingChecked ·commit main
Files of Agentic video understanding
Show the full text95 lines
Agentic video understanding
Hireable understanding layer. The model takes a goal and decides what to watch, at what speed, and through which modality (frames, audio, transcript), fetching only the moments needed. Vendor claims: up to ~66% lower cost and ~88% fewer tokens vs static fixed-FPS ingest, with higher accuracy.
What this is / is not
Is: goal → watch only what you need → timestamps + quotes + confidence.
Is not: a video editor. Do not cut, overlay, caption-burn, render, schedule, post, email, or write CRM from this skill. Hand cuts to Overlap, FFmpeg, or net-new-video-editor. Approvals stay with the calling lane.
When to use
- Pre-call / sales-call mining: buyer objection, next step, competitive mention
- Shortform scoring: find a 3-second standalone hook and in/out points
- Longform / X research: named-person + contrast moments in podcast or YouTube tape
- Talent review: bar evidence in a Loom or trial recording
- Client audit: every mention of a keyword across a discovery recording
Skip when the job is already a clean transcript and you only need text search.
Inputs
| Field | Required | Notes |
|---|---|---|
source |
yes | URL or local media path the runtime can read |
goal |
yes | One sentence retrieval goal |
keywords |
no | Extra strings to bias retrieval |
max_moments |
no | Default 5 |
modality |
no | auto (default), frames, audio, or transcript |
Process
- Restate the goal as 1–3 retrieval queries. Done when each query is falsifiable (you would know if a moment matched).
- Call Gemini agentic video understanding (Gemini API or AI Studio) with
source, queries,max_moments, and modality preference. Prefer the agentic path over fixed-FPS full ingest when available. Done when the API returns candidate windows or an explicit empty set. - Normalize moments into the output schema below. Flag paraphrase vs verbatim. Drop fabricated timestamps. Done when every kept moment has
t_start,t_end,modality,quote,why,confidence. - Stop and hand off to the caller. Do not cut, overlay, schedule, publish, email, or CRM-write.
Output schema
Markdown for humans, optional JSON for machines:
{
"goal": "",
"source": "",
"moments": [
{
"t_start": "MM:SS",
"t_end": "MM:SS",
"modality": "frames|audio|transcript",
"quote": "",
"verbatim": true,
"why": "",
"confidence": 0.0
}
],
"empty_reason": null,
"tokens_note": "agentic path used|fallback static ingest"
}
Hard gates
- No full fixed-FPS ingest when the agentic path is available
- No invented timestamps or quotes
- No dumping full transcripts or client PII into public artifacts
- No cut / render / overlay / schedule / publish / send from this skill
Setup
- Gemini API key or Google AI Studio access: https://ai.studio
- See Google’s developer guide for agentic video understanding in Gemini
- Env:
GEMINI_API_KEY(or the project’s existing Google AI credential)
Caller one-liners
- Pre-call:
goal="exact next-step commitment and any pricing pushback" - Shortform:
goal="best 3-second standalone hook; return in/out for one clip" - Talent:
goal="evidence they hit the role bar on X; max 5 moments" - Audit:
goal="every mention of Reddit, AEO, or budget"
Completion
Done when the caller has the schema above (or a documented empty set) and this skill has performed no side effects beyond the Gemini read.
| 1 | |
| 2 | name Agentic video understanding |
| 3 | description >- |
| 4 | Use when an agent must extract moments, quotes, objections, hooks, or evidence |
| 5 | from long video or audio cheaper than full-frame ingest — sales calls, |
| 6 | podcasts, YouTube episodes, Loom trials, discovery recordings. Goal-directed |
| 7 | watch via Gemini agentic video understanding (frames, audio, or transcript). |
| 8 | Not for cutting, overlays, rendering, scheduling, or publishing. |
| 9 | |
| 10 | |
| 11 | # Agentic video understanding |
| 12 | |
| 13 | Hireable understanding layer. The model takes a goal and decides what to watch, at what speed, and through which modality (frames, audio, transcript), fetching only the moments needed. Vendor claims: up to ~66% lower cost and ~88% fewer tokens vs static fixed-FPS ingest, with higher accuracy. |
| 14 | |
| 15 | ## What this is / is not |
| 16 | |
| 17 | **Is:** goal → watch only what you need → timestamps + quotes + confidence. |
| 18 | |
| 19 | **Is not:** a video editor. Do not cut, overlay, caption-burn, render, schedule, post, email, or write CRM from this skill. Hand cuts to Overlap, FFmpeg, or `net-new-video-editor`. Approvals stay with the calling lane. |
| 20 | |
| 21 | ## When to use |
| 22 | |
| 23 | Pre-call / sales-call mining: buyer objection, next step, competitive mention |
| 24 | Shortform scoring: find a 3-second standalone hook and in/out points |
| 25 | Longform / X research: named-person + contrast moments in podcast or YouTube tape |
| 26 | Talent review: bar evidence in a Loom or trial recording |
| 27 | Client audit: every mention of a keyword across a discovery recording |
| 28 | |
| 29 | Skip when the job is already a clean transcript and you only need text search. |
| 30 | |
| 31 | ## Inputs |
| 32 | |
| 33 | | Field | Required | Notes | |
| 34 | |-------|----------|-------| |
| 35 | | `source` | yes | URL or local media path the runtime can read | |
| 36 | | `goal` | yes | One sentence retrieval goal | |
| 37 | | `keywords` | no | Extra strings to bias retrieval | |
| 38 | | `max_moments` | no | Default 5 | |
| 39 | | `modality` | no | `auto` (default), `frames`, `audio`, or `transcript` | |
| 40 | |
| 41 | ## Process |
| 42 | |
| 43 | **Restate the goal** as 1–3 retrieval queries. Done when each query is falsifiable (you would know if a moment matched). |
| 44 | **Call Gemini agentic video understanding** (Gemini API or AI Studio) with `source`, queries, `max_moments`, and modality preference. Prefer the agentic path over fixed-FPS full ingest when available. Done when the API returns candidate windows or an explicit empty set. |
| 45 | **Normalize moments** into the output schema below. Flag paraphrase vs verbatim. Drop fabricated timestamps. Done when every kept moment has `t_start`, `t_end`, `modality`, `quote`, `why`, `confidence`. |
| 46 | **Stop and hand off** to the caller. Do not cut, overlay, schedule, publish, email, or CRM-write. |
| 47 | |
| 48 | ## Output schema |
| 49 | |
| 50 | Markdown for humans, optional JSON for machines: |
| 51 | |
| 52 | |
| 53 | { |
| 54 | "goal": "", |
| 55 | "source": "", |
| 56 | "moments": [ |
| 57 | { |
| 58 | "t_start": "MM:SS", |
| 59 | "t_end": "MM:SS", |
| 60 | "modality": "frames|audio|transcript", |
| 61 | "quote": "", |
| 62 | "verbatim": true, |
| 63 | "why": "", |
| 64 | "confidence": 0.0 |
| 65 | } |
| 66 | ], |
| 67 | "empty_reason": null, |
| 68 | "tokens_note": "agentic path used|fallback static ingest" |
| 69 | } |
| 70 | |
| 71 | |
| 72 | ## Hard gates |
| 73 | |
| 74 | No full fixed-FPS ingest when the agentic path is available |
| 75 | No invented timestamps or quotes |
| 76 | No dumping full transcripts or client PII into public artifacts |
| 77 | No cut / render / overlay / schedule / publish / send from this skill |
| 78 | |
| 79 | ## Setup |
| 80 | |
| 81 | Gemini API key or Google AI Studio access: https://ai.studio |
| 82 | See Google’s developer guide for agentic video understanding in Gemini |
| 83 | Env: `GEMINI_API_KEY` (or the project’s existing Google AI credential) |
| 84 | |
| 85 | ## Caller one-liners |
| 86 | |
| 87 | Pre-call: `goal="exact next-step commitment and any pricing pushback"` |
| 88 | Shortform: `goal="best 3-second standalone hook; return in/out for one clip"` |
| 89 | Talent: `goal="evidence they hit the role bar on X; max 5 moments"` |
| 90 | Audit: `goal="every mention of Reddit, AEO, or budget"` |
| 91 | |
| 92 | ## Completion |
| 93 | |
| 94 | Done when the caller has the schema above (or a documented empty set) and this skill has performed no side effects beyond the Gemini read. |
| 95 |
Discussion
Browse more free Claude skills.