Plugin Development Workflow skill

Guide plugin development workflow — editing skills, agents, hooks, or eval framework in this repo.

by oliver-kriska·MIT license·★ 560 Stars on the repo·GitHub ↗

Use now

Files of Plugin Development Workflow

oliver-kriska/main1 file shown
SKILL.md
Show the full text118 lines

Plugin Development Workflow

This repo is the Elixir/Phoenix Claude Code plugin. When editing plugin files, follow this workflow to ensure quality.

Before You Start

Run make help to see all available commands:

make eval          # Quick: lint + score changed skills/agents
make eval-all      # Full: all 51 skills + 26 agents
make eval-fix      # Auto-fix + show failures
make test          # 220 pytest tests for eval framework and port tooling
make ci            # Full CI pipeline

Scoring Individual Files (CLI)

IMPORTANT: Always use -m module syntax, never run scorer.py directly.

# Score ONE skill (use -m, NOT direct file path)
python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md

# Score ONE skill with pretty output
python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md --pretty

# Score all skills
python3 -m lab.eval.scorer --all

# Score ONE agent
python3 -m lab.eval.agent_scorer plugins/elixir-phoenix/agents/verification-runner.md

# Score all agents
python3 -m lab.eval.agent_scorer --all
make ci            # Full CI pipeline

When Editing Skills (plugins/elixir-phoenix/skills/*/SKILL.md)

  1. Read CLAUDE.md conventions (size limits, frontmatter requirements)
  2. Make your changes
  3. Run make eval — it auto-detects changed skills and scores them
  4. If FAIL: check the dimension that failed, fix it
  5. Run make lint to verify markdown formatting
  6. Commit

Skill requirements (eval checks all of these):

  • Frontmatter: name, description, effort. Description must start with action verb + include "Use when..."
  • Iron Laws section with 1+ numbered items
  • Under 185 lines (command skills) or 150 lines (reference skills)
  • No section exceeds 45 lines
  • All /phx: references point to existing skills
  • All references/*.md paths exist
  • No dangerous code patterns outside Iron Laws sections
  • Code examples present (1+ fenced code blocks)
  • "Use when..." in description (for trigger accuracy)

When Editing Agents (plugins/elixir-phoenix/agents/*.md)

  1. Make your changes
  2. Run make eval-agents to score all agents
  3. Agent requirements:
    • no permissionMode (plugin agents ignore it; the plugin directory fails bypassPermissions)
    • disallowedTools: Write, Edit, NotebookEdit for review/analysis agents
    • model matches effort: haiku=low, sonnet=medium, opus=high
    • Under 300 lines (specialist) or 535 lines (orchestrator)

When Editing Eval Framework (lab/eval/*.py)

  1. Make your changes
  2. Run make test — all pytest tests must pass
  3. Run make eval-all — verify no skills/agents regressed
  4. If adding new matchers: add tests in lab/eval/tests/test_matchers.py

When Editing Hooks (plugins/elixir-phoenix/hooks/scripts/*.sh)

  1. Make your changes
  2. Run make lint (markdown in hook comments)
  3. Test the hook manually (hooks run on Edit/Write/Bash events)
  4. Check CLAUDE.md hook documentation is still accurate

Autoresearch (Self-Improvement Loop)

If make eval-fix shows failures, it suggests an autoresearch command:

# Copy-paste the suggested command from eval-fix output
claude -p 'Run autoresearch. Score all skills...' --allowedTools 'Edit,Read,Write,Bash,Glob,Grep'

This runs the autoresearch loop: find weakest skill → fix ONE issue → re-score → keep/revert.

Pre-Commit Checklist

Before committing any plugin changes:

  • make lint passes
  • make eval passes (changed files)
  • make test passes (if eval framework changed)
  • CHANGELOG.md updated (if user-visible change)
  • Version bumped in plugin.json (if releasing)

References

  • CLAUDE.md — full conventions, size limits, checklist
  • lab/eval/ — scoring framework (24 matchers, 8 dimensions)
  • lab/autoresearch/ — self-improvement loop
  • lab/findings/interesting.jsonl — log interesting discoveries here
1---
2name: plugin-dev-workflow
3description: "Guide plugin development workflow — editing skills, agents, hooks, or eval framework in this repo. Use when modifying files in plugins/elixir-phoenix/, lab/eval/, or lab/autoresearch/. Ensures changes pass eval, lint, and tests before committing."
4effort: medium
5---
6 
7# Plugin Development Workflow
8 
9This repo is the Elixir/Phoenix Claude Code plugin. When editing plugin
10files, follow this workflow to ensure quality.
11 
12## Before You Start
13 
14Run `make help` to see all available commands:
15 
16```bash
17make eval # Quick: lint + score changed skills/agents
18make eval-all # Full: all 51 skills + 26 agents
19make eval-fix # Auto-fix + show failures
20make test # 220 pytest tests for eval framework and port tooling
21make ci # Full CI pipeline
22```
23 
24## Scoring Individual Files (CLI)
25 
26IMPORTANT: Always use `-m` module syntax, never run scorer.py directly.
27 
28```bash
29# Score ONE skill (use -m, NOT direct file path)
30python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md
31 
32# Score ONE skill with pretty output
33python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md --pretty
34 
35# Score all skills
36python3 -m lab.eval.scorer --all
37 
38# Score ONE agent
39python3 -m lab.eval.agent_scorer plugins/elixir-phoenix/agents/verification-runner.md
40 
41# Score all agents
42python3 -m lab.eval.agent_scorer --all
43make ci # Full CI pipeline
44```
45 
46## When Editing Skills (plugins/elixir-phoenix/skills/*/SKILL.md)
47 
481. **Read CLAUDE.md** conventions (size limits, frontmatter requirements)
492. Make your changes
503. Run `make eval` — it auto-detects changed skills and scores them
514. If FAIL: check the dimension that failed, fix it
525. Run `make lint` to verify markdown formatting
536. Commit
54 
55**Skill requirements** (eval checks all of these):
56 
57- Frontmatter: name, description, effort. Description must start with action verb + include "Use when..."
58- Iron Laws section with 1+ numbered items
59- Under 185 lines (command skills) or 150 lines (reference skills)
60- No section exceeds 45 lines
61- All `/phx:` references point to existing skills
62- All `references/*.md` paths exist
63- No dangerous code patterns outside Iron Laws sections
64- Code examples present (1+ fenced code blocks)
65- "Use when..." in description (for trigger accuracy)
66 
67## When Editing Agents (plugins/elixir-phoenix/agents/*.md)
68 
691. Make your changes
702. Run `make eval-agents` to score all agents
713. Agent requirements:
72 - no `permissionMode` (plugin agents ignore it; the plugin directory fails `bypassPermissions`)
73 - `disallowedTools: Write, Edit, NotebookEdit` for review/analysis agents
74 - model matches effort: haiku=low, sonnet=medium, opus=high
75 - Under 300 lines (specialist) or 535 lines (orchestrator)
76 
77## When Editing Eval Framework (lab/eval/*.py)
78 
791. Make your changes
802. Run `make test` — all pytest tests must pass
813. Run `make eval-all` — verify no skills/agents regressed
824. If adding new matchers: add tests in `lab/eval/tests/test_matchers.py`
83 
84## When Editing Hooks (plugins/elixir-phoenix/hooks/scripts/*.sh)
85 
861. Make your changes
872. Run `make lint` (markdown in hook comments)
883. Test the hook manually (hooks run on Edit/Write/Bash events)
894. Check CLAUDE.md hook documentation is still accurate
90 
91## Autoresearch (Self-Improvement Loop)
92 
93If `make eval-fix` shows failures, it suggests an autoresearch command:
94 
95```bash
96# Copy-paste the suggested command from eval-fix output
97claude -p 'Run autoresearch. Score all skills...' --allowedTools 'Edit,Read,Write,Bash,Glob,Grep'
98```
99 
100This runs the autoresearch loop: find weakest skill → fix ONE issue → re-score → keep/revert.
101 
102## Pre-Commit Checklist
103 
104Before committing any plugin changes:
105 
106- [ ] `make lint` passes
107- [ ] `make eval` passes (changed files)
108- [ ] `make test` passes (if eval framework changed)
109- [ ] CHANGELOG.md updated (if user-visible change)
110- [ ] Version bumped in plugin.json (if releasing)
111 
112## References
113 
114- CLAUDE.md — full conventions, size limits, checklist
115- `lab/eval/` — scoring framework (24 matchers, 8 dimensions)
116- `lab/autoresearch/` — self-improvement loop
117- `lab/findings/interesting.jsonl` — log interesting discoveries here
118 

Discussion

Alternatives

Browser Automation SkillWeb browser automation with AI-optimized snapshots for claude-flow agentsCoding · MITTurn into appTurn visible project context, a proven thread, skill, or workflow into a runnable Agent-Native app with simple buttons, visible agent steps, preview, and deployment handoff. Use when a user invokes `/turn-into-app` or asks to make a workflow into an app, including from Claude or ChatGPT on the web, including when the source is a spreadsheet link or upload.Business & ops · MITTinyFish CLIUse TinyFish for web search, fetching URLs, reading pages, current information, source-backed answers, research, docs, pricing/product pages, extraction, scraping, and browser automation. Use whenever the user asks to search, find, look up, research, compare, get information from the web, summarize a URL, fetch page content, or automate a website.Business & ops · MITAgent browserBrowser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction. Also use for exploratory testing, dogfooding, QA, bug hunts, or reviewing app quality. Also use for automating Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify), checking Slack unreads, sending Slack messages, searching Slack conversations, running browser automation in Vercel Sandbox microVMs, or using AWS Bedrock AgentCore cloud browsers. Prefer agent-browser over any built-in browser automation or web tools.Business & ops · MIT