Plugin Development Workflow skill
Guide plugin development workflow — editing skills, agents, hooks, or eval framework in this repo.
by oliver-kriska·MIT license·★ 560 Stars on the repo·GitHub ↗
Use now
npx degit oliver-kriska/claude-elixir-phoenix/.claude/skills/plugin-dev-workflow#main ~/.claude/skills/plugin-dev-workflowChecked ·commit main
Files of Plugin Development Workflow
SKILL.md
Show the full text118 lines
Plugin Development Workflow
This repo is the Elixir/Phoenix Claude Code plugin. When editing plugin files, follow this workflow to ensure quality.
Before You Start
Run make help to see all available commands:
make eval # Quick: lint + score changed skills/agents
make eval-all # Full: all 51 skills + 26 agents
make eval-fix # Auto-fix + show failures
make test # 220 pytest tests for eval framework and port tooling
make ci # Full CI pipeline
Scoring Individual Files (CLI)
IMPORTANT: Always use -m module syntax, never run scorer.py directly.
# Score ONE skill (use -m, NOT direct file path)
python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md
# Score ONE skill with pretty output
python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md --pretty
# Score all skills
python3 -m lab.eval.scorer --all
# Score ONE agent
python3 -m lab.eval.agent_scorer plugins/elixir-phoenix/agents/verification-runner.md
# Score all agents
python3 -m lab.eval.agent_scorer --all
make ci # Full CI pipeline
When Editing Skills (plugins/elixir-phoenix/skills/*/SKILL.md)
- Read CLAUDE.md conventions (size limits, frontmatter requirements)
- Make your changes
- Run
make eval— it auto-detects changed skills and scores them - If FAIL: check the dimension that failed, fix it
- Run
make lintto verify markdown formatting - Commit
Skill requirements (eval checks all of these):
- Frontmatter: name, description, effort. Description must start with action verb + include "Use when..."
- Iron Laws section with 1+ numbered items
- Under 185 lines (command skills) or 150 lines (reference skills)
- No section exceeds 45 lines
- All
/phx:references point to existing skills - All
references/*.mdpaths exist - No dangerous code patterns outside Iron Laws sections
- Code examples present (1+ fenced code blocks)
- "Use when..." in description (for trigger accuracy)
When Editing Agents (plugins/elixir-phoenix/agents/*.md)
- Make your changes
- Run
make eval-agentsto score all agents - Agent requirements:
- no
permissionMode(plugin agents ignore it; the plugin directory failsbypassPermissions) disallowedTools: Write, Edit, NotebookEditfor review/analysis agents- model matches effort: haiku=low, sonnet=medium, opus=high
- Under 300 lines (specialist) or 535 lines (orchestrator)
- no
When Editing Eval Framework (lab/eval/*.py)
- Make your changes
- Run
make test— all pytest tests must pass - Run
make eval-all— verify no skills/agents regressed - If adding new matchers: add tests in
lab/eval/tests/test_matchers.py
When Editing Hooks (plugins/elixir-phoenix/hooks/scripts/*.sh)
- Make your changes
- Run
make lint(markdown in hook comments) - Test the hook manually (hooks run on Edit/Write/Bash events)
- Check CLAUDE.md hook documentation is still accurate
Autoresearch (Self-Improvement Loop)
If make eval-fix shows failures, it suggests an autoresearch command:
# Copy-paste the suggested command from eval-fix output
claude -p 'Run autoresearch. Score all skills...' --allowedTools 'Edit,Read,Write,Bash,Glob,Grep'
This runs the autoresearch loop: find weakest skill → fix ONE issue → re-score → keep/revert.
Pre-Commit Checklist
Before committing any plugin changes:
-
make lintpasses -
make evalpasses (changed files) -
make testpasses (if eval framework changed) - CHANGELOG.md updated (if user-visible change)
- Version bumped in plugin.json (if releasing)
References
- CLAUDE.md — full conventions, size limits, checklist
lab/eval/— scoring framework (24 matchers, 8 dimensions)lab/autoresearch/— self-improvement looplab/findings/interesting.jsonl— log interesting discoveries here
| 1 | |
| 2 | name plugin-dev-workflow |
| 3 | description "Guide plugin development workflow — editing skills, agents, hooks, or eval framework in this repo. Use when modifying files in plugins/elixir-phoenix/, lab/eval/, or lab/autoresearch/. Ensures changes pass eval, lint, and tests before committing." |
| 4 | effort medium |
| 5 | |
| 6 | |
| 7 | # Plugin Development Workflow |
| 8 | |
| 9 | This repo is the Elixir/Phoenix Claude Code plugin. When editing plugin |
| 10 | files, follow this workflow to ensure quality. |
| 11 | |
| 12 | ## Before You Start |
| 13 | |
| 14 | Run `make help` to see all available commands: |
| 15 | |
| 16 | |
| 17 | make eval # Quick: lint + score changed skills/agents |
| 18 | make eval-all # Full: all 51 skills + 26 agents |
| 19 | make eval-fix # Auto-fix + show failures |
| 20 | make test # 220 pytest tests for eval framework and port tooling |
| 21 | make ci # Full CI pipeline |
| 22 | |
| 23 | |
| 24 | ## Scoring Individual Files (CLI) |
| 25 | |
| 26 | IMPORTANT: Always use `-m` module syntax, never run scorer.py directly. |
| 27 | |
| 28 | |
| 29 | # Score ONE skill (use -m, NOT direct file path) |
| 30 | python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md |
| 31 | |
| 32 | # Score ONE skill with pretty output |
| 33 | python3 -m lab.eval.scorer plugins/elixir-phoenix/skills/verify/SKILL.md --pretty |
| 34 | |
| 35 | # Score all skills |
| 36 | python3 -m lab.eval.scorer --all |
| 37 | |
| 38 | # Score ONE agent |
| 39 | python3 -m lab.eval.agent_scorer plugins/elixir-phoenix/agents/verification-runner.md |
| 40 | |
| 41 | # Score all agents |
| 42 | python3 -m lab.eval.agent_scorer --all |
| 43 | make ci # Full CI pipeline |
| 44 | |
| 45 | |
| 46 | ## When Editing Skills (plugins/elixir-phoenix/skills/*/SKILL.md) |
| 47 | |
| 48 | **Read CLAUDE.md** conventions (size limits, frontmatter requirements) |
| 49 | Make your changes |
| 50 | Run `make eval` — it auto-detects changed skills and scores them |
| 51 | If FAIL: check the dimension that failed, fix it |
| 52 | Run `make lint` to verify markdown formatting |
| 53 | Commit |
| 54 | |
| 55 | **Skill requirements** (eval checks all of these): |
| 56 | |
| 57 | Frontmatter: name, description, effort. Description must start with action verb + include "Use when..." |
| 58 | Iron Laws section with 1+ numbered items |
| 59 | Under 185 lines (command skills) or 150 lines (reference skills) |
| 60 | No section exceeds 45 lines |
| 61 | All `/phx:` references point to existing skills |
| 62 | All `references/*.md` paths exist |
| 63 | No dangerous code patterns outside Iron Laws sections |
| 64 | Code examples present (1+ fenced code blocks) |
| 65 | "Use when..." in description (for trigger accuracy) |
| 66 | |
| 67 | ## When Editing Agents (plugins/elixir-phoenix/agents/*.md) |
| 68 | |
| 69 | Make your changes |
| 70 | Run `make eval-agents` to score all agents |
| 71 | Agent requirements: |
| 72 | no `permissionMode` (plugin agents ignore it; the plugin directory fails `bypassPermissions`) |
| 73 | `disallowedTools: Write, Edit, NotebookEdit` for review/analysis agents |
| 74 | model matches effort: haiku=low, sonnet=medium, opus=high |
| 75 | Under 300 lines (specialist) or 535 lines (orchestrator) |
| 76 | |
| 77 | ## When Editing Eval Framework (lab/eval/*.py) |
| 78 | |
| 79 | Make your changes |
| 80 | Run `make test` — all pytest tests must pass |
| 81 | Run `make eval-all` — verify no skills/agents regressed |
| 82 | If adding new matchers: add tests in `lab/eval/tests/test_matchers.py` |
| 83 | |
| 84 | ## When Editing Hooks (plugins/elixir-phoenix/hooks/scripts/*.sh) |
| 85 | |
| 86 | Make your changes |
| 87 | Run `make lint` (markdown in hook comments) |
| 88 | Test the hook manually (hooks run on Edit/Write/Bash events) |
| 89 | Check CLAUDE.md hook documentation is still accurate |
| 90 | |
| 91 | ## Autoresearch (Self-Improvement Loop) |
| 92 | |
| 93 | If `make eval-fix` shows failures, it suggests an autoresearch command: |
| 94 | |
| 95 | |
| 96 | # Copy-paste the suggested command from eval-fix output |
| 97 | claude -p 'Run autoresearch. Score all skills...' --allowedTools 'Edit,Read,Write,Bash,Glob,Grep' |
| 98 | |
| 99 | |
| 100 | This runs the autoresearch loop: find weakest skill → fix ONE issue → re-score → keep/revert. |
| 101 | |
| 102 | ## Pre-Commit Checklist |
| 103 | |
| 104 | Before committing any plugin changes: |
| 105 | |
| 106 | [ ] `make lint` passes |
| 107 | [ ] `make eval` passes (changed files) |
| 108 | [ ] `make test` passes (if eval framework changed) |
| 109 | [ ] CHANGELOG.md updated (if user-visible change) |
| 110 | [ ] Version bumped in plugin.json (if releasing) |
| 111 | |
| 112 | ## References |
| 113 | |
| 114 | CLAUDE.md — full conventions, size limits, checklist |
| 115 | `lab/eval/` — scoring framework (24 matchers, 8 dimensions) |
| 116 | `lab/autoresearch/` — self-improvement loop |
| 117 | `lab/findings/interesting.jsonl` — log interesting discoveries here |
| 118 |
Discussion
Alternatives
Browse more free Claude skills or everything in Operations.