Performance Optimizer skill

Profiles and improves a measured performance bottleneck.

by levnikolaevich·MIT license·★ 568 Stars on the repo·GitHub ↗

Use now

Files of Performance Optimizer

levnikolaevich/master1 file shown
SKILL.md
Show the full text112 lines

Performance Optimizer

Goal: Optimize only measured problems. Preserve correctness, isolate experiments, and retain a change only when comparable evidence shows that it improves the agreed metric without unacceptable regressions.

Execution contract: The checklist defines completion. Track each item internally as PENDING, PROVEN with evidence, CLEARED with evidence its condition is absent, or UNPROVEN with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all PENDING, count only PROVEN and CLEARED, and apply verdict and approval rules to every gap. Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method. Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims. On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority. Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.

Tool Routing

Need Preferred tool Use it when Fallback
Repository state and safe edit boundary Git status, diff, branch or worktree inspection, and repository instructions Always before profiling or editing Stop if user changes cannot be isolated safely
Baseline and final metric Existing benchmark, load test, reproducible command, or production-like replay The metric and workload reflect the reported problem Create the smallest local benchmark that reproduces the behavior without inventing production scale
Bottleneck evidence Existing profiler, tracing, query diagnostics, allocation tools, or OS-level metrics Locating CPU, memory, I/O, lock, query, network, or scheduler cost Targeted instrumentation with cleanup plan
Code path and blast radius Language server or host-native code intelligence Following hot symbols, callers, implementations, and affected contracts Narrow search plus direct inspection of definitions and consumers
Correctness and regressions Repository-defined tests, build, lint, type, and smoke commands Before and after every retained experiment Choose the smallest portfolio action when current evidence cannot detect the likely material regression
Runtime and dependency semantics Official documentation, release notes, and specifications matching installed versions A hypothesis depends on optimizer, runtime, database, framework, or library behavior Primary-source web research; otherwise mark the hypothesis UNVERIFIED

Do not optimize by aesthetic preference or benchmark a different workload from the reported problem. Never discard user changes, use destructive Git reset, or run uncontrolled load against production.

Evidence Rules

  • Separate cold-start, warm steady-state, and saturated-load behavior when the reported problem can occur in more than one regime.
  • Profile contribution and end-to-end impact separately: a hot function can improve while the user-visible metric does not.
  • Treat profiler estimates, synthetic workloads, and production observations as different evidence classes and label them.
  • Correctness, resource safety, and operational stability are hard constraints, not secondary metrics.

Checklist

1. Define the Problem and Protect the Workspace
  • Resolve the user-visible problem, workload, environment, primary metric, overall target, minimum improvement required to keep an experiment, and hard constraints before editing.
  • Confirm a measurable performance symptom and distinguish its cause from incorrect results or missing observability. Configuration, capacity, and dependencies may be valid bottlenecks; fix them only within the approved scope.
  • Read repository instructions and inspect Git state, branches, uncommitted changes, ignored artifacts, and available isolation mechanisms.
  • Preserve user work and isolate experiments in a safe branch or worktree when changes, benchmarks, or generated artifacts could interfere.
  • Start a run-owned resource ledger with every created absolute path, worktree, process ID, cache, profile, and temporary artifact; never register pre-existing resources as cleanup targets.
  • Identify correctness, security, memory, cost, compatibility, and operational constraints that no optimization may violate.
  • Locate existing benchmarks, profiles, performance tests, production traces, service-level objectives, and known environmental variability.
2. Establish a Reproducible Baseline
  • Use the same metric type as the observed problem: latency distribution, throughput, CPU, memory, allocation, I/O, query count, lock wait, or another direct measure.
  • Make the workload representative and deterministic enough to compare, including data size, concurrency, cache state, warmup, and build mode.
  • Cover the operating points that could reverse the conclusion--at minimum the reported case plus relevant data-size or concurrency boundaries--without inventing synthetic scale.
  • Choose a bounded comparison budget sufficient to assess material noise; record raw results, an appropriate center/percentile, spread, failures, and environment. Report inconclusive measurements rather than repeating until a gain appears.
  • When drift or noise is material, interleave or randomize baseline and candidate runs and prefer paired comparisons over one block of "before" followed by one block of "after."
  • Verify that the benchmark detects an intentionally slower or obviously changed path when practical; a benchmark insensitive to behavior cannot validate optimization.
  • Run relevant correctness tests before editing so pre-existing failures are not attributed to experiments.
  • Stop and report BLOCKED if the problem cannot be reproduced and no trustworthy production evidence can define a safe proxy.
3. Profile and Form Hypotheses
  • Profile the end-to-end path before focusing on a function, query, allocation, lock, or network call.
  • Build a ranked cost map with measured contribution, call frequency, inclusive and exclusive cost where available, and affected workload.
  • Trace the top costs to implementation, callers, data shape, concurrency model, configuration, and external dependencies.
  • Distinguish root bottlenecks from downstream symptoms, measurement overhead, debug builds, cold starts, and one-time initialization.
  • If profiling crosses services or processes whose code is in scope, align traces/correlation IDs and follow the measured downstream path; do not label an accessible internal service "external" and stop at its latency.
  • Estimate profiler or instrumentation perturbation and confirm the final end-to-end result without invasive instrumentation.
  • Research official runtime, framework, database, and dependency behavior only when it can confirm or reject a concrete hypothesis.
  • Check existing platform and dependency capabilities before proposing custom caches, pools, schedulers, serializers, or data structures.
  • Write a small ordered hypothesis set; for each state expected metric change, mechanism, affected files, risk, dependencies, and verification.
  • Reject hypotheses that lack a measurable mechanism, require speculative scale, or cannot be rolled back independently.
4. Execute Atomic Keep-or-Discard Experiments
  • Test value and boundary: Require every test to detect a concrete defect in this product's business logic and name the protected business outcome. Prefer E2E through user or external-system boundaries; use integration or unit tests only for business scenarios difficult to exercise reliably through E2E. Reject platform, trivial-wiring, implementation-detail, and duplicate proof with no distinct business failure signal.
  • Map each risky hypothesis to existing proof and the material regression it could cause; implement KEEP, ADD, UPDATE, MERGE, DELETE, or justified NO_TEST within the approved test scope to produce the smallest trustworthy safety evidence, remove superseded testware, and retire temporary characterization proof when its trigger ends.
  • For caching, batching, parallelism, pooling, or retry changes, explicitly protect invalidation, ordering, idempotency, cancellation, backpressure, timeout, and bounded-resource semantics that the faster path could violate.
  • Apply the smallest coherent change that tests one mechanism; group changes only when their effects are intentionally inseparable.
  • Keep instrumentation bounded, low-overhead, and easy to remove; never leave secrets or sensitive payloads in diagnostic output.
  • Run focused correctness checks after the edit. Attribute failures to the experiment, baseline, or environment; repair bounded experiment defects and recheck, or discard when correctness cannot be established.
  • Repeat the exact baseline benchmark under comparable conditions and preserve raw results.
  • Inspect the diff for accidental cleanup, unrelated refactoring, generated churn, debug flags, changed benchmark inputs, and hidden configuration changes.
  • Mark KEEP only when the experiment meets the predeclared minimum improvement beyond noise and every hard constraint passes.
  • Mark DISCARD and revert only that experiment when the keep threshold is missed, results regress, or safety becomes uncertain; never lower the threshold after observing results.
  • After a kept change, establish the new compound baseline before testing the next hypothesis.
5. Stop, Verify, and Report
  • Continue only when new measurement supports another hypothesis; stop at the agreed target, diminishing returns, exhausted safe options, or a missing prerequisite; report explicitly whether the target was reached.
  • Confirm that build, lint, type, test, smoke, benchmark, and operational evidence covers the final retained state and all required gates. Reuse passing evidence for that state; rerun checks only where later changes or unresolved failures invalidate it.
  • Run-owned cleanup: Remove only run-owned ledger entries: verify absolute paths remain inside approved temporary roots, stop exact recorded process IDs, preserve dirty or pre-existing worktrees, and retain evidence artifacts intentionally reported.
  • Confirm that the benchmark definition and acceptance threshold did not drift during the run.
  • Reconcile the hypothesis ledger with retained edits and raw results, including discarded experiments.
  • Bind measured improvement to the workload, environment, baseline and retained code/configuration; distinguish benchmark improvement from proven production impact.
  • Use DELIVERED only when retained improvements reach the agreed overall target with every constraint passing; use PARTIAL when a verified improvement is kept but the overall target remains unmet. Use NO_CHANGE when all experiments are discarded and the baseline is restored; use BLOCKED when a safety prerequisite, reproducible baseline, or safe restoration path is unavailable.

Self-Check

  • Reconcile before returning. Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.

Output Contract

Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:

  1. Result: The exact skill-specific verdict token first, then the supported outcome.
  2. Scope: Reviewed/changed scope, exclusions, baseline, and material assumptions.
  3. Evidence: Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
  4. Verification: Checks/results, unavailable evidence, and applicable cleanup/external state.
  5. Completion: Checklist: X/Y complete; Incomplete: None or each UNPROVEN item's reason, outcome impact, and exact next action; residual risks and required decisions.

Skill-specific evidence: Target workload, metric, acceptance threshold, correctness constraints, environment, sampling and variance method; comparable baseline/final distributions and deltas. Record every hypothesis, mechanism, KEEP / DISCARD, measurements, and verification. Include affected test portfolio actions, residual bottlenecks, and run-owned raw samples, commands, configuration, final diff, and cleanup evidence with paths/hashes when available.

1---
2name: ln-44-performance-optimizer
3description: "Profiles and improves a measured performance bottleneck; retains only verified improvements."
4---
5 
6# Performance Optimizer
7 
8**Goal:** Optimize only measured problems. Preserve correctness, isolate experiments, and retain a change only when comparable evidence shows that it improves the agreed metric without unacceptable regressions.
9 
10**Execution contract:** The checklist defines completion. Track each item internally as `PENDING`, `PROVEN` with evidence, `CLEARED` with evidence its condition is absent, or `UNPROVEN` with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all `PENDING`, count only `PROVEN` and `CLEARED`, and apply verdict and approval rules to every gap.
11Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method.
12Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims.
13On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority.
14Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.
15 
16 
17## Tool Routing
18 
19| Need | Preferred tool | Use it when | Fallback |
20|---|---|---|---|
21| Repository state and safe edit boundary | Git status, diff, branch or worktree inspection, and repository instructions | Always before profiling or editing | Stop if user changes cannot be isolated safely |
22| Baseline and final metric | Existing benchmark, load test, reproducible command, or production-like replay | The metric and workload reflect the reported problem | Create the smallest local benchmark that reproduces the behavior without inventing production scale |
23| Bottleneck evidence | Existing profiler, tracing, query diagnostics, allocation tools, or OS-level metrics | Locating CPU, memory, I/O, lock, query, network, or scheduler cost | Targeted instrumentation with cleanup plan |
24| Code path and blast radius | Language server or host-native code intelligence | Following hot symbols, callers, implementations, and affected contracts | Narrow search plus direct inspection of definitions and consumers |
25| Correctness and regressions | Repository-defined tests, build, lint, type, and smoke commands | Before and after every retained experiment | Choose the smallest portfolio action when current evidence cannot detect the likely material regression |
26| Runtime and dependency semantics | Official documentation, release notes, and specifications matching installed versions | A hypothesis depends on optimizer, runtime, database, framework, or library behavior | Primary-source web research; otherwise mark the hypothesis `UNVERIFIED` |
27 
28Do not optimize by aesthetic preference or benchmark a different workload from the reported problem. Never discard user changes, use destructive Git reset, or run uncontrolled load against production.
29 
30## Evidence Rules
31 
32- Separate cold-start, warm steady-state, and saturated-load behavior when the reported problem can occur in more than one regime.
33- Profile contribution and end-to-end impact separately: a hot function can improve while the user-visible metric does not.
34- Treat profiler estimates, synthetic workloads, and production observations as different evidence classes and label them.
35- Correctness, resource safety, and operational stability are hard constraints, not secondary metrics.
36 
37## Checklist
38 
39### 1. Define the Problem and Protect the Workspace
40 
41- [ ] Resolve the user-visible problem, workload, environment, primary metric, overall target, minimum improvement required to keep an experiment, and hard constraints before editing.
42- [ ] Confirm a measurable performance symptom and distinguish its cause from incorrect results or missing observability. Configuration, capacity, and dependencies may be valid bottlenecks; fix them only within the approved scope.
43- [ ] Read repository instructions and inspect Git state, branches, uncommitted changes, ignored artifacts, and available isolation mechanisms.
44- [ ] Preserve user work and isolate experiments in a safe branch or worktree when changes, benchmarks, or generated artifacts could interfere.
45- [ ] Start a run-owned resource ledger with every created absolute path, worktree, process ID, cache, profile, and temporary artifact; never register pre-existing resources as cleanup targets.
46- [ ] Identify correctness, security, memory, cost, compatibility, and operational constraints that no optimization may violate.
47- [ ] Locate existing benchmarks, profiles, performance tests, production traces, service-level objectives, and known environmental variability.
48 
49### 2. Establish a Reproducible Baseline
50 
51- [ ] Use the same metric type as the observed problem: latency distribution, throughput, CPU, memory, allocation, I/O, query count, lock wait, or another direct measure.
52- [ ] Make the workload representative and deterministic enough to compare, including data size, concurrency, cache state, warmup, and build mode.
53- [ ] Cover the operating points that could reverse the conclusion--at minimum the reported case plus relevant data-size or concurrency boundaries--without inventing synthetic scale.
54- [ ] Choose a bounded comparison budget sufficient to assess material noise; record raw results, an appropriate center/percentile, spread, failures, and environment. Report inconclusive measurements rather than repeating until a gain appears.
55- [ ] When drift or noise is material, interleave or randomize baseline and candidate runs and prefer paired comparisons over one block of "before" followed by one block of "after."
56- [ ] Verify that the benchmark detects an intentionally slower or obviously changed path when practical; a benchmark insensitive to behavior cannot validate optimization.
57- [ ] Run relevant correctness tests before editing so pre-existing failures are not attributed to experiments.
58- [ ] Stop and report `BLOCKED` if the problem cannot be reproduced and no trustworthy production evidence can define a safe proxy.
59 
60### 3. Profile and Form Hypotheses
61 
62- [ ] Profile the end-to-end path before focusing on a function, query, allocation, lock, or network call.
63- [ ] Build a ranked cost map with measured contribution, call frequency, inclusive and exclusive cost where available, and affected workload.
64- [ ] Trace the top costs to implementation, callers, data shape, concurrency model, configuration, and external dependencies.
65- [ ] Distinguish root bottlenecks from downstream symptoms, measurement overhead, debug builds, cold starts, and one-time initialization.
66- [ ] If profiling crosses services or processes whose code is in scope, align traces/correlation IDs and follow the measured downstream path; do not label an accessible internal service "external" and stop at its latency.
67- [ ] Estimate profiler or instrumentation perturbation and confirm the final end-to-end result without invasive instrumentation.
68- [ ] Research official runtime, framework, database, and dependency behavior only when it can confirm or reject a concrete hypothesis.
69- [ ] Check existing platform and dependency capabilities before proposing custom caches, pools, schedulers, serializers, or data structures.
70- [ ] Write a small ordered hypothesis set; for each state expected metric change, mechanism, affected files, risk, dependencies, and verification.
71- [ ] Reject hypotheses that lack a measurable mechanism, require speculative scale, or cannot be rolled back independently.
72 
73### 4. Execute Atomic Keep-or-Discard Experiments
74 
75- [ ] **Test value and boundary:** Require every test to detect a concrete defect in this product's business logic and name the protected business outcome. Prefer E2E through user or external-system boundaries; use integration or unit tests only for business scenarios difficult to exercise reliably through E2E. Reject platform, trivial-wiring, implementation-detail, and duplicate proof with no distinct business failure signal.
76- [ ] Map each risky hypothesis to existing proof and the material regression it could cause; implement `KEEP`, `ADD`, `UPDATE`, `MERGE`, `DELETE`, or justified `NO_TEST` within the approved test scope to produce the smallest trustworthy safety evidence, remove superseded testware, and retire temporary characterization proof when its trigger ends.
77- [ ] For caching, batching, parallelism, pooling, or retry changes, explicitly protect invalidation, ordering, idempotency, cancellation, backpressure, timeout, and bounded-resource semantics that the faster path could violate.
78- [ ] Apply the smallest coherent change that tests one mechanism; group changes only when their effects are intentionally inseparable.
79- [ ] Keep instrumentation bounded, low-overhead, and easy to remove; never leave secrets or sensitive payloads in diagnostic output.
80- [ ] Run focused correctness checks after the edit. Attribute failures to the experiment, baseline, or environment; repair bounded experiment defects and recheck, or discard when correctness cannot be established.
81- [ ] Repeat the exact baseline benchmark under comparable conditions and preserve raw results.
82- [ ] Inspect the diff for accidental cleanup, unrelated refactoring, generated churn, debug flags, changed benchmark inputs, and hidden configuration changes.
83- [ ] Mark `KEEP` only when the experiment meets the predeclared minimum improvement beyond noise and every hard constraint passes.
84- [ ] Mark `DISCARD` and revert only that experiment when the keep threshold is missed, results regress, or safety becomes uncertain; never lower the threshold after observing results.
85- [ ] After a kept change, establish the new compound baseline before testing the next hypothesis.
86 
87### 5. Stop, Verify, and Report
88 
89- [ ] Continue only when new measurement supports another hypothesis; stop at the agreed target, diminishing returns, exhausted safe options, or a missing prerequisite; report explicitly whether the target was reached.
90- [ ] Confirm that build, lint, type, test, smoke, benchmark, and operational evidence covers the final retained state and all required gates. Reuse passing evidence for that state; rerun checks only where later changes or unresolved failures invalidate it.
91- [ ] **Run-owned cleanup:** Remove only run-owned ledger entries: verify absolute paths remain inside approved temporary roots, stop exact recorded process IDs, preserve dirty or pre-existing worktrees, and retain evidence artifacts intentionally reported.
92- [ ] Confirm that the benchmark definition and acceptance threshold did not drift during the run.
93- [ ] Reconcile the hypothesis ledger with retained edits and raw results, including discarded experiments.
94- [ ] Bind measured improvement to the workload, environment, baseline and retained code/configuration; distinguish benchmark improvement from proven production impact.
95- [ ] Use `DELIVERED` only when retained improvements reach the agreed overall target with every constraint passing; use `PARTIAL` when a verified improvement is kept but the overall target remains unmet. Use `NO_CHANGE` when all experiments are discarded and the baseline is restored; use `BLOCKED` when a safety prerequisite, reproducible baseline, or safe restoration path is unavailable.
96 
97## Self-Check
98 
99- [ ] **Reconcile before returning.** Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.
100 
101## Output Contract
102 
103Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:
104 
1051. **Result:** The exact skill-specific verdict token first, then the supported outcome.
1062. **Scope:** Reviewed/changed scope, exclusions, baseline, and material assumptions.
1073. **Evidence:** Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
1084. **Verification:** Checks/results, unavailable evidence, and applicable cleanup/external state.
1095. **Completion:** `Checklist: X/Y complete`; `Incomplete: None` or each `UNPROVEN` item's reason, outcome impact, and exact next action; residual risks and required decisions.
110 
111**Skill-specific evidence:** Target workload, metric, acceptance threshold, correctness constraints, environment, sampling and variance method; comparable baseline/final distributions and deltas. Record every hypothesis, mechanism, `KEEP / DISCARD`, measurements, and verification. Include affected test portfolio actions, residual bottlenecks, and run-owned raw samples, commands, configuration, final diff, and cleanup evidence with paths/hashes when available.
112 

Discussion

Alternatives

Skill CreatorCreate new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.Coding · Apache-2.0Professional Full-Stack Developer for Network Mapping & Monitoring ApplicationAct as a professional full-stack developer tasked with building a web application for mapping and monitoring networks using Mikrotik Netwatch API. Implement multi-user role-based management to handle devices, monitor their status, and manage user subscriptions.Coding · CC0-1.0Prompt refinerHigh-end Prompt Engineering & Prompt Refiner skill. Transforms raw or messy user requests into concise, token-efficient, high-performance master prompts for systems like GPT, Claude, and Gemini. Use when you want to optimize or redesign a prompt so it solves the problem reliably while minimizing tokens.Data & AI · CC0-1.0Constraint driven developmentEstablishes a project's quality bar as a written contract and stops agents quietly lowering it. Interviews the user on which dimensions matter, supplies sane default thresholds when they have no number in mind, records everything in CONSTRAINTS.md, and watches the diff for a weakened bar — new @ts-ignore or eslint-disable suppressions, skipped or deleted tests, assertions stripped out, unimplemented stubs, thresholds edited down. Use when no quality bar is written down, when the user says "set up constraints" or "define our standards", when the user wants dimensions they care about — accessibility, web performance, coverage — set up as enforced constraints, when an agent keeps silencing checks or skipping tests to get to green, when you need a coverage or performance threshold and don't know what number to pick, or when an agent writes more code than anyone will read.Coding · MIT