pm-ai-shipping: code-review becomes the top-level skill; perf + security are sub-cases
Restructure, per the ask that code review be the parent and the other two dimensions its sub-cases: - SKILL.md gains a "one engine, three anchors" section. Correctness is the core and stays inline; performance and security move to their own reference files, loaded only when selected. - references/performance-review.md (new) — a universal, stack-agnostic core (repeated work, growth relationships, retention, copying, contention, amplification) plus the three-part bar for a performance finding. Defers the database/web checklist to /performance-audit-static instead of restating it. - references/security-review.md (new) — trust boundaries and sinks for code with no web surface, and the one rule that INVERTS relative to correctness: attacker-equals-victim refutes a security finding but never a correctness one. Defers the full procedure to /security-audit-static. - Both audit commands now say they are the specialisation behind their sub-case, so the narrow entry points still lead back to the skill. ship-check gains two stages it was missing: - Step 3, correctness review — the pass neither audit performs: logic and state defects that compile clean and pass the suite. - Step 6, independent unsteered review — a fresh session of a second model (Codex or equivalent), given no checklist and no prior findings, with the subject computed from a diff rather than described. Every finding is hand-verified against the code before it enters the packet, since an unsteered reviewer carries no refutation discipline of its own. The packet reports whether it ran clean or did not run at all - those are different signals. Also carries the working-tree edits already in progress: model-and-orchestration guidance on both audits, the OWASP A02/A06/A09 backstop, CSP in the output-encoding bullet, the prompt-injection/agent-abuse bullet, and the Audit Provenance section (now also naming the second model). No version bump - not tested against the benchmark yet. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01G42vsxSKL7je39AsHZ5aJm
This commit is contained in:
co-authored by
Claude Opus 5
parent
2e662ac04d
commit
18032bc9f7
@@ -20,6 +20,12 @@ This is a static review of code and queries, not a load test.
|
||||
|
||||
Audit **$ARGUMENTS**. If empty, review the whole repository, prioritizing list and dashboard views, frequently hit endpoints, and large tables.
|
||||
|
||||
## Model and orchestration
|
||||
|
||||
- **Run every subagent on the strongest model available** — Fable or Mythos when you have access, otherwise Opus 4.8. Match the **effort level of the current session** when the surface exposes it.
|
||||
- **Flat fan-out for large scopes.** For a big repo, fan out with parallel subagents — one per view/route/table cluster running the three checks below — then rank the merged findings yourself. One level is the target; nest a second only when a cluster is too big for one agent's context. Don't reach for a self-generating workflow.
|
||||
- **Reroutes are unlikely here, but report them if they happen.** Unlike the security audit, performance work rarely trips Fable's safety classifiers. If a cluster does get rerouted to Opus 4.8, note it in the report so the reader knows the model mix.
|
||||
|
||||
## The audit
|
||||
|
||||
### 1. Over-fetch in view payloads
|
||||
@@ -56,4 +62,5 @@ End with what's already efficient (say it explicitly) and what needs runtime pro
|
||||
- Rank by impact-per-effort — one missing index on a hot table usually beats ten micro-optimizations.
|
||||
- Don't flag theoretical inefficiency with no growth path; flag what breaks as rows or traffic scale.
|
||||
- This command covers performance only. For authorization, injection, and data-exposure risks, use `/security-audit-static`.
|
||||
- This is the data-backed-application specialisation of the **code-review** skill's performance sub-case. For logic and state defects, or for a review across several dimensions at once, use `/pm-ai-shipping:code-review`.
|
||||
- For an end-to-end pass with documentation and a shipping packet, use `/ship-check`.
|
||||
|
||||
@@ -23,6 +23,12 @@ This is a review, not a guarantee: it produces code-review findings, not confirm
|
||||
|
||||
Audit **$ARGUMENTS**. If empty, audit the whole repository, prioritizing request handlers, auth, data access, background jobs, and anything that renders, fetches, executes, logs, or stores user-controlled data. For non-trivial scopes, fan out with parallel subagents — one per function/module cluster, each running the mapping and inspection (steps 1–3); then merge candidates and run the self-refute (step 4) yourself over the full set.
|
||||
|
||||
## Model and orchestration
|
||||
|
||||
- **Run every subagent on the strongest model available** — Fable or Mythos when you have access, otherwise Opus 4.8. Match the **effort level of the current session** when the surface exposes it. This is recall-first work: a missed cross-file flow is the costly failure, so don't let a cluster silently drop to a cheaper model or a lower effort.
|
||||
- **Expect reroutes, and report them.** A security audit is exactly the content Fable's safety classifiers screen for, so some subagents will be **automatically rerouted to Opus 4.8**. That is fine for this work — but say so. Note in the report which clusters ran on the fallback model, so the reader knows the audit's model mix instead of assuming one model saw everything.
|
||||
- **Flat fan-out, not a workflow.** One level of parallel subagents (parent → cluster auditors → merge) is the target. Nest a second level **only** when a single cluster is too big for one agent's context. A deep org chart or a self-generating workflow adds coordination cost without improving recall here.
|
||||
|
||||
## The audit (small engine, strong constraint)
|
||||
|
||||
### 1. Map entry points to trust boundaries and sinks
|
||||
@@ -45,7 +51,9 @@ For each finding, try to disprove it. Default to **keep** unless you find cited
|
||||
|
||||
Name the **attacker** and the **victim**: refute if the only victim is the attacker on their own machine/account/tenant/data and no shared system or privilege boundary is crossed; keep if the impact reaches other users, tenants, shared infrastructure, billing, email reputation, secrets, or compliance-sensitive data. **Never apply attacker-equals-victim refutation to SSRF/outbound-network sinks, shared billing or quota sinks, data-exposure findings, cross-tenant or cross-principal flows, or server-side execution/rendering** — those harm someone other than the attacker by definition. Never refute a finding merely because the code is pre-existing — pre-existing bugs are the point. Do not speculate.
|
||||
|
||||
### 5. Report only what survives
|
||||
### 5. Report only what survives — with an OWASP Top 10 backstop
|
||||
|
||||
Before writing the report, map every surviving finding to its OWASP Top 10 category, and flag any category with **zero** findings as an explicit "not covered — double-check" line. This catches the classes this engine underweights: **A02 cryptographic failures** (plaintext or weakly-hashed credentials, tokens, or PII at rest; predictable tokens; missing encryption on sensitive columns), **A06 vulnerable and outdated components** (a dependency with a *reachable* exploit path — not version-drift noise), and **A09 logging and monitoring failures** (auth failures, access-control denials, and privileged actions that leave no trace for detection). The backstop is a coverage check, not a mandate to invent findings — an honest "no evidence found in A02" is a valid result.
|
||||
|
||||
## High-miss checklist (technology-shaped, not stack-specific)
|
||||
|
||||
@@ -55,7 +63,8 @@ Apply these — they're where AI-built apps most often fail:
|
||||
- **Auth-provider drift** — claims from an external identity provider (e.g. Clerk) trusted without verifying how they map to data scope.
|
||||
- **Gate/action field mismatch** — permission checked on one ID, action performed on an independent ID never proven to belong to it.
|
||||
- **Forgeable request signals** — endpoints gated by `?source=cron`, `?bot=1`, guessable headers, or unsigned webhook-like payloads instead of real auth. Raise severity when the endpoint mutates data, sends email, or triggers paid usage.
|
||||
- **Output encoding vs. input validation** — user data interpolated into HTML, `<title>`, attributes, JSON-LD, SQL, or Markdown must be encoded for *that* sink; input validation doesn't count. (XSS, CSP gaps.)
|
||||
- **Output encoding vs. input validation, and CSP** — user data interpolated into HTML, `<title>`, attributes, JSON-LD, SQL, or Markdown must be encoded for *that* sink; input validation doesn't count. Check the Content-Security-Policy itself: weak or missing directives, `unsafe-inline`, wildcard sources, inline event handlers — recommend a stricter policy that still supports app features. (XSS, CSP.)
|
||||
- **Prompt injection and agent abuse (AI apps)** — treat the model as both a sink and a source. Untrusted content (fetched pages, uploaded files, DB rows, tool output) reaching an LLM prompt; attacker text driving a privileged tool call or agent action (confused deputy); system-prompt or secret exfiltration; and unvalidated LLM *output* flowing into a downstream sink (SQL, shell, HTML, a follow-on tool call).
|
||||
- **SSRF / renderer abuse** — attacker-influenced URLs, HTML, SVG, or Markdown reaching an outbound fetch or a renderer (headless browser, PDF/OG-image generator).
|
||||
- **Parser / validator differentials** — the validator accepts a value the consumer interprets differently: unanchored regex, `startsWith`/substring allowlists, URL-parser disagreement, encoding/case/slash/path-normalization mismatch, or validation on one representation and execution on another.
|
||||
- **Fail-open paths** — error, `catch`, timeout, cancellation, cache-miss, stale-cache, feature-flag, or boundary-value branches that default to *allow*. AI code loves a permissive fallback.
|
||||
@@ -83,4 +92,5 @@ End with: the root-cause theme across findings; **what is well-built — say it
|
||||
|
||||
- Don't report generic hardening with no concrete impact, outdated deps without a reachable path, or test/mock code unless it ships. Logic and authorization bugs with no classic sink still count.
|
||||
- This command covers security only. For over-fetching, indexes, and caching, use `/performance-audit-static`.
|
||||
- This is the specialised procedure behind the **code-review** skill's security sub-case. For logic and state defects, or for a review across several dimensions at once, use `/pm-ai-shipping:code-review`.
|
||||
- For an end-to-end pass that documents first and produces a shipping packet, use `/ship-check`.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
description: Turn a vibe-coded repo into a reviewer-ready shipping packet — document the app, wire agent context, run security and performance audits, map test coverage, and compile the results
|
||||
description: Turn a vibe-coded repo into a reviewer-ready shipping packet — document the app, wire agent context, run correctness, security and performance reviews, add an independent unsteered pass, map test coverage, and compile the results
|
||||
argument-hint: "<repo path or area; defaults to the whole repository>"
|
||||
---
|
||||
|
||||
@@ -29,19 +29,39 @@ Ensure the system docs exist and are current (run `/document-app` if they're mis
|
||||
|
||||
Create or refresh `CLAUDE.md` (and a thin `AGENTS.md` pointing to it) **derived from** the system docs — the operating instructions the next AI coding agent inherits: what the system is, the trust boundaries, what may and may not be touched, where the guardrails are. This is a different artifact from the system docs: instructions, not description.
|
||||
|
||||
### Step 3: Security audit
|
||||
### Step 3: Correctness review
|
||||
|
||||
Run the security pass (`/security-audit-static`), applying the **intended-vs-implemented** skill to flag where the code diverges from `permissions.md`, `flows.md`, and `architecture.md`. Summarize surviving findings.
|
||||
Apply the **code-review** skill with `dimensions=correctness`. This is the pass the other two audits do not perform: logic and state defects that compile clean, pass the suite, and violate an agreement between two places that each look reasonable alone. Run its forced probes rather than reading through — authority reconciliation (a *requested* value still driving state where the authority returned something different) and identity correlation (results joined to their originating entity by an unstable key) are the classes strong agents miss most, and they are missed at the *look*, not at the fix.
|
||||
|
||||
### Step 4: Performance audit
|
||||
Fan out over flows, never over files. Summarize surviving findings.
|
||||
|
||||
### Step 4: Security audit
|
||||
|
||||
Run the security pass (`/security-audit-static`), applying the **intended-vs-implemented** skill to flag where the code diverges from `permissions.md`, `flows.md`, and `architecture.md`. Summarize surviving findings, and **carry through the model mix it reports** — which clusters ran on the strongest model and which were rerouted to the fallback (Opus 4.8) by Fable's classifiers.
|
||||
|
||||
### Step 5: Performance audit
|
||||
|
||||
Run the performance pass (`/performance-audit-static`) — over-fetching, missing indexes, caching. Summarize findings.
|
||||
|
||||
### Step 5: Derive the test-coverage map
|
||||
### Step 6: Independent unsteered review
|
||||
|
||||
Run `/derive-tests` to turn the documented rules — and the gaps the audits just surfaced — into a coverage map (`tests.md`): which rules are pinned by tests that exist *today*, which are only proposed, which are guarded-live or manual, and which have no verification at all. Running this **after** the audits is deliberate: each confirmed finding becomes a concrete regression test to pin, so the same gap can't silently reopen on the next AI edit. This is the operational form of "documented == implemented," and the unverified boundary rules feed straight into the launch-blocker assessment below.
|
||||
Everything above is *steered*: each pass looks for the classes its own checklist names, which is exactly why each pass is blind in the same places twice. This step is the backstop, and on a real release it is the highest-yield step in this sequence.
|
||||
|
||||
### Step 6: Compile the shipping packet
|
||||
Hand the subject to a **fresh session of a different model** — Codex (`codex exec`) is the usual choice, but any capable second model works — under three rules:
|
||||
|
||||
1. **Fresh, never a resume.** Not the thread that wrote the code, and not one that has seen the earlier findings. A session that already argued the code is correct will argue it again.
|
||||
2. **No checklist and no pointer to prior findings.** The value here is what an unprimed reader notices. Giving it the audit output converts an independent sample into a confirmation pass.
|
||||
3. **Define the subject mechanically, not in prose.** Diff against the last release tag or the deployed branch, plus the working tree — e.g. `git log --oneline <last-tag>..HEAD` and `git status`. A described subject drifts; a computed one does not.
|
||||
|
||||
**Verify every finding against the code by hand before it enters the packet.** An unsteered reviewer has no refutation discipline imposed on it, so it will produce confident findings that the code already prevents. Apply the **code-review** skill's keep/drop rule to each one: a finding survives only with a supported obligation, a feasible execution, a concrete contradiction, an observable consequence, and a counterargument you actually checked.
|
||||
|
||||
Distinguish defects the change **introduced** from defects it merely **revealed** — both belong in the packet, but only the first blocks the change itself. On a release pass, repeat the loop until a round surfaces no introduced findings above Low.
|
||||
|
||||
### Step 7: Derive the test-coverage map
|
||||
|
||||
Run `/derive-tests` to turn the documented rules — and the gaps the reviews just surfaced — into a coverage map (`tests.md`): which rules are pinned by tests that exist *today*, which are only proposed, which are guarded-live or manual, and which have no verification at all. Running this **after** the reviews is deliberate: each confirmed finding becomes a concrete regression test to pin, so the same gap can't silently reopen on the next AI edit. This is the operational form of "documented == implemented," and the unverified boundary rules feed straight into the launch-blocker assessment below.
|
||||
|
||||
### Step 8: Compile the shipping packet
|
||||
|
||||
```
|
||||
## Shipping Packet: [repo / area]
|
||||
@@ -55,12 +75,21 @@ CLAUDE.md / AGENTS.md: [created / updated / already current]
|
||||
### Test Coverage
|
||||
[Rules pinned by tests that exist today · proposed but not yet written · guarded-live/manual · and the documented rules nothing verifies yet]
|
||||
|
||||
### Correctness Summary
|
||||
[Surviving findings, each: Expectation · Trigger · Defect · Impact · Remedy, citing every participant]
|
||||
|
||||
### Security Summary
|
||||
[Counts by severity + the surviving findings, each: Risk · Attack · Impact · Fix]
|
||||
|
||||
### Performance Summary
|
||||
[Findings by view/route/table, each: Recommendation · Effort · Priority]
|
||||
|
||||
### Independent Review
|
||||
[Which model and session ran it, how the subject was computed, how many findings it returned, how many survived hand-verification — and the ones that survived. Note whether the last round was clean.]
|
||||
|
||||
### Audit Provenance
|
||||
[Which model each audit actually ran on, any clusters Fable's classifiers rerouted to the fallback (Opus 4.8), and the second model used in Step 6 — so the reviewer knows how much of the work saw the strongest model vs. the fallback, and that at least one pass was genuinely independent]
|
||||
|
||||
### Launch Blockers
|
||||
[Unresolved Critical/High items — including any boundary rule that is both unverified and unaudited — that should stop a ship]
|
||||
|
||||
@@ -73,4 +102,5 @@ CLAUDE.md / AGENTS.md: [created / updated / already current]
|
||||
- This is a handoff compiler: the value is sequencing plus synthesis, not re-deriving each audit.
|
||||
- If documentation is missing, the packet says so loudly — an audit without documented intent is incomplete, and the inventory makes that visible rather than hiding it.
|
||||
- Findings are code-review results, not confirmed exploits; the packet is a basis for human sign-off, not a substitute for it.
|
||||
- Run the specialist commands directly (`/document-app`, `/derive-tests`, `/security-audit-static`, `/performance-audit-static`) when you only need one stage.
|
||||
- Step 6 is skippable only when no second model is available — say so in the packet rather than omitting the section, because "not run" and "run clean" are very different signals to a reviewer.
|
||||
- Run the specialist commands directly (`/document-app`, `/derive-tests`, `/pm-ai-shipping:code-review`, `/security-audit-static`, `/performance-audit-static`) when you only need one stage.
|
||||
|
||||
Reference in New Issue
Block a user