bf63fd19309513a02d415f1d75a12169043c35f4
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
76c51f3ae8 |
code-review: fold in the classes only the natural fix history shows
I had only used one repo's per-category survival table plus the other's aggregate numbers, and had not looked at the shipping repos' fix history at all. Reading all 105 planted bugs per-bug, and four weeks of real fixes, changed four things. Taxonomy is now thirteen lenses: - Lens 10 (representation and information loss) is promoted to the highest-frequency class in every corpus and given three named sub-shapes: the nullish family (pending/absent/empty/zero/false/failed collapsing into each other), projection and field-set drift (a producer quietly stops emitting a field, consumers degrade instead of failing), and unresolved values stored as resolved ones. Plus the cast/any/suppression tell - an annotation on a boundary marks where two sides disagreed and someone silenced the compiler. - Lens 12 gains reachability: a predicate nothing can satisfy, a handler never wired, a scheduler never started. Reads as correct code; common in the wild. - Lens 13, verification and observability, is new: the check that cannot fail, the oracle measuring the wrong thing, the effect whose absence nothing would notice. It carries a note on WHY it is new - a planted defect is detectable by construction, so silent failure is systematically absent from planted corpora and heavily represented in real fix histories. A checklist trained only on planted bugs will never prompt you to look here. Refutation gains "absorption is not prevention": a cache that usually holds, a retry that usually succeeds, a default that is usually right - none of those refute a finding, they postpone it. Drop only on a mechanism that makes the execution impossible. Corollary: "works nearly always" describes a race. Parallelism gains two constraints: - One model. Fan-out is for coverage, not a second opinion; workers run the coordinator's model. A single foreign worker makes a measured result unattributable. The independent second-model pass stays where it belongs, as an explicit /ship-check step. - Read-only workers. Read, search, navigate - no writes, edits or mutating commands. A worker that can edit drifts from reviewing into silently fixing, and the tree must end identical to how it started or findings cannot be checked against it. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01G42vsxSKL7je39AsHZ5aJm |
||
|
|
18032bc9f7 |
pm-ai-shipping: code-review becomes the top-level skill; perf + security are sub-cases
Restructure, per the ask that code review be the parent and the other two dimensions its sub-cases: - SKILL.md gains a "one engine, three anchors" section. Correctness is the core and stays inline; performance and security move to their own reference files, loaded only when selected. - references/performance-review.md (new) — a universal, stack-agnostic core (repeated work, growth relationships, retention, copying, contention, amplification) plus the three-part bar for a performance finding. Defers the database/web checklist to /performance-audit-static instead of restating it. - references/security-review.md (new) — trust boundaries and sinks for code with no web surface, and the one rule that INVERTS relative to correctness: attacker-equals-victim refutes a security finding but never a correctness one. Defers the full procedure to /security-audit-static. - Both audit commands now say they are the specialisation behind their sub-case, so the narrow entry points still lead back to the skill. ship-check gains two stages it was missing: - Step 3, correctness review — the pass neither audit performs: logic and state defects that compile clean and pass the suite. - Step 6, independent unsteered review — a fresh session of a second model (Codex or equivalent), given no checklist and no prior findings, with the subject computed from a diff rather than described. Every finding is hand-verified against the code before it enters the packet, since an unsteered reviewer carries no refutation discipline of its own. The packet reports whether it ran clean or did not run at all - those are different signals. Also carries the working-tree edits already in progress: model-and-orchestration guidance on both audits, the OWASP A02/A06/A09 backstop, CSP in the output-encoding bullet, the prompt-injection/agent-abuse bullet, and the Audit Provenance section (now also naming the second model). No version bump - not tested against the benchmark yet. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01G42vsxSKL7je39AsHZ5aJm |
||
|
|
2e662ac04d |
pm-ai-shipping: add the code-review skill (correctness / performance / security, each optional)
NOT RELEASED. No version bump - versions stay at 2.0.0 across all 9 plugins and marketplace.json, because bumping is the release action and this ships only after it has been tested. On a branch for the same reason. WHY A SKILL AND NOT A FOURTH COMMAND. /security-audit-static is already mature - sink analysis, self-refutation with attacker/victim rules, OWASP backstop, fan-out. Rebuilding that inside something new would duplicate it. The hole in this plugin is CORRECTNESS: there is no bug-finding review at all. So this is one skill with three independently activated dimensions that defers to the existing command for security and points at intended-vs-implemented for the doc-vs-code axis. THE ANCHOR IS THE AGREEMENT, NOT THE FILE. The defects reviewers miss are rarely visible inside one file - they are disagreements between two participants that each read sensibly alone. Engine: map a flow, identify an obligation, inspect EVERY participant, construct a violating execution, trace the consequence, refute, report. Two lenses get a forced probe rather than a checklist mention: authority reconciliation (a requested value is not an applied value) and identity correlation (is the key unique, stable and live under overlap and reuse). Refutation discipline is deliberately stricter than the security command's: a correctness defect can harm only the person who triggered it and still be serious, so the attacker/victim test does not transfer, and 'keep unless disproved' is too permissive. Keep / Drop / Unresolved, with unresolved kept out of the findings list. Parallelism fans out over complete flows, never over files - partitioning by file is exactly the split that hides cross-boundary defects. Overlapping reads are allowed and encouraged. Coverage reports work performed in four states; zero findings is not 'not covered'. Co-designed with GPT-6 Astra (Codex CLI). Contains no project-specific content: no repo names, no paths, no bug identifiers, no defect text - verified by scan. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01G42vsxSKL7je39AsHZ5aJm |