code-review: fold in the classes only the natural fix history shows

I had only used one repo's per-category survival table plus the other's
aggregate numbers, and had not looked at the shipping repos' fix history at
all. Reading all 105 planted bugs per-bug, and four weeks of real fixes,
changed four things.

Taxonomy is now thirteen lenses:

- Lens 10 (representation and information loss) is promoted to the
  highest-frequency class in every corpus and given three named sub-shapes:
  the nullish family (pending/absent/empty/zero/false/failed collapsing into
  each other), projection and field-set drift (a producer quietly stops
  emitting a field, consumers degrade instead of failing), and unresolved
  values stored as resolved ones. Plus the cast/any/suppression tell - an
  annotation on a boundary marks where two sides disagreed and someone
  silenced the compiler.
- Lens 12 gains reachability: a predicate nothing can satisfy, a handler never
  wired, a scheduler never started. Reads as correct code; common in the wild.
- Lens 13, verification and observability, is new: the check that cannot fail,
  the oracle measuring the wrong thing, the effect whose absence nothing would
  notice. It carries a note on WHY it is new - a planted defect is detectable
  by construction, so silent failure is systematically absent from planted
  corpora and heavily represented in real fix histories. A checklist trained
  only on planted bugs will never prompt you to look here.

Refutation gains "absorption is not prevention": a cache that usually holds, a
retry that usually succeeds, a default that is usually right - none of those
refute a finding, they postpone it. Drop only on a mechanism that makes the
execution impossible. Corollary: "works nearly always" describes a race.

Parallelism gains two constraints:

- One model. Fan-out is for coverage, not a second opinion; workers run the
  coordinator's model. A single foreign worker makes a measured result
  unattributable. The independent second-model pass stays where it belongs,
  as an explicit /ship-check step.
- Read-only workers. Read, search, navigate - no writes, edits or mutating
  commands. A worker that can edit drifts from reviewing into silently fixing,
  and the tree must end identical to how it started or findings cannot be
  checked against it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01G42vsxSKL7je39AsHZ5aJm
This commit is contained in:
Pawel Huryn
2026-09-13 15:29:49 +02:00
co-authored by Claude Opus 5
parent 18032bc9f7
commit 76c51f3ae8
2 changed files with 79 additions and 6 deletions
+25 -1
View File
@@ -130,7 +130,7 @@ checklist will not help; each needs an explicit probe:
removal and reuse. A label, a position or arrival order is suspicious exactly when those
properties can fail. Probe: run two operations concurrently and complete them out of order.
The full set of twelve diagnostic lenses, each with a detection tell, is in
The full set of thirteen diagnostic lenses, each with a detection tell, is in
`references/correctness-taxonomy.md`. They are overlapping lenses, not a quota to fill.
## Refutation: the discipline that makes this worth running
@@ -158,6 +158,15 @@ can harm only the person who triggered it and still be serious.** Equally, "keep
too permissive here — an ungrounded suspicion with no constructed execution is not a finding. When
both sub-cases are active, apply each test only to its own dimension.
**Absorption is not prevention.** The most expensive refutation mistake is finding something
downstream that happens to hide the defect - a cache that usually holds the value, a retry that
usually succeeds, a default that is usually right - and dropping the finding. That is not a
guarantee, it is a coincidence with good odds, and it fails the day the absorber is cold, evicted or
reconfigured. Drop only on a mechanism that makes the execution *impossible*, and say which mechanism
it was. For the same reason, **"it works nearly always" describes a race, not a refutation** - a
timing window that usually resolves correctly is a finding, and the fact that you had to reason about
which side usually wins is the evidence.
Passing tests, unfamiliar code, a suspicious name, a missing test and a sibling difference are
evidence to investigate — none of them is proof, and none is refutation. Deduplicate by violated
agreement and root cause, never by file. There is no findings quota; zero supported findings is a
@@ -187,6 +196,21 @@ drops to a cheaper model or a lower effort is the cheapest way to lose one. If a
or downgraded, say which in the report — a reader who assumes one model saw everything will
misjudge the coverage.
**One model, unless told otherwise.** Fan-out here is for coverage, not for a second opinion: every
worker runs the same model as the coordinator. Do not bring in a second model to cross-check findings
unless explicitly asked. Mixing models makes the result unattributable — when this skill is being
measured, or compared across models, a single foreign worker invalidates the number. The independent
second-model pass is a separate, explicitly-invoked step (`/ship-check` Step 6), never something this
skill reaches for on its own.
**Give workers read-only tools.** A review worker needs to read, search and navigate — nothing more.
Withhold file writes, edits and any mutating command; allow read, search and non-mutating inspection.
Three reasons, in order of importance: a worker that can edit will drift from reviewing into
"helpfully" fixing and stop reporting what it silently repaired; the repository under review must end
the run byte-identical to how it started, or the findings cannot be checked against it; and a
read-only worker cannot damage a working tree it misunderstood. If the host cannot restrict tools,
say so in the prompt and verify the tree is unchanged when the run ends.
## Report
Lead with supported findings, ordered by impact. Keep severity separate from evidential strength.