fix(rerank): derive the doc cap from tokens (measured floor 1.92), cap the query #24

Merged
justin merged 1 commits from fix/rerank-token-budget into main 2026-09-10 22:00:30 -04:00
Owner

Ports docs-mcp-template#4. Companion: crop-chem-docs#9.

This corpus was the badly exposed one — the prod rerank failures traced here, not to crop-chem.

The bug

RERANK_DOC_MAX_CHARS was a flat 2000 characters standing in for the reranker's 1024-token pair limit, and the comment asserted that was "≈ 500-700 tokens" with headroom for the query. Measured against {RERANK_URL}/tokenize, that estimate is wrong for this corpus by nearly 2×:

chars/token floor  1.92
tokens @2000-char cap:  max 1042   p99 1031   p95 1017
41 of 250 sampled chunks (16%) exceeded 1000 tokens

Identifier-dense variety/trial tables tokenize far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the entire batch when any one (query, doc) pair exceeds bert.context_length = 1024 — a hard ceiling on a BERT cross-encoder with learned absolute position embeddings, not something a bigger --ubatch-size lifts — one oversized chunk silently dropped that whole query to fused order.

Observed in production as 1034 / 1040 / 1043-token rejections, which match this corpus's distribution almost exactly.

What changed

  • RERANK_DOC_MAX_CHARS derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN (1.90) / a query reserve → 1523 chars. Still generous precisely because this corpus's density is now accounted for rather than guessed — a naive "just make it smaller" fix would have cut far more content for no extra safety.
  • The query is truncated too; previously only the document was, though it's the pair that must fit.
  • eval/retrievers.py default moved in step, so the harness measures what production actually runs.
  • Budgets the pair to ~94% of the ceiling rather than exactly 1024, since chars-per-token is a sampled estimate.

Truncation stays scoring-only; full chunk text still goes to the caller.

Eval — no regression, failures eliminated

21 golden queries, k=5:

Retriever Passed Recall P@1 MRR
hybrid+rerank 21/21 100.00% 90.48% 0.905
bm25 20/21 95.24% 85.71% 0.857
hybrid 19/21 90.48% 61.90% 0.698
dense 11/21 52.38% 33.33% 0.373

Identical to the documented baseline, and oversize rerank rejections went 3 → 0 in the same run that previously produced 3.

The hybrid row — 61.90% P@1 — is what every query hit by an oversized chunk was silently getting. That ~29-point gap is the cost this PR removes.

Related

The sidecar itself was also mis-launched (no -b/-ub, so llama.cpp clamped its physical batch to 512 and rejected anything over half the model's capacity). Fixed host-side already; this PR is the client half. Tracked in Drawbar/drawbar-backend#342.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe

Ports [docs-mcp-template#4](https://git.jpaul.io/justin/docs-mcp-template/pulls/4). Companion: [crop-chem-docs#9](https://git.jpaul.io/justin/crop-chem-docs/pulls/9). **This corpus was the badly exposed one** — the prod rerank failures traced here, not to crop-chem. ## The bug `RERANK_DOC_MAX_CHARS` was a flat **2000 characters** standing in for the reranker's **1024-token** pair limit, and the comment asserted that was *"≈ 500-700 tokens"* with headroom for the query. Measured against `{RERANK_URL}/tokenize`, that estimate is wrong for this corpus by nearly 2×: ``` chars/token floor 1.92 tokens @2000-char cap: max 1042 p99 1031 p95 1017 41 of 250 sampled chunks (16%) exceeded 1000 tokens ``` Identifier-dense variety/trial tables tokenize far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the **entire batch** when any one `(query, doc)` pair exceeds `bert.context_length = 1024` — a hard ceiling on a BERT cross-encoder with learned absolute position embeddings, not something a bigger `--ubatch-size` lifts — one oversized chunk silently dropped that **whole query** to fused order. Observed in production as 1034 / 1040 / 1043-token rejections, which match this corpus's distribution almost exactly. ## What changed - `RERANK_DOC_MAX_CHARS` derived from `RERANK_CTX_TOKENS` / `RERANK_CHARS_PER_TOKEN` (1.90) / a query reserve → **1523 chars**. Still generous precisely *because* this corpus's density is now accounted for rather than guessed — a naive "just make it smaller" fix would have cut far more content for no extra safety. - The **query is truncated too**; previously only the document was, though it's the pair that must fit. - `eval/retrievers.py` default moved in step, so the harness measures what production actually runs. - Budgets the pair to ~94% of the ceiling rather than exactly 1024, since chars-per-token is a sampled estimate. Truncation stays scoring-only; full chunk text still goes to the caller. ## Eval — no regression, failures eliminated 21 golden queries, k=5: | Retriever | Passed | Recall | P@1 | MRR | |---|---|---|---|---| | **hybrid+rerank** | 21/21 | 100.00% | **90.48%** | **0.905** | | bm25 | 20/21 | 95.24% | 85.71% | 0.857 | | hybrid | 19/21 | 90.48% | 61.90% | 0.698 | | dense | 11/21 | 52.38% | 33.33% | 0.373 | Identical to the documented baseline, **and oversize rerank rejections went 3 → 0** in the same run that previously produced 3. The `hybrid` row — **61.90% P@1** — is what every query hit by an oversized chunk was silently getting. That ~29-point gap is the cost this PR removes. ## Related The sidecar itself was also mis-launched (no `-b`/`-ub`, so llama.cpp clamped its physical batch to 512 and rejected anything over *half* the model's capacity). Fixed host-side already; this PR is the client half. Tracked in Drawbar/drawbar-backend#342. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
justin added 1 commit 2026-09-10 21:52:17 -04:00
Ports the docs-mcp-template fix. This corpus was the badly exposed one.

RERANK_DOC_MAX_CHARS was a flat 2000 CHARACTERS standing in for the reranker's
1024-TOKEN pair limit, and the comment claimed that was "≈ 500-700 tokens" with
headroom for the query. Measured against {RERANK_URL}/tokenize, that estimate
is wrong for this corpus by nearly 2x:

  chars/token floor 1.92   tokens @2000-char cap: max 1042, p99 1031, p95 1017
  41 of 250 sampled chunks (16%) exceeded 1000 tokens

Identifier-dense variety/trial tables tokenise far worse than prose (crop-chem's
floor is 1.47). Since llama.cpp 500s the ENTIRE batch when any one (query, doc)
pair exceeds bert.context_length=1024 — a hard architectural ceiling on a BERT
cross-encoder, not something a bigger --ubatch-size lifts — one oversized chunk
silently dropped that whole query to fused order. Observed in prod as 1034/1040/
1043-token rejections.

The cap is now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query
reserve (1523 chars here — still generous, because this corpus's density is
accounted for rather than guessed), budgeting the PAIR to ~94% of the ceiling.
The query is truncated too; previously only the document was. eval/retrievers.py
default moved in step so the harness measures what production runs.

Eval (21 golden queries, k=5) — no regression, failures eliminated:
  hybrid+rerank  21/21  Recall 100%  P@1 90.48%  MRR 0.905   (unchanged)
  oversize rerank rejections during the run: 3 -> 0
  (same run's no-rerank `hybrid` row: P@1 61.90% — the cliff this avoids)

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
justin merged commit 4de5aaa02f into main 2026-09-10 22:00:30 -04:00
justin deleted branch fix/rerank-token-budget 2026-09-10 22:00:30 -04:00
Sign in to join this conversation.