This corpus was the badly exposed one — the prod rerank failures traced here, not to crop-chem.
The bug
RERANK_DOC_MAX_CHARS was a flat 2000 characters standing in for the reranker's 1024-token pair limit, and the comment asserted that was "≈ 500-700 tokens" with headroom for the query. Measured against {RERANK_URL}/tokenize, that estimate is wrong for this corpus by nearly 2×:
Identifier-dense variety/trial tables tokenize far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the entire batch when any one (query, doc) pair exceeds bert.context_length = 1024 — a hard ceiling on a BERT cross-encoder with learned absolute position embeddings, not something a bigger --ubatch-size lifts — one oversized chunk silently dropped that whole query to fused order.
Observed in production as 1034 / 1040 / 1043-token rejections, which match this corpus's distribution almost exactly.
What changed
RERANK_DOC_MAX_CHARS derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN (1.90) / a query reserve → 1523 chars. Still generous precisely because this corpus's density is now accounted for rather than guessed — a naive "just make it smaller" fix would have cut far more content for no extra safety.
The query is truncated too; previously only the document was, though it's the pair that must fit.
eval/retrievers.py default moved in step, so the harness measures what production actually runs.
Budgets the pair to ~94% of the ceiling rather than exactly 1024, since chars-per-token is a sampled estimate.
Truncation stays scoring-only; full chunk text still goes to the caller.
Eval — no regression, failures eliminated
21 golden queries, k=5:
Retriever
Passed
Recall
P@1
MRR
hybrid+rerank
21/21
100.00%
90.48%
0.905
bm25
20/21
95.24%
85.71%
0.857
hybrid
19/21
90.48%
61.90%
0.698
dense
11/21
52.38%
33.33%
0.373
Identical to the documented baseline, and oversize rerank rejections went 3 → 0 in the same run that previously produced 3.
The hybrid row — 61.90% P@1 — is what every query hit by an oversized chunk was silently getting. That ~29-point gap is the cost this PR removes.
Related
The sidecar itself was also mis-launched (no -b/-ub, so llama.cpp clamped its physical batch to 512 and rejected anything over half the model's capacity). Fixed host-side already; this PR is the client half. Tracked in Drawbar/drawbar-backend#342.
Ports [docs-mcp-template#4](https://git.jpaul.io/justin/docs-mcp-template/pulls/4). Companion: [crop-chem-docs#9](https://git.jpaul.io/justin/crop-chem-docs/pulls/9).
**This corpus was the badly exposed one** — the prod rerank failures traced here, not to crop-chem.
## The bug
`RERANK_DOC_MAX_CHARS` was a flat **2000 characters** standing in for the reranker's **1024-token** pair limit, and the comment asserted that was *"≈ 500-700 tokens"* with headroom for the query. Measured against `{RERANK_URL}/tokenize`, that estimate is wrong for this corpus by nearly 2×:
```
chars/token floor 1.92
tokens @2000-char cap: max 1042 p99 1031 p95 1017
41 of 250 sampled chunks (16%) exceeded 1000 tokens
```
Identifier-dense variety/trial tables tokenize far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the **entire batch** when any one `(query, doc)` pair exceeds `bert.context_length = 1024` — a hard ceiling on a BERT cross-encoder with learned absolute position embeddings, not something a bigger `--ubatch-size` lifts — one oversized chunk silently dropped that **whole query** to fused order.
Observed in production as 1034 / 1040 / 1043-token rejections, which match this corpus's distribution almost exactly.
## What changed
- `RERANK_DOC_MAX_CHARS` derived from `RERANK_CTX_TOKENS` / `RERANK_CHARS_PER_TOKEN` (1.90) / a query reserve → **1523 chars**. Still generous precisely *because* this corpus's density is now accounted for rather than guessed — a naive "just make it smaller" fix would have cut far more content for no extra safety.
- The **query is truncated too**; previously only the document was, though it's the pair that must fit.
- `eval/retrievers.py` default moved in step, so the harness measures what production actually runs.
- Budgets the pair to ~94% of the ceiling rather than exactly 1024, since chars-per-token is a sampled estimate.
Truncation stays scoring-only; full chunk text still goes to the caller.
## Eval — no regression, failures eliminated
21 golden queries, k=5:
| Retriever | Passed | Recall | P@1 | MRR |
|---|---|---|---|---|
| **hybrid+rerank** | 21/21 | 100.00% | **90.48%** | **0.905** |
| bm25 | 20/21 | 95.24% | 85.71% | 0.857 |
| hybrid | 19/21 | 90.48% | 61.90% | 0.698 |
| dense | 11/21 | 52.38% | 33.33% | 0.373 |
Identical to the documented baseline, **and oversize rerank rejections went 3 → 0** in the same run that previously produced 3.
The `hybrid` row — **61.90% P@1** — is what every query hit by an oversized chunk was silently getting. That ~29-point gap is the cost this PR removes.
## Related
The sidecar itself was also mis-launched (no `-b`/`-ub`, so llama.cpp clamped its physical batch to 512 and rejected anything over *half* the model's capacity). Fixed host-side already; this PR is the client half. Tracked in Drawbar/drawbar-backend#342.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
Ports the docs-mcp-template fix. This corpus was the badly exposed one.
RERANK_DOC_MAX_CHARS was a flat 2000 CHARACTERS standing in for the reranker's
1024-TOKEN pair limit, and the comment claimed that was "≈ 500-700 tokens" with
headroom for the query. Measured against {RERANK_URL}/tokenize, that estimate
is wrong for this corpus by nearly 2x:
chars/token floor 1.92 tokens @2000-char cap: max 1042, p99 1031, p95 1017
41 of 250 sampled chunks (16%) exceeded 1000 tokens
Identifier-dense variety/trial tables tokenise far worse than prose (crop-chem's
floor is 1.47). Since llama.cpp 500s the ENTIRE batch when any one (query, doc)
pair exceeds bert.context_length=1024 — a hard architectural ceiling on a BERT
cross-encoder, not something a bigger --ubatch-size lifts — one oversized chunk
silently dropped that whole query to fused order. Observed in prod as 1034/1040/
1043-token rejections.
The cap is now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query
reserve (1523 chars here — still generous, because this corpus's density is
accounted for rather than guessed), budgeting the PAIR to ~94% of the ceiling.
The query is truncated too; previously only the document was. eval/retrievers.py
default moved in step so the harness measures what production runs.
Eval (21 golden queries, k=5) — no regression, failures eliminated:
hybrid+rerank 21/21 Recall 100% P@1 90.48% MRR 0.905 (unchanged)
oversize rerank rejections during the run: 3 -> 0
(same run's no-rerank `hybrid` row: P@1 61.90% — the cliff this avoids)
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
justin
merged commit 4de5aaa02f into main2026-09-10 22:00:30 -04:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Ports docs-mcp-template#4. Companion: crop-chem-docs#9.
This corpus was the badly exposed one — the prod rerank failures traced here, not to crop-chem.
The bug
RERANK_DOC_MAX_CHARSwas a flat 2000 characters standing in for the reranker's 1024-token pair limit, and the comment asserted that was "≈ 500-700 tokens" with headroom for the query. Measured against{RERANK_URL}/tokenize, that estimate is wrong for this corpus by nearly 2×:Identifier-dense variety/trial tables tokenize far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the entire batch when any one
(query, doc)pair exceedsbert.context_length = 1024— a hard ceiling on a BERT cross-encoder with learned absolute position embeddings, not something a bigger--ubatch-sizelifts — one oversized chunk silently dropped that whole query to fused order.Observed in production as 1034 / 1040 / 1043-token rejections, which match this corpus's distribution almost exactly.
What changed
RERANK_DOC_MAX_CHARSderived fromRERANK_CTX_TOKENS/RERANK_CHARS_PER_TOKEN(1.90) / a query reserve → 1523 chars. Still generous precisely because this corpus's density is now accounted for rather than guessed — a naive "just make it smaller" fix would have cut far more content for no extra safety.eval/retrievers.pydefault moved in step, so the harness measures what production actually runs.Truncation stays scoring-only; full chunk text still goes to the caller.
Eval — no regression, failures eliminated
21 golden queries, k=5:
Identical to the documented baseline, and oversize rerank rejections went 3 → 0 in the same run that previously produced 3.
The
hybridrow — 61.90% P@1 — is what every query hit by an oversized chunk was silently getting. That ~29-point gap is the cost this PR removes.Related
The sidecar itself was also mis-launched (no
-b/-ub, so llama.cpp clamped its physical batch to 512 and rejected anything over half the model's capacity). Fixed host-side already; this PR is the client half. Tracked in Drawbar/drawbar-backend#342.🤖 Generated with Claude Code
https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
Ports the docs-mcp-template fix. This corpus was the badly exposed one. RERANK_DOC_MAX_CHARS was a flat 2000 CHARACTERS standing in for the reranker's 1024-TOKEN pair limit, and the comment claimed that was "≈ 500-700 tokens" with headroom for the query. Measured against {RERANK_URL}/tokenize, that estimate is wrong for this corpus by nearly 2x: chars/token floor 1.92 tokens @2000-char cap: max 1042, p99 1031, p95 1017 41 of 250 sampled chunks (16%) exceeded 1000 tokens Identifier-dense variety/trial tables tokenise far worse than prose (crop-chem's floor is 1.47). Since llama.cpp 500s the ENTIRE batch when any one (query, doc) pair exceeds bert.context_length=1024 — a hard architectural ceiling on a BERT cross-encoder, not something a bigger --ubatch-size lifts — one oversized chunk silently dropped that whole query to fused order. Observed in prod as 1034/1040/ 1043-token rejections. The cap is now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query reserve (1523 chars here — still generous, because this corpus's density is accounted for rather than guessed), budgeting the PAIR to ~94% of the ceiling. The query is truncated too; previously only the document was. eval/retrievers.py default moved in step so the harness measures what production runs. Eval (21 golden queries, k=5) — no regression, failures eliminated: hybrid+rerank 21/21 Recall 100% P@1 90.48% MRR 0.905 (unchanged) oversize rerank rejections during the run: 3 -> 0 (same run's no-rerank `hybrid` row: P@1 61.90% — the cliff this avoids) Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe