Both rerank call sites here — _rerank_pool in docs_mcp/server.py and RerankedRetriever in rag/retrieval.py — truncated docs to a flat 2000 characters as a stand-in for the reranker's 1024-token pair limit.
jina-reranker-v2 is a BERT cross-encoder with bert.context_length = 1024 and learned absolute position embeddings, so 1024 is a hard ceiling, not a tunable. llama.cpp 500s the entire batch if any one pair exceeds it, silently dropping that query to fused order.
Measured for this corpus
Via {RERANK_URL}/tokenize over 250 real chunks: floor 1.47 chars/token (EPA/Bayer label prose), worst observed 997 tokens at the old 2000-char cap.
That's under the ceiling on its own — which is exactly why this corpus looked fine — but the reranker scores (query, doc)pairs, so prepending a query pushes it over. This repo was less exposed than seed (0/250 chunks over 1000 tokens vs seed's 41/250), so the change here is largely defensive; the cap is now correct by construction rather than by luck.
What changed
RERANK_DOC_MAX_CHARS derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN (1.45, just under the measured 1.47) / a query reserve → 1091 chars, budgeting the pair to ~94% of the ceiling instead of exactly 1024.
The query is truncated too — previously only the document was, though it's the pair that must fit.
Both call sites now share the same derived constants, so rag/retrieval.py (what the eval harness exercises) and docs_mcp/server.py (what production runs) can no longer drift apart.
Truncation stays scoring-only; full label text is still what goes back to the user.
Eval — no regression
35 golden queries, k=5, pool=50, hybrid+rerank:
MRR
Recall@5
nDCG@5
Errors
before
0.667
0.643
0.627
0
after
0.667
0.643
0.627
0
Cutting the cap 2000 → 1091 chars cost nothing measurable.
Caveat worth recording: my first "after" run was invalid. The image ships precompiled __pycache__/*.pyc and Python loaded the stale bytecode over the docker cp'd source, so it measured the old code — and identical numbers are exactly what a no-op produces. The table above is from a re-run with __pycache__ cleared and rag.retrieval.RERANK_DOC_MAX_CHARS == 1091 verified as actually loaded.
Ports [docs-mcp-template#4](https://git.jpaul.io/justin/docs-mcp-template/pulls/4) to this clone. Companion: [seed-mcp](https://git.jpaul.io/justin/seed-mcp/pulls) (same fix, different measured ratio).
## The bug
Both rerank call sites here — `_rerank_pool` in `docs_mcp/server.py` and `RerankedRetriever` in `rag/retrieval.py` — truncated docs to a flat **2000 characters** as a stand-in for the reranker's **1024-token** pair limit.
`jina-reranker-v2` is a BERT cross-encoder with `bert.context_length = 1024` and learned absolute position embeddings, so 1024 is a hard ceiling, not a tunable. llama.cpp 500s the **entire batch** if any one pair exceeds it, silently dropping that query to fused order.
## Measured for this corpus
Via `{RERANK_URL}/tokenize` over 250 real chunks: **floor 1.47 chars/token** (EPA/Bayer label prose), worst observed **997 tokens** at the old 2000-char cap.
That's under the ceiling *on its own* — which is exactly why this corpus looked fine — but the reranker scores `(query, doc)` **pairs**, so prepending a query pushes it over. This repo was less exposed than seed (0/250 chunks over 1000 tokens vs seed's 41/250), so the change here is largely defensive; the cap is now correct by construction rather than by luck.
## What changed
- `RERANK_DOC_MAX_CHARS` derived from `RERANK_CTX_TOKENS` / `RERANK_CHARS_PER_TOKEN` (1.45, just under the measured 1.47) / a query reserve → **1091 chars**, budgeting the pair to ~94% of the ceiling instead of exactly 1024.
- The **query is truncated too** — previously only the document was, though it's the pair that must fit.
- Both call sites now share the same derived constants, so `rag/retrieval.py` (what the eval harness exercises) and `docs_mcp/server.py` (what production runs) can no longer drift apart.
Truncation stays scoring-only; full label text is still what goes back to the user.
## Eval — no regression
35 golden queries, k=5, pool=50, hybrid+rerank:
| | MRR | Recall@5 | nDCG@5 | Errors |
|---|---|---|---|---|
| before | 0.667 | 0.643 | 0.627 | 0 |
| **after** | **0.667** | **0.643** | **0.627** | **0** |
Cutting the cap 2000 → 1091 chars cost nothing measurable.
**Caveat worth recording:** my first "after" run was invalid. The image ships precompiled `__pycache__/*.pyc` and Python loaded the **stale bytecode** over the `docker cp`'d source, so it measured the old code — and identical numbers are exactly what a no-op produces. The table above is from a re-run with `__pycache__` cleared and `rag.retrieval.RERANK_DOC_MAX_CHARS == 1091` verified as actually loaded.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
Ports the docs-mcp-template fix. Both rerank call sites here (_rerank_pool in
docs_mcp/server.py and RerankedRetriever in rag/retrieval.py) truncated docs to
a flat 2000 CHARACTERS as a stand-in for the reranker's 1024-TOKEN pair limit.
jina-reranker-v2 is a BERT cross-encoder with bert.context_length=1024 and
learned absolute position embeddings — 1024 is a hard ceiling, not a tunable —
and llama.cpp 500s the ENTIRE batch if any one pair exceeds it, silently
dropping that query to fused order.
Measured floor for this corpus via {RERANK_URL}/tokenize: 1.47 chars/token
(EPA/Bayer label prose), worst observed 997 tokens at the old 2000-char cap —
under the ceiling alone, but over it once the query is prepended. The cap is
now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query reserve
(1091 chars here), budgeting the PAIR to ~94% rather than exactly 1024.
The query is now truncated too; previously only the document was, though it is
the pair that must fit.
Eval (hybrid+rerank, 35 golden queries, k=5, pool=50) — no regression:
before MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors
after MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors
Note when re-testing in a running container: the image ships precompiled
__pycache__/*.pyc and Python will load the STALE bytecode over a docker cp'd
source edit. rm -rf /app/<pkg>/__pycache__ first or you measure the old code.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Ports docs-mcp-template#4 to this clone. Companion: seed-mcp (same fix, different measured ratio).
The bug
Both rerank call sites here —
_rerank_poolindocs_mcp/server.pyandRerankedRetrieverinrag/retrieval.py— truncated docs to a flat 2000 characters as a stand-in for the reranker's 1024-token pair limit.jina-reranker-v2is a BERT cross-encoder withbert.context_length = 1024and learned absolute position embeddings, so 1024 is a hard ceiling, not a tunable. llama.cpp 500s the entire batch if any one pair exceeds it, silently dropping that query to fused order.Measured for this corpus
Via
{RERANK_URL}/tokenizeover 250 real chunks: floor 1.47 chars/token (EPA/Bayer label prose), worst observed 997 tokens at the old 2000-char cap.That's under the ceiling on its own — which is exactly why this corpus looked fine — but the reranker scores
(query, doc)pairs, so prepending a query pushes it over. This repo was less exposed than seed (0/250 chunks over 1000 tokens vs seed's 41/250), so the change here is largely defensive; the cap is now correct by construction rather than by luck.What changed
RERANK_DOC_MAX_CHARSderived fromRERANK_CTX_TOKENS/RERANK_CHARS_PER_TOKEN(1.45, just under the measured 1.47) / a query reserve → 1091 chars, budgeting the pair to ~94% of the ceiling instead of exactly 1024.rag/retrieval.py(what the eval harness exercises) anddocs_mcp/server.py(what production runs) can no longer drift apart.Truncation stays scoring-only; full label text is still what goes back to the user.
Eval — no regression
35 golden queries, k=5, pool=50, hybrid+rerank:
Cutting the cap 2000 → 1091 chars cost nothing measurable.
Caveat worth recording: my first "after" run was invalid. The image ships precompiled
__pycache__/*.pycand Python loaded the stale bytecode over thedocker cp'd source, so it measured the old code — and identical numbers are exactly what a no-op produces. The table above is from a re-run with__pycache__cleared andrag.retrieval.RERANK_DOC_MAX_CHARS == 1091verified as actually loaded.🤖 Generated with Claude Code
https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe
Ports the docs-mcp-template fix. Both rerank call sites here (_rerank_pool in docs_mcp/server.py and RerankedRetriever in rag/retrieval.py) truncated docs to a flat 2000 CHARACTERS as a stand-in for the reranker's 1024-TOKEN pair limit. jina-reranker-v2 is a BERT cross-encoder with bert.context_length=1024 and learned absolute position embeddings — 1024 is a hard ceiling, not a tunable — and llama.cpp 500s the ENTIRE batch if any one pair exceeds it, silently dropping that query to fused order. Measured floor for this corpus via {RERANK_URL}/tokenize: 1.47 chars/token (EPA/Bayer label prose), worst observed 997 tokens at the old 2000-char cap — under the ceiling alone, but over it once the query is prepended. The cap is now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query reserve (1091 chars here), budgeting the PAIR to ~94% rather than exactly 1024. The query is now truncated too; previously only the document was, though it is the pair that must fit. Eval (hybrid+rerank, 35 golden queries, k=5, pool=50) — no regression: before MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors after MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors Note when re-testing in a running container: the image ships precompiled __pycache__/*.pyc and Python will load the STALE bytecode over a docker cp'd source edit. rm -rf /app/<pkg>/__pycache__ first or you measure the old code. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe