Ports the docs-mcp-template fix. Both rerank call sites here (_rerank_pool in
docs_mcp/server.py and RerankedRetriever in rag/retrieval.py) truncated docs to
a flat 2000 CHARACTERS as a stand-in for the reranker's 1024-TOKEN pair limit.
jina-reranker-v2 is a BERT cross-encoder with bert.context_length=1024 and
learned absolute position embeddings — 1024 is a hard ceiling, not a tunable —
and llama.cpp 500s the ENTIRE batch if any one pair exceeds it, silently
dropping that query to fused order.
Measured floor for this corpus via {RERANK_URL}/tokenize: 1.47 chars/token
(EPA/Bayer label prose), worst observed 997 tokens at the old 2000-char cap —
under the ceiling alone, but over it once the query is prepended. The cap is
now derived from RERANK_CTX_TOKENS / RERANK_CHARS_PER_TOKEN / a query reserve
(1091 chars here), budgeting the PAIR to ~94% rather than exactly 1024.
The query is now truncated too; previously only the document was, though it is
the pair that must fit.
Eval (hybrid+rerank, 35 golden queries, k=5, pool=50) — no regression:
before MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors
after MRR 0.667 Recall@5 0.643 nDCG@5 0.627 0 errors
Note when re-testing in a running container: the image ships precompiled
__pycache__/*.pyc and Python will load the STALE bytecode over a docker cp'd
source edit. rm -rf /app/<pkg>/__pycache__ first or you measure the old code.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01AiYH8nxc6DgUTdwHnP9PEe