Files
crop-chem-docs/deploy/rerank-docker.md
T
claudeandClaude Opus 5 c363a221e1 docs(deploy): make deploy/ describe the deployment that actually exists
deploy/docker-compose.yml was unedited docs-mcp-template boilerplate —
untouched since the scaffold commit, still carrying <product>,
<registry> and <owner> placeholders — describing a standalone stack
that has never existed. crop-chem-docs runs as the `chem-mcp` service
inside Drawbar's parent compose.

It also set MCP_ALLOWED_HOSTS, which no code in this repo reads. The
knob is MCP_DISABLE_DNS_REBINDING_PROTECTION. Anyone who trusted the
old file and set an allowlist would have gotten a 421 on every request
with nothing in the logs to explain it.

- deploy/docker-compose.yml: replaced with the real chem-mcp block, a
  copy of what runs in Drawbar/drawbar-backend. Verified structurally
  identical to the parent (image, environment, expose, extra_hosts,
  restart, labels all equal). Carries the why for each setting: the
  :latest-vs-corpus-tag Watchtower trap (#339), the rebind-protection
  rationale, and that the OLLAMA_URL override is load-bearing because
  Drawbar's own ollama service is commented out — the image default
  http://ollama:11434 does not resolve in that stack, so without the
  override every search_docs call fails to embed its query.
- deploy/drawbar-compose-snippet.md: deleted. It was a second,
  differently-wrong copy (service name `crop-chem-docs`, ports
  8001:8000, and "No environment block needed — the image's defaults
  handle it", which is false on both the rebind and Ollama counts).
  Its still-true content (verification commands) moved into the compose
  file; the tag scheme and deploy chain were already in the README.
- deploy/rerank-docker.md: RERANK_URL said http://10.10.1.65:8082. In
  production the MCP reaches the sidecar by compose service name
  (http://llama-rerank:8080, baked into the image). Documents the
  network-attach gotcha that makes rerank fail silently, and keeps the
  host-IP form for local dev.
- README.md: file tree updated for the deleted file; Watchtower poll
  interval corrected 5 min -> 60s (WATCHTOWER_POLL_INTERVAL=60, as
  configured on trashpanda).

Verified: no <product>/<registry>/<owner> placeholders remain in
deploy/ or README, and all six env vars set in the block are ones the
server actually reads.

Closes #5

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FFBDnRWHispovmJVK9rXc9
2026-09-10 21:43:56 -04:00

103 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Reranker sidecar — llama.cpp + jina-reranker-v2-base
Phase 6 setup. The MCP server reads `RERANK_URL` and, when set, pipes
the top-50 dense (or hybrid) chunks through this sidecar before
returning to the LLM. See `docs_mcp/server.py:_rerank_pool`.
## Production deploy — trashpanda (Tesla P4, 8 GB VRAM)
This is where the reranker lives. Same box that runs the Drawbar
backend + Cloudflare Tunnel, so the MCP server can reach it on the
internal LAN.
```bash
ssh [email protected] \
'docker run -d --name llama-rerank --restart unless-stopped --gpus all \
-p 8082:8080 \
ghcr.io/ggml-org/llama.cpp:server-cuda \
-hf gpustack/jina-reranker-v2-base-multilingual-GGUF:Q8_0 \
--reranking --host 0.0.0.0 --port 8080 -ngl 99'
```
Key flags:
- `--gpus all` — pass through the Tesla P4
- `server-cuda` image — CUDA-built llama.cpp (not the CPU-only `:server`)
- `-ngl 99` — offload all layers to GPU
- `-hf <repo>` — auto-download from HuggingFace on first start (~280 MB,
cached in the container volume)
- `--reranking` — enables `/v1/rerank` endpoint
- `--restart unless-stopped` — survives reboot
VRAM usage: ~280 MB model + CUDA context. Well under the 8 GB the
Tesla P4 has, leaves room for nomic-embed-text (~560 MB) if you
later co-host it.
## Configure the MCP server
In production the MCP reaches this sidecar **over the Drawbar compose
network by service name**, not by host IP — `RERANK_URL` is baked into
the image as:
```
RERANK_URL=http://llama-rerank:8080
```
so the deployed `chem-mcp` service sets nothing. See
`deploy/docker-compose.yml`.
That only resolves if the `llama-rerank` container is attached to
`drawbar-backend_default`. If it is on the default bridge network
instead, the name resolves via public DNS to an unrelated IP and
connection-refuses — and `search_docs` falls back to dense+BM25
**silently**. Attach it with:
```bash
docker network connect drawbar-backend_default llama-rerank
```
For local dev outside that network, point at the published port
directly:
```bash
export RERANK_URL=http://10.10.1.65:8082
```
## Verify
```bash
curl http://10.10.1.65:8082/v1/rerank -H 'Content-Type: application/json' -d '{
"query": "soybean herbicide for waterhemp",
"documents": [
"Roundup Custom for fallow burndown",
"Sencor metribuzin controls waterhemp in soybean pre-emergence"
]
}'
```
Expect index=1 (the Sencor doc) at score ~0.8, index=0 at a strongly
negative score, in under 1 s.
## Performance reference
| Mode | Pool | Wall time |
|---|---|---|
| CPU (local 28-thread Xeon) | 50 docs | ~23 s |
| GPU (Tesla P4 on trashpanda) | 50 docs | ~0.7-1.5 s |
| GPU (Tesla P4) | 20 docs | ~0.4 s |
The Tesla P4 is Pascal-era (8.1 TFLOPs FP32) so a modern Ampere or
Ada Lovelace GPU would be ~3-5× faster, but for the row-crop label
corpus query rate the P4 is plenty.
## Troubleshooting
- **Model not on GPU?** Check `docker logs llama-rerank | grep CUDA` —
you should see `CUDA0 : Tesla P4 (8109 MiB, ... free)` and tensor
load lines. If you see CPU-only init, you forgot `--gpus all` or
used `:server` instead of `:server-cuda`.
- **Conflict with Ollama on the same GPU?** No — both processes can
share the GPU, CUDA handles VRAM partitioning. nomic-embed-text +
jina-reranker-v2-base together use ~840 MB on the 8 GB card.
- **First rerank call is slow (~4 s)?** Warm-up. Subsequent calls are
~0.7 s for 50 docs.