Files
crop-chem-docs/deploy/rerank-docker.md
T

103 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Reranker sidecar — llama.cpp + jina-reranker-v2-base
Phase 6 setup. The MCP server reads `RERANK_URL` and, when set, pipes
the top-50 dense (or hybrid) chunks through this sidecar before
returning to the LLM. See `docs_mcp/server.py:_rerank_pool`.
## Production deploy — trashpanda (Tesla P4, 8 GB VRAM)
This is where the reranker lives. Same box that runs the Drawbar
backend + Cloudflare Tunnel, so the MCP server can reach it on the
internal LAN.
```bash
ssh [email protected] \
'docker run -d --name llama-rerank --restart unless-stopped --gpus all \
-p 8082:8080 \
ghcr.io/ggml-org/llama.cpp:server-cuda \
-hf gpustack/jina-reranker-v2-base-multilingual-GGUF:Q8_0 \
--reranking --host 0.0.0.0 --port 8080 -ngl 99'
```
Key flags:
- `--gpus all` — pass through the Tesla P4
- `server-cuda` image — CUDA-built llama.cpp (not the CPU-only `:server`)
- `-ngl 99` — offload all layers to GPU
- `-hf <repo>` — auto-download from HuggingFace on first start (~280 MB,
cached in the container volume)
- `--reranking` — enables `/v1/rerank` endpoint
- `--restart unless-stopped` — survives reboot
VRAM usage: ~280 MB model + CUDA context. Well under the 8 GB the
Tesla P4 has, leaves room for nomic-embed-text (~560 MB) if you
later co-host it.
## Configure the MCP server
In production the MCP reaches this sidecar **over the Drawbar compose
network by service name**, not by host IP — `RERANK_URL` is baked into
the image as:
```
RERANK_URL=http://llama-rerank:8080
```
so the deployed `chem-mcp` service sets nothing. See
`deploy/docker-compose.yml`.
That only resolves if the `llama-rerank` container is attached to
`drawbar-backend_default`. If it is on the default bridge network
instead, the name resolves via public DNS to an unrelated IP and
connection-refuses — and `search_docs` falls back to dense+BM25
**silently**. Attach it with:
```bash
docker network connect drawbar-backend_default llama-rerank
```
For local dev outside that network, point at the published port
directly:
```bash
export RERANK_URL=http://10.10.1.65:8082
```
## Verify
```bash
curl http://10.10.1.65:8082/v1/rerank -H 'Content-Type: application/json' -d '{
"query": "soybean herbicide for waterhemp",
"documents": [
"Roundup Custom for fallow burndown",
"Sencor metribuzin controls waterhemp in soybean pre-emergence"
]
}'
```
Expect index=1 (the Sencor doc) at score ~0.8, index=0 at a strongly
negative score, in under 1 s.
## Performance reference
| Mode | Pool | Wall time |
|---|---|---|
| CPU (local 28-thread Xeon) | 50 docs | ~23 s |
| GPU (Tesla P4 on trashpanda) | 50 docs | ~0.7-1.5 s |
| GPU (Tesla P4) | 20 docs | ~0.4 s |
The Tesla P4 is Pascal-era (8.1 TFLOPs FP32) so a modern Ampere or
Ada Lovelace GPU would be ~3-5× faster, but for the row-crop label
corpus query rate the P4 is plenty.
## Troubleshooting
- **Model not on GPU?** Check `docker logs llama-rerank | grep CUDA` —
you should see `CUDA0 : Tesla P4 (8109 MiB, ... free)` and tensor
load lines. If you see CPU-only init, you forgot `--gpus all` or
used `:server` instead of `:server-cuda`.
- **Conflict with Ollama on the same GPU?** No — both processes can
share the GPU, CUDA handles VRAM partitioning. nomic-embed-text +
jina-reranker-v2-base together use ~840 MB on the 8 GB card.
- **First rerank call is slow (~4 s)?** Warm-up. Subsequent calls are
~0.7 s for 50 docs.