docs(deploy): make deploy/ describe the deployment that actually exists (#8)

Co-authored-by: claude <[email protected]>
This commit was merged in pull request #8.
This commit is contained in:
2026-09-10 21:45:01 -04:00
committed by claude
parent 254c4df71d
commit 053884f9fb
4 changed files with 135 additions and 250 deletions
+4 -4
View File
@@ -91,9 +91,9 @@ PRODUCT_NAME=crop_chem python -m docs_mcp.server --transport stdio
├── requirements.txt ├── requirements.txt
├── Dockerfile ├── Dockerfile
├── deploy/ ├── deploy/
│ ├── docker-compose.yml # Drop-in compose for Drawbar │ ├── docker-compose.yml # The chem-mcp service block to merge
│ ├── drawbar-compose-snippet.md # Notes on the parent compose merge │ │ # into Drawbar's parent compose
│ └── rerank-docker.md # llama-rerank service deployment │ └── rerank-docker.md # llama-rerank sidecar deployment
├── .gitea/workflows/ ├── .gitea/workflows/
│ ├── refresh.yml # Monthly cron: scrape + index + image push │ ├── refresh.yml # Monthly cron: scrape + index + image push
│ └── image-only.yml # On-demand code-only ship cycle │ └── image-only.yml # On-demand code-only ship cycle
@@ -140,7 +140,7 @@ Same Watchtower auto-deploy chain as seed-mcp. On every push to `main` that touc
2. Rebuilds Chroma + BM25 (~few min on the GPU pool) 2. Rebuilds Chroma + BM25 (~few min on the GPU pool)
3. `docker build` + push three tags to the LAN registry 3. `docker build` + push three tags to the LAN registry
4. Links the package to the repo via Gitea API 4. Links the package to the repo via Gitea API
5. Watchtower on trashpanda polls `:latest` every 5 min → recreates `drawbar-backend-chem-mcp-1` 5. Watchtower on trashpanda polls `:latest` every 60s → recreates `drawbar-backend-chem-mcp-1`
Corpus refresh runs monthly via `refresh.yml`. EPA PPLS is the slow source — ~hours at 1 req/sec at full scale. Corpus refresh runs monthly via `refresh.yml`. EPA PPLS is the slow source — ~hours at 1 req/sec at full scale.
+106 -93
View File
@@ -1,106 +1,119 @@
# Hosting stack for a docs MCP server. # crop-chem-docs service block to MERGE into Drawbar's parent compose
# file at /home/justin/drawbar/drawbar-backend/docker-compose.yml on
# trashpanda (10.10.1.65).
# #
# Replace <product> below with your product name on first deploy. # This is NOT a standalone stack — do not `docker compose up` this file
# Volumes: usage logs are mounted to a host path so they survive # on its own. The MCP is one service inside the Drawbar backend stack,
# Watchtower-driven container recreates. # where it is reached over the internal docker network as
# `chem-mcp:8080` by drawbar-backend-api (CHEM_MCP_BASE_URL). Its tools
# land in the advisor's catalog under the `chem:` prefix via the
# mcp_client multiplex. Sibling seed-mcp sits alongside it.
# #
# This template assumes a reverse proxy / Cloudflare Tunnel terminates # Keep this block in sync with the parent compose. It is a copy of what
# TLS in front of port 8000. Adjust if your infra differs. # actually runs, so that this repo's deploy/ documents reality rather
# than an aspiration.
services: services:
# The MCP server. Watchtower auto-pulls on :latest changes. # crop-chem-docs — ~4.1K US row-crop pesticide / herbicide labels
<product>-docs-mcp: # (EPA PPLS + Bayer), ~219K chunks. The advisor consults it for label
image: <registry>/<owner>/<product>-docs-mcp:latest # rates, REI/PHI and rotation restrictions. Chroma + BM25 indexes are
container_name: <product>-docs-mcp # baked into the image, so there's no cold-start corpus build, no DB
restart: unless-stopped # and no auth.
ports: chem-mcp:
- "8000:8000" # :latest, NOT a corpus- tag. Watchtower only re-pulls the tag the
# container is already running, and corpus-YYYY.MM.DD tags are minted
# once and never re-pushed — so pinning one while keeping the
# watchtower label makes the opt-in silently inert. That is how prod
# sat on the May 2026 corpus through two refreshes until 2026-09-10
# (Drawbar/drawbar-backend#339). To freeze a snapshot deliberately,
# pin the corpus tag AND drop the watchtower label.
image: git.jpaul.io/justin/crop-chem-docs:latest
environment: environment:
PRODUCT_NAME: "<product>"
PRODUCT_DOCS_URL: "https://docs.example.com"
# Streamable-HTTP transport. Stateless mode is required for
# production: clients don't lose sessions when Watchtower
# recreates the container.
MCP_TRANSPORT: streamable-http MCP_TRANSPORT: streamable-http
MCP_HOST: 0.0.0.0 MCP_HOST: 0.0.0.0
MCP_PORT: "8000" MCP_PORT: "8080"
# DNS-rebinding protection rejects any Host header that isn't in
# If you run MetaMCP or another gateway in front and reach # its (empty by default) allowlist, with a 421. On an internal
# this container via its compose DNS name (e.g. <product>-docs-mcp:8000), # docker network the caller's Host is `chem-mcp:8080` — exactly
# add that hostname here. "*" disables the rebind check entirely. # what gets rejected. Safe to disable here: the container is
MCP_ALLOWED_HOSTS: "<product>-docs-mcp,localhost,127.0.0.1" # `expose`d only, never published to a host port, so it is only
# reachable from inside the compose network.
# Phase 6 — reranker sidecar (jina-reranker-v2-base via llama.cpp). #
RERANK_URL: http://<product>-rerank:8080 # NOTE: this is the only knob this server has for it. There is no
RERANK_POOL: "200" # MCP_ALLOWED_HOSTS — the code does not read such a variable.
RERANK_TIMEOUT: "30" MCP_DISABLE_DNS_REBINDING_PROTECTION: "1"
# Query-time embeddings hit Ollama. Drawbar's own `ollama` compose
# Phase 8 — hybrid retrieval (BM25 + dense + RRF). Set true # service is commented out, so the image default
# only after the eval harness shows the dense-only path # (OLLAMA_URL=http://ollama:11434) does NOT resolve in this stack —
# missing technical-term queries that BM25 catches. # this override is load-bearing, not cosmetic. Without it every
HYBRID_SEARCH: "true" # search_docs call fails to embed its query.
OLLAMA_URL: ${CHEM_OLLAMA_URL:-http://host.docker.internal:11434}
# Phase 10 — usage telemetry. EMBED_MODEL: ${CHEM_EMBED_MODEL:-nomic-embed-text}
USAGE_LOG_DIR: /app/var/logs # Not set here on purpose — these come from the image's ENV
USAGE_LOG_KEEP_DAYS: "90" # defaults (see Dockerfile) and are correct for this stack:
volumes: # PRODUCT_NAME=crop_chem
# Usage logs persist across container recreates. # HYBRID_SEARCH=true
- ./<product>-docs-mcp-logs:/app/var/logs # RERANK_URL=http://llama-rerank:8080
depends_on: # Override any of them here if the stack's service names differ.
- <product>-rerank # Hybrid + rerank is the eval-validated config (MRR 0.672 vs 0.544
# for BM25 alone; see eval/results/with_rerank.md). Hybrid WITHOUT
# rerank is worse than BM25 alone — don't ship that combination.
extra_hosts:
- "host.docker.internal:host-gateway"
expose:
- "8080"
restart: unless-stopped
labels: labels:
# Watchtower polls *only* containers with this label set true. # Watchtower auto-pulls :latest on push from CI. The label is
# required because the Drawbar stack's watchtower runs in
# label-mode (WATCHTOWER_LABEL_ENABLE=true); it polls every 60s.
com.centurylinklabs.watchtower.enable: "true" com.centurylinklabs.watchtower.enable: "true"
networks:
- mcp
# Reranker sidecar — llama.cpp serving jina-reranker-v2-base.
# Requires GPU access; adjust runtime/devices for your hardware.
<product>-rerank:
image: ghcr.io/ggml-org/llama.cpp:server-cuda
container_name: <product>-rerank
restart: unless-stopped
# Mount the GGUF model from the host. Download from huggingface
# (gguf-org/jina-reranker-v2-base-multilingual-GGUF) first.
volumes:
- /path/to/models:/models:ro
command: >
--model /models/jina-reranker-v2-base.Q8_0.gguf
--reranking
--host 0.0.0.0
--port 8080
--n-gpu-layers 99
--ctx-size 4096
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
networks:
- mcp
# Watchtower — auto-pulls :latest on push. # ─── llama-rerank ────────────────────────────────────────────────────
# Only watches containers labeled `com.centurylinklabs.watchtower.enable=true`. #
watchtower: # The reranker is a SHARED sidecar (chem-mcp and seed-mcp both use it),
image: containrrr/watchtower:latest # and it is not declared in the parent compose — it runs as a standalone
container_name: watchtower # container. See deploy/rerank-docker.md for how to stand it up.
restart: unless-stopped #
volumes: # The gotcha: it must be attached to the `drawbar-backend_default`
- /var/run/docker.sock:/var/run/docker.sock # network or `RERANK_URL=http://llama-rerank:8080` resolves via public
environment: # DNS to an unrelated IP and connection-refuses. The MCP then falls back
WATCHTOWER_POLL_INTERVAL: "300" # 5 min # to dense+BM25 SILENTLY — retrieval quality craters with no error in
WATCHTOWER_LABEL_ENABLE: "true" # the log. This bit chem-mcp through 2026-05-25. To fix or re-fix:
WATCHTOWER_CLEANUP: "true" # remove old images after pull #
# If your registry requires auth, mount a docker config: # docker network connect drawbar-backend_default llama-rerank
# volumes: #
# - ./registry-auth.json:/config.json:ro # It is idempotent, but it does NOT survive the container being
networks: # recreated. Better: bring llama-rerank into the parent compose so the
- mcp # attachment is declarative.
#
# Confirm it is actually engaged — the header says which mode ran:
#
# docker exec drawbar-backend-chem-mcp-1 python -c \
# "from docs_mcp.server import search_docs; \
# print(search_docs('soybean herbicide for waterhemp', k=2))"
#
# Expect `mode=hybrid-rrf+rerank`. If it reads `mode=hybrid-rrf`, the
# sidecar is unreachable and you are serving degraded results.
networks: # ─── verifying a deploy ──────────────────────────────────────────────
mcp: #
driver: bridge # docker exec drawbar-backend-chem-mcp-1 python -c \
# "from docs_mcp.server import corpus_status; print(corpus_status())"
#
# Expect the label/chunk counts and the active feature flags. To confirm
# the transport itself, from the api container:
#
# docker exec drawbar-backend-api-1 python -c \
# "import urllib.request, json; \
# req=urllib.request.Request('http://chem-mcp:8080/mcp', \
# data=json.dumps({'jsonrpc':'2.0','id':1,'method':'initialize', \
# 'params':{'protocolVersion':'2025-06-18','capabilities':{}, \
# 'clientInfo':{'name':'smoke','version':'0'}}}).encode(), \
# headers={'Content-Type':'application/json', \
# 'Accept':'application/json, text/event-stream'}); \
# print(urllib.request.urlopen(req, timeout=15).status)"
#
# Expect 200.
-148
View File
@@ -1,148 +0,0 @@
# Drawbar deploy — `crop-chem-docs` MCP server snippet
Drop these two services into Drawbar's `docker-compose.yml`. Targets
the trashpanda stack: shared Docker network with the existing
Drawbar services + the Cloudflare Tunnel.
## Pre-reqs (one-time on the deploy host)
1. **Docker login to the Gitea registry:**
```bash
docker login git.jpaul.io -u justin # PAT for password
```
2. **NVIDIA Container Toolkit** — already installed on trashpanda
(the existing standalone `llama-rerank` container ran with
`--gpus all` fine).
3. **If a standalone `llama-rerank` container is already running**
(left over from earlier setup), remove it so the compose service
can bind the same name:
```bash
docker rm -f llama-rerank
```
## Compose services
```yaml
services:
# ---- Reranker sidecar -----------------------------------------
# jina-reranker-v2-base-multilingual via llama.cpp on the Tesla P4.
# Internal port only (no host port mapping needed — the MCP reaches
# it via Docker DNS). ~280 MB GPU VRAM at idle, ~500 MB during a
# 50-doc rerank. Co-exists fine with any other GPU users on the P4.
llama-rerank:
image: ghcr.io/ggml-org/llama.cpp:server-cuda
container_name: llama-rerank
restart: unless-stopped
command:
- "-hf"
- "gpustack/jina-reranker-v2-base-multilingual-GGUF:Q8_0"
- "--reranking"
- "--host"
- "0.0.0.0"
- "--port"
- "8080"
- "-ngl"
- "99" # offload all layers to GPU
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
# Model cache survives container recreates; first start downloads
# the GGUF (~280 MB) from HuggingFace.
volumes:
- llama-rerank-cache:/root/.cache/huggingface
networks:
- default
# ---- MCP server ------------------------------------------------
crop-chem-docs:
image: git.jpaul.io/justin/crop-chem-docs:corpus-2026.05.24
# :latest for dev / Watchtower auto-pull
container_name: crop-chem-docs
restart: unless-stopped
ports:
- "8001:8000" # MCP server (streamable-http). Adjust host port.
# No environment block needed — the image's defaults handle it:
# OLLAMA_URL=http://ollama:11434
# RERANK_URL=http://llama-rerank:8080
# HYBRID_SEARCH=true
# PRODUCT_NAME=crop_chem
# Override here only if your services have different names.
depends_on:
- llama-rerank
networks:
- default
labels:
com.centurylinklabs.watchtower.enable: "true"
volumes:
llama-rerank-cache:
```
## Note on the existing `ollama` service
The Dockerfile default is `OLLAMA_URL=http://ollama:11434` — that
assumes there's an `ollama` service in the same compose stack. If
trashpanda's Ollama is a host-mode process (not a compose service),
override the env in the `crop-chem-docs` block:
```yaml
environment:
OLLAMA_URL: "http://host.docker.internal:11434"
extra_hosts:
- "host.docker.internal:host-gateway"
```
Or just add Ollama itself to the compose stack as a sibling service.
## Test once both are up
```bash
docker compose up -d llama-rerank crop-chem-docs
# Wait ~10s for both to come up, then:
docker exec crop-chem-docs python -c \
"from docs_mcp.server import corpus_status; print(corpus_status())"
```
Expect: `# crop-chem-docs corpus status`, 4,159 labels, 216,467
chunks, BM25 db present, `RERANK_URL=http://llama-rerank:8080`,
`HYBRID_SEARCH=on`.
Then a live search to verify hybrid+rerank:
```bash
docker exec crop-chem-docs python -c \
"from docs_mcp.server import search_docs; print(search_docs('soybean herbicide for waterhemp', k=2))"
```
Expect: 2 hits with Sencor/Tackle/Warrant in top-2, `mode=hybrid-rrf+rerank` in the header.
## What the MCP container exposes
| Tool | What it does |
|---|---|
| `search_docs` | Hybrid+rerank pesticide-label search with optional filters |
| `get_page` | Full label markdown + metadata by `(source, source_key)` |
| `list_versions` | Discover sources, product classes, signal words, registrants |
| `corpus_status` | Counts + freshness; useful for health probes |
| `crop_chem_api_lessons` | Curated agronomy / label-handling knowledge — call before recommending |
## Tag scheme
| Tag | When | Use for |
|---|---|---|
| `:latest` | Every monthly refresh + every code push | Dev / Watchtower auto-pull |
| `:<sha12>` | Every build | Rollback pin |
| `:corpus-YYYY.MM.DD` | Every build | **Production pin** (frozen corpus version) |
## Updating the corpus
- **Monthly cron** — 1st @ 06:00 UTC, full re-scrape of Bayer + EPA PPLS,
reindex, image push. Watchtower pulls the new `:latest` automatically.
- **Manual** — Gitea Actions UI → `Monthly corpus refresh` → `Run workflow`.
Optional `sources` input for single-source refresh (e.g., `bayer` only).
+25 -5
View File
@@ -34,13 +34,33 @@ later co-host it.
## Configure the MCP server ## Configure the MCP server
```bash In production the MCP reaches this sidecar **over the Drawbar compose
export RERANK_URL=http://10.10.1.65:8082 network by service name**, not by host IP — `RERANK_URL` is baked into
# search_docs now reranks the hybrid pool through the GPU before returning the image as:
```
RERANK_URL=http://llama-rerank:8080
``` ```
In production (the MetaMCP-fronted Drawbar deploy), this is baked so the deployed `chem-mcp` service sets nothing. See
into the MCP server's container env. `deploy/docker-compose.yml`.
That only resolves if the `llama-rerank` container is attached to
`drawbar-backend_default`. If it is on the default bridge network
instead, the name resolves via public DNS to an unrelated IP and
connection-refuses — and `search_docs` falls back to dense+BM25
**silently**. Attach it with:
```bash
docker network connect drawbar-backend_default llama-rerank
```
For local dev outside that network, point at the published port
directly:
```bash
export RERANK_URL=http://10.10.1.65:8082
```
## Verify ## Verify