Files
crop-chem-docs/.gitea/workflows/refresh.yml
T
claudeandClaude Opus 5 83af1ce3f0 deps: migrate to mcp 2.x (lift the <2 pin)
mcp 2.0.0 removed `mcp.server.fastmcp`; the `<2` ceiling in 23c7a3d
stopped the bleeding but left this server on the previous major.
Ports to the 2.x API, following the same shape seed-mcp shipped.

- requirements.txt: `mcp[fastmcp]>=1.0.0,<2` -> `mcp>=2,<3` (2.x has
  no [fastmcp] extra)
- server.py: FastMCP -> mcp.server.mcpserver.MCPServer; `mcp.settings`
  is gone, so host/port/stateless_http/transport_security are now
  run() kwargs
- stateless_http is gated to streamable-http. Under 1.x it was a
  constructor arg and so applied to sse too; sse is a dev-only path
  here, and this matches seed-mcp.
- CI: both workflows now run `python -c "import docs_mcp.server"`
  against the built image before pushing it. This is the durable
  half — a green build can no longer ship a non-importing image.
- CLAUDE.md / README.md: correct the FastMCP references. PLAN.md is
  left alone; it tracks the upstream template, not this repo.

Verification (mcp 1.27.1 -> 2.2.0):
- tools/list dumped in wire format is byte-for-byte identical,
  md5 cea0326f59b55abbbf47c9679252d38f both sides. All 5 tools
  (search_docs, get_page, list_versions, corpus_status,
  crop_chem_api_lessons) unchanged, so routing is unaffected.
- streamable-http: `initialize` -> 200, no mcp-session-id header
  (stateless_http confirmed active); tools/list and a real
  tools/call over the wire both succeed.
- stdio: initialize + tools/list both fine.
- Image built; the new CI smoke command passes against it.
  corpus_status reads the baked indexes (4,164 labels / 216,467
  chunks) and a live `search_docs` returns hits with mode=hybrid-rrf
  against Ollama on .0.125.

No eval numbers: this touches transport and packaging only. The
chunker, embedder, Chroma/BM25 stores, RRF and the reranker are all
untouched, and the live search above confirms retrieval still runs.

Note: `serverInfo.version` now reports "" instead of the SDK version
(2.x behaviour for an unversioned server). Cosmetic, but visible to
clients. httpx2 + opentelemetry-api come in as 2.x deps and coexist
with the existing httpx 0.28.1 pin, as expected.

Closes #4

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FFBDnRWHispovmJVK9rXc9
2026-09-10 19:47:32 -04:00

202 lines
8.6 KiB
YAML

name: Monthly corpus refresh
# Runs the full pipeline: scrape all sources → rebuild indexes →
# push image. Cron'd once a month (1st @ 06:00 UTC). Skip the
# reindex + image-push if the scrape produced no diff against the
# committed corpus.
#
# Runtime budget: act_runner kills the job container at exactly 3 h
# (it runs as `/bin/sleep 10800`), which is what killed every run from
# 2026-06-01 through 2026-09-01. The EPA step is incremental and
# parallel so a steady-state refresh is ~15-20 min: cached not-row-crop
# verdicts cost no request, and a product already on disk is only
# re-downloaded when EPA reports a new label date. `timeout-minutes`
# below fails the job legibly before the container disappears.
on:
schedule:
- cron: "0 6 1 * *" # 1st of each month, 06:00 UTC
workflow_dispatch:
inputs:
force_build:
description: "Rebuild indexes + push image even if corpus is unchanged"
type: boolean
default: false
sources:
description: "Sources to scrape (comma-separated, blank = all)"
type: string
default: ""
env:
# Self-hosted Gitea registry on the same LAN as the runner.
REGISTRY_PUSH: 192.168.0.2:1234
REGISTRY_PULL: git.jpaul.io
IMAGE: ${{ github.repository }}
# Embedder pool for the reindex step. Two Ollama instances on the
# Gitea/runner host (one per GPU) + the Windows Ollama. Trashpanda's
# Ollama is production-shared; CI doesn't hit it.
OLLAMA_URL: http://192.168.0.2:11434,http://192.168.0.2:11435,http://192.168.0.125:11434
EMBED_MODEL: nomic-embed-text
PRODUCT_NAME: crop_chem
jobs:
refresh:
runs-on: docker
# Below act_runner's own 3 h container lifetime, so an overrun fails
# as a timeout instead of "container ... does not exist".
timeout-minutes: 170
container:
image: catthehacker/ubuntu:act-latest
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Verify image name resolves
# ${{ github.event.repository.name }} silently resolves to EMPTY on
# schedule events (Gitea >=1.27), which built the tag
# "192.168.0.2:1234/<owner>/:latest" and failed the build only AFTER
# the full scrape. IMAGE now comes from github.repository; this step
# fails in seconds if it ever goes empty again.
run: |
echo "IMAGE=${IMAGE}"
case "${IMAGE}" in
*/?*) ;;
*) echo "ERROR: IMAGE did not resolve to owner/repo"; exit 1 ;;
esac
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install dependencies
run: |
python -m pip install -q --upgrade pip
python -m pip install -q -r requirements.txt
# ---- Phase 1: scrape ---------------------------------------
- name: Scrape Bayer
if: ${{ inputs.sources == '' || contains(inputs.sources, 'bayer') }}
run: python -m scrape.runner --source bayer --force
- name: Scrape EPA PPLS
if: ${{ inputs.sources == '' || contains(inputs.sources, 'epa_ppls') }}
# Deliberately NOT --force: that re-downloaded all ~11.4K candidate
# registrations every month (~15-20 h of work) and never finished.
# Without it the run is incremental — one cheap API call per product,
# a PDF only when the label date actually moved — and the committed
# filter cache skips the ~7.3K non-row-crop products outright.
# Workers share one 5 req/sec ceiling, so this is faster without
# being ruder. Local full re-fetch: add --force --no-filter-cache.
run: python -m scrape.runner --source epa_ppls --workers 6
# ---- Commit corpus changes + retry-on-race -----------------
- name: Commit corpus changes (if any)
id: commit
run: |
git config user.name "crop-chem-docs-refresh"
git config user.email "[email protected]"
# The filter-verdict cache is committed so next month starts warm,
# but it changes on every run — only a real corpus diff may trigger
# the reindex + image build, so `changed` is decided on the corpus
# paths alone.
git add sources.json corpus scrape/state
if git diff --cached --quiet -- sources.json corpus; then
echo "no corpus changes — skipping reindex and image build"
echo "changed=false" >> "$GITHUB_OUTPUT"
else
echo "changed=true" >> "$GITHUB_OUTPUT"
fi
if git diff --cached --quiet; then
echo "nothing staged — no commit"
exit 0
fi
ts=$(date -u +"%Y-%m-%dT%H:%MZ")
n_bayer=$(find corpus/bayer -name '*.json' 2>/dev/null | wc -l)
n_epa=$(find corpus/epa_ppls -name '*.json' 2>/dev/null | wc -l)
git commit -m "monthly refresh: ${ts} — bayer=${n_bayer} epa_ppls=${n_epa}"
attempt=1
while [ $attempt -le 3 ]; do
if git push; then
echo "pushed corpus changes (attempt $attempt)"
break
fi
if [ $attempt -eq 3 ]; then
echo "push still failing after 3 attempts"; exit 1
fi
git fetch origin main
git rebase origin/main || { echo "rebase conflict"; exit 1; }
attempt=$((attempt + 1))
done
# ---- Rebuild Chroma + BM25 ---------------------------------
- name: Rebuild indexes
if: steps.commit.outputs.changed == 'true' || inputs.force_build == true
run: python -m rag.index --rebuild
# ---- Build & push image ------------------------------------
- name: Log in to Gitea container registry
if: steps.commit.outputs.changed == 'true' || inputs.force_build == true
run: echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${REGISTRY_PUSH}" -u "${{ github.repository_owner }}" --password-stdin
- name: Build & push image
if: steps.commit.outputs.changed == 'true' || inputs.force_build == true
# Tags: :latest (Watchtower target), :<sha12> (rollback pin),
# :corpus-<YYYY.MM.DD> (links image to corpus version so
# Drawbar can pin to a specific corpus snapshot).
run: |
SHA_TAG=$(echo "$GITHUB_SHA" | cut -c1-12)
CORPUS_TAG="corpus-$(date -u +%Y.%m.%d)"
docker build \
-t "${REGISTRY_PUSH}/${IMAGE}:latest" \
-t "${REGISTRY_PUSH}/${IMAGE}:${SHA_TAG}" \
-t "${REGISTRY_PUSH}/${IMAGE}:${CORPUS_TAG}" \
.
# Smoke: the image must at least import its server module before we
# publish it. mcp 2.x renamed mcp.server.fastmcp -> mcp.server.mcpserver,
# so an unpinned dep produced a green build that crash-looped in prod.
# Import is side-effect-free here (lazy singletons), so this needs no
# Ollama/Chroma.
docker run --rm --entrypoint python \
"${REGISTRY_PUSH}/${IMAGE}:latest" -c "import docs_mcp.server"
docker push "${REGISTRY_PUSH}/${IMAGE}:latest"
docker push "${REGISTRY_PUSH}/${IMAGE}:${SHA_TAG}"
docker push "${REGISTRY_PUSH}/${IMAGE}:${CORPUS_TAG}"
- name: Link container package to this repo
if: steps.commit.outputs.changed == 'true' || inputs.force_build == true
env:
GITEA_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
OWNER="${{ github.repository_owner }}"
PKG="${GITHUB_REPOSITORY##*/}"
BODY=$(mktemp)
CODE=$(curl -sS -o "$BODY" -w "%{http_code}" -X POST \
-H "Authorization: token ${GITEA_TOKEN}" \
"https://${REGISTRY_PULL}/api/v1/packages/${OWNER}/container/${PKG}/-/link/${PKG}")
echo "link http=$CODE body=$(cat "$BODY")"
case "$CODE" in
201) echo "linked package to ${OWNER}/${PKG}" ;;
400) echo "already linked — ok" ;;
*) echo "unexpected status $CODE"; exit 1 ;;
esac
- name: Prune old container versions
# GC requires broader scope than REGISTRY_TOKEN's push perms
# (HTTP 403 on /packages/.../versions). Non-critical housekeeping.
# TODO: issue separate PAT with admin:package scope.
if: steps.commit.outputs.changed == 'true' || inputs.force_build == true
continue-on-error: true
env:
GITEA_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
python scripts/registry_gc.py \
--owner "${{ github.repository_owner }}" \
--package "${GITHUB_REPOSITORY##*/}" \
--keep-days 180 \
--keep-latest 6