The monthly refresh has failed every month since 2026-06-01, always at
exactly 3:00:00, on two runner hosts and three act_runner versions. The
job container is created as `/bin/sleep 10800` (act_runner's default
`runner.timeout`), so at 3 h it vanishes mid-step and the run dies with
the misleading `container "GITEA-ACTIONS-TASK-..." does not exist` /
`docker daemon ping ... context deadline exceeded`.
It was never going to fit. `--force` re-fetched all 11,415 candidate
registrations every month. Profiling the 2026-09-01 log (2,198 products
in 2 h 44 m before the kill):
- ~7,300 discarded as not-row-crop at a 1.31 s median, essentially all
of it the hard-coded REQUEST_DELAY_SECONDS=1.1 sleep — ~2.8 h/month
spent deciding to throw things away.
- ~4,070 written at a 6.6 s median but a 15.0 s MEAN: individual PDFs
burned minutes (100-1262 = 689 s, 241-441 = 602 s).
- Extrapolated full run: 15-20 h, 5-7x the cap.
Three changes, measured end-to-end against a throwaway corpus:
1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling
in at ~50 KB/s never trips them and download_pdf() could hang as long
as EPA kept dribbling. It now streams with an overall 180 s deadline
and 2 attempts instead of 4. Verified live: the 33.8 MB 241-441 label
tripped the deadline at 8.9 MB on attempt 1 and completed on the
retry, where before it cost 600 s.
2. Make it incremental. A committed filter cache remembers not-row-crop
verdicts (TTL 180 d, with a deterministic +/-30 d per-product jitter
so a cold run's verdicts do not all expire in the same month and
resurrect this bug). Products already on disk are re-downloaded only
when EPA reports a new label acceptance date or PDF URL, so dropping
--force costs no freshness — unlike the old skip-if-exists check,
which never noticed a revision.
3. Parallelise. --workers (default 6) with a shared 5 req/sec ceiling,
replacing the per-request sleeps: faster without being ruder.
Measured on 300 products: cold 4.6 min -> warm 1.0 min, 0 errors
(0 wrote / 80 unchanged / 220 filtered-cached). A steady-state monthly
refresh should land at ~15-20 min against a 170-minute job timeout,
which now fails legibly instead of letting the container disappear.
The workflow commits scrape/state so the next run starts warm, but
decides `changed` from the corpus paths alone — otherwise the cache
would fake a corpus diff every month and trigger a needless reindex
and image push.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
The README had never been customized after cloning the
docs-mcp-template — title said "docs-mcp-template" and it read as
the template's generic introduction with no mention of EPA PPLS,
the Bayer scraper, the ~4k label corpus, or the production deploy.
Replace with a crop-chem-docs-specific README that covers:
- Corpus inventory: 4,159 indexed pages (91 Bayer + 4,068 EPA PPLS)
- MCP tool catalog with crop_chem_api_lessons specifics
- Eval baseline from eval/results/with_rerank.md showing
hybrid+rerank wins (MRR 0.672) over BM25-only (0.544) and that
hybrid-without-rerank actively HURTS (0.114) — same pattern
seed-mcp found independently
- Note that the deployed rerank was silently failing through
2026-05-25 due to the llama-rerank Docker network gotcha;
fixed and re-running eval is on the followup list
- Quick-start commands
- Repo layout reference
- Infrastructure: registry, embedder pool, shared llama-rerank
sidecar, PRODUCT_NAME=crop_chem
- Cross-link to the sibling seed-mcp project
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>