perf(epa_ppls): make the monthly refresh fit the runner's 3 h budget (#3)
Image rebuild (skip scrape) / build (push) Successful in 1h46m17s

Co-authored-by: claude <[email protected]>
This commit was merged in pull request #3.
This commit is contained in:
2026-09-01 16:40:48 -04:00
committed by claude
parent 0f296a0ec4
commit 98842d1ed6
316 changed files with 129846 additions and 84529 deletions
+7 -2
View File
@@ -59,9 +59,14 @@ pip install -r requirements.txt
# Sample-scrape to verify wiring:
python -m scrape.runner --source bayer --limit 5
# Full refresh (be polite — bayer is small, epa_ppls is hours):
# Incremental refresh — what CI runs monthly (~15-20 min for epa_ppls):
# one cheap API call per product, a PDF only when the label date moved,
# and cached "not row-crop" verdicts skipped outright.
python -m scrape.runner --source bayer --force
python -m scrape.runner --source epa_ppls --force
python -m scrape.runner --source epa_ppls --workers 6
# Full re-fetch of every label (hours — does NOT fit CI's 3 h runner cap):
python -m scrape.runner --source epa_ppls --force --no-filter-cache
# Rebuild Chroma + BM25:
OLLAMA_URL=http://192.168.0.125:11434 PRODUCT_NAME=crop_chem \