perf(epa_ppls): make the monthly refresh fit the runner's 3 h budget (#3)
Image rebuild (skip scrape) / build (push) Successful in 1h46m17s
Image rebuild (skip scrape) / build (push) Successful in 1h46m17s
Co-authored-by: claude <[email protected]>
This commit was merged in pull request #3.
This commit is contained in:
@@ -59,9 +59,14 @@ pip install -r requirements.txt
|
||||
# Sample-scrape to verify wiring:
|
||||
python -m scrape.runner --source bayer --limit 5
|
||||
|
||||
# Full refresh (be polite — bayer is small, epa_ppls is hours):
|
||||
# Incremental refresh — what CI runs monthly (~15-20 min for epa_ppls):
|
||||
# one cheap API call per product, a PDF only when the label date moved,
|
||||
# and cached "not row-crop" verdicts skipped outright.
|
||||
python -m scrape.runner --source bayer --force
|
||||
python -m scrape.runner --source epa_ppls --force
|
||||
python -m scrape.runner --source epa_ppls --workers 6
|
||||
|
||||
# Full re-fetch of every label (hours — does NOT fit CI's 3 h runner cap):
|
||||
python -m scrape.runner --source epa_ppls --force --no-filter-cache
|
||||
|
||||
# Rebuild Chroma + BM25:
|
||||
OLLAMA_URL=http://192.168.0.125:11434 PRODUCT_NAME=crop_chem \
|
||||
|
||||
Reference in New Issue
Block a user