perf(epa_ppls): make the monthly refresh fit the runner's 3 h budget
The monthly refresh has failed every month since 2026-06-01, always at
exactly 3:00:00, on two runner hosts and three act_runner versions. The
job container is created as `/bin/sleep 10800` (act_runner's default
`runner.timeout`), so at 3 h it vanishes mid-step and the run dies with
the misleading `container "GITEA-ACTIONS-TASK-..." does not exist` /
`docker daemon ping ... context deadline exceeded`.
It was never going to fit. `--force` re-fetched all 11,415 candidate
registrations every month. Profiling the 2026-09-01 log (2,198 products
in 2 h 44 m before the kill):
- ~7,300 discarded as not-row-crop at a 1.31 s median, essentially all
of it the hard-coded REQUEST_DELAY_SECONDS=1.1 sleep — ~2.8 h/month
spent deciding to throw things away.
- ~4,070 written at a 6.6 s median but a 15.0 s MEAN: individual PDFs
burned minutes (100-1262 = 689 s, 241-441 = 602 s).
- Extrapolated full run: 15-20 h, 5-7x the cap.
Three changes, measured end-to-end against a throwaway corpus:
1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling
in at ~50 KB/s never trips them and download_pdf() could hang as long
as EPA kept dribbling. It now streams with an overall 180 s deadline
and 2 attempts instead of 4. Verified live: the 33.8 MB 241-441 label
tripped the deadline at 8.9 MB on attempt 1 and completed on the
retry, where before it cost 600 s.
2. Make it incremental. A committed filter cache remembers not-row-crop
verdicts (TTL 180 d, with a deterministic +/-30 d per-product jitter
so a cold run's verdicts do not all expire in the same month and
resurrect this bug). Products already on disk are re-downloaded only
when EPA reports a new label acceptance date or PDF URL, so dropping
--force costs no freshness — unlike the old skip-if-exists check,
which never noticed a revision.
3. Parallelise. --workers (default 6) with a shared 5 req/sec ceiling,
replacing the per-request sleeps: faster without being ruder.
Measured on 300 products: cold 4.6 min -> warm 1.0 min, 0 errors
(0 wrote / 80 unchanged / 220 filtered-cached). A steady-state monthly
refresh should land at ~15-20 min against a 170-minute job timeout,
which now fails legibly instead of letting the container disappear.
The workflow commits scrape/state so the next run starts warm, but
decides `changed` from the corpus paths alone — otherwise the cache
would fake a corpus diff every month and trigger a needless reindex
and image push.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
This commit is contained in:
@@ -5,11 +5,13 @@ name: Monthly corpus refresh
|
||||
# reindex + image-push if the scrape produced no diff against the
|
||||
# committed corpus.
|
||||
#
|
||||
# Bayer takes ~30 min; EPA PPLS takes ~7 h with row-crop +
|
||||
# registrant filters. The whole monthly job is ~8-9 h end-to-end.
|
||||
# If that's too long for the runner you can:
|
||||
# - Run just one source: workflow_dispatch with sources="bayer"
|
||||
# - Limit EPA at the scraper: edit the step to add "--limit 5000"
|
||||
# Runtime budget: act_runner kills the job container at exactly 3 h
|
||||
# (it runs as `/bin/sleep 10800`), which is what killed every run from
|
||||
# 2026-06-01 through 2026-09-01. The EPA step is incremental and
|
||||
# parallel so a steady-state refresh is ~15-20 min: cached not-row-crop
|
||||
# verdicts cost no request, and a product already on disk is only
|
||||
# re-downloaded when EPA reports a new label date. `timeout-minutes`
|
||||
# below fails the job legibly before the container disappears.
|
||||
|
||||
on:
|
||||
schedule:
|
||||
@@ -42,6 +44,9 @@ env:
|
||||
jobs:
|
||||
refresh:
|
||||
runs-on: docker
|
||||
# Below act_runner's own 3 h container lifetime, so an overrun fails
|
||||
# as a timeout instead of "container ... does not exist".
|
||||
timeout-minutes: 170
|
||||
container:
|
||||
image: catthehacker/ubuntu:act-latest
|
||||
steps:
|
||||
@@ -80,9 +85,14 @@ jobs:
|
||||
|
||||
- name: Scrape EPA PPLS
|
||||
if: ${{ inputs.sources == '' || contains(inputs.sources, 'epa_ppls') }}
|
||||
# Row-crop + registrant filters keep this to ~16K PDFs / ~7h.
|
||||
# Pass --no-row-crop-filter or --no-registrant-filter to broaden.
|
||||
run: python -m scrape.runner --source epa_ppls --force
|
||||
# Deliberately NOT --force: that re-downloaded all ~11.4K candidate
|
||||
# registrations every month (~15-20 h of work) and never finished.
|
||||
# Without it the run is incremental — one cheap API call per product,
|
||||
# a PDF only when the label date actually moved — and the committed
|
||||
# filter cache skips the ~7.3K non-row-crop products outright.
|
||||
# Workers share one 5 req/sec ceiling, so this is faster without
|
||||
# being ruder. Local full re-fetch: add --force --no-filter-cache.
|
||||
run: python -m scrape.runner --source epa_ppls --workers 6
|
||||
|
||||
# ---- Commit corpus changes + retry-on-race -----------------
|
||||
- name: Commit corpus changes (if any)
|
||||
@@ -90,13 +100,21 @@ jobs:
|
||||
run: |
|
||||
git config user.name "crop-chem-docs-refresh"
|
||||
git config user.email "[email protected]"
|
||||
git add sources.json corpus
|
||||
if git diff --cached --quiet; then
|
||||
# The filter-verdict cache is committed so next month starts warm,
|
||||
# but it changes on every run — only a real corpus diff may trigger
|
||||
# the reindex + image build, so `changed` is decided on the corpus
|
||||
# paths alone.
|
||||
git add sources.json corpus scrape/state
|
||||
if git diff --cached --quiet -- sources.json corpus; then
|
||||
echo "no corpus changes — skipping reindex and image build"
|
||||
echo "changed=false" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "changed=true" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
if git diff --cached --quiet; then
|
||||
echo "nothing staged — no commit"
|
||||
exit 0
|
||||
fi
|
||||
echo "changed=true" >> "$GITHUB_OUTPUT"
|
||||
ts=$(date -u +"%Y-%m-%dT%H:%MZ")
|
||||
n_bayer=$(find corpus/bayer -name '*.json' 2>/dev/null | wc -l)
|
||||
n_epa=$(find corpus/epa_ppls -name '*.json' 2>/dev/null | wc -l)
|
||||
|
||||
Reference in New Issue
Block a user