perf(epa_ppls): make the monthly refresh fit the runner's 3 h budget #3

Merged
claude merged 2 commits from fix/epa-ppls-fit-in-runner-budget into main 2026-09-01 16:40:50 -04:00
Contributor

The monthly refresh has failed every month since 2026-06-01, always at exactly 3:00:00, across two runner hosts and three act_runner versions. The job container runs as /bin/sleep 10800 — act_runner's default runner.timeout — so at 3 h it vanishes mid-step and the run dies with the misleading container "GITEA-ACTIONS-TASK-…" does not exist.

It was never going to fit: --force re-fetched all 11,415 candidate registrations monthly. Profiling the Sep 1 log (2,198 products in 2h44m before the kill):

  • ~7,300 discarded as not-row-crop at a 1.31 s median — essentially all of it the hard-coded REQUEST_DELAY_SECONDS = 1.1 sleep. ~2.8 h/month spent deciding to throw things away.
  • ~4,070 written at a 6.6 s median but a 15.0 s mean — single PDFs burning minutes (100-1262 = 689 s, 241-441 = 602 s).
  • Extrapolated full run: 15-20 h, 5-7× the cap.

Changes

  1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling at ~50 KB/s never trips them. download_pdf now streams with an overall 180 s deadline and 2 attempts (was 4). Verified live: the 33.8 MB 241-441 label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry — previously 600 s.
  2. Incremental refresh. A committed filter cache remembers not-row-crop verdicts (TTL 180 d with a deterministic ±30 d per-product jitter, so a cold run's verdicts do not all expire in the same month and resurrect this bug). Products on disk are re-fetched only when EPA reports a new label date or PDF URL — so dropping --force costs no freshness, unlike the old skip-if-exists check which never noticed a revision.
  3. Parallelism. --workers (default 6) behind a shared 5 req/sec ceiling, replacing the per-request sleeps.

Measured

300 products against a throwaway corpus: cold 4.6 min → warm 1.0 min, 0 errors (0 wrote / 80 unchanged / 220 filtered-cached). --force and --workers 1 both still exercised. Steady state should land ~15-20 min against the new 170-minute job timeout, which fails legibly instead of letting the container disappear.

The workflow commits scrape/state so the next run starts warm, but decides changed from the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex + image push.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR

The monthly refresh has failed every month since 2026-06-01, always at **exactly 3:00:00**, across two runner hosts and three act_runner versions. The job container runs as `/bin/sleep 10800` — act_runner's default `runner.timeout` — so at 3 h it vanishes mid-step and the run dies with the misleading `container "GITEA-ACTIONS-TASK-…" does not exist`. It was never going to fit: `--force` re-fetched all **11,415** candidate registrations monthly. Profiling the Sep 1 log (2,198 products in 2h44m before the kill): - ~7,300 discarded as not-row-crop at a 1.31 s median — essentially all of it the hard-coded `REQUEST_DELAY_SECONDS = 1.1` sleep. ~2.8 h/month spent deciding to throw things away. - ~4,070 written at a 6.6 s median but a **15.0 s mean** — single PDFs burning minutes (`100-1262` = 689 s, `241-441` = 602 s). - Extrapolated full run: **15-20 h**, 5-7× the cap. ### Changes 1. **Bound PDF downloads.** httpx timeouts are per-read, so a body trickling at ~50 KB/s never trips them. `download_pdf` now streams with an overall 180 s deadline and 2 attempts (was 4). Verified live: the 33.8 MB `241-441` label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry — previously 600 s. 2. **Incremental refresh.** A committed filter cache remembers not-row-crop verdicts (TTL 180 d with a deterministic ±30 d per-product jitter, so a cold run's verdicts do not all expire in the same month and resurrect this bug). Products on disk are re-fetched only when EPA reports a new label date or PDF URL — so dropping `--force` costs no freshness, unlike the old skip-if-exists check which never noticed a revision. 3. **Parallelism.** `--workers` (default 6) behind a shared 5 req/sec ceiling, replacing the per-request sleeps. ### Measured 300 products against a throwaway corpus: **cold 4.6 min → warm 1.0 min**, 0 errors (`0 wrote / 80 unchanged / 220 filtered-cached`). `--force` and `--workers 1` both still exercised. Steady state should land ~15-20 min against the new 170-minute job timeout, which fails legibly instead of letting the container disappear. The workflow commits `scrape/state` so the next run starts warm, but decides `changed` from the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex + image push. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
claude added 1 commit 2026-09-01 14:13:26 -04:00
The monthly refresh has failed every month since 2026-06-01, always at
exactly 3:00:00, on two runner hosts and three act_runner versions. The
job container is created as `/bin/sleep 10800` (act_runner's default
`runner.timeout`), so at 3 h it vanishes mid-step and the run dies with
the misleading `container "GITEA-ACTIONS-TASK-..." does not exist` /
`docker daemon ping ... context deadline exceeded`.

It was never going to fit. `--force` re-fetched all 11,415 candidate
registrations every month. Profiling the 2026-09-01 log (2,198 products
in 2 h 44 m before the kill):

  - ~7,300 discarded as not-row-crop at a 1.31 s median, essentially all
    of it the hard-coded REQUEST_DELAY_SECONDS=1.1 sleep — ~2.8 h/month
    spent deciding to throw things away.
  - ~4,070 written at a 6.6 s median but a 15.0 s MEAN: individual PDFs
    burned minutes (100-1262 = 689 s, 241-441 = 602 s).
  - Extrapolated full run: 15-20 h, 5-7x the cap.

Three changes, measured end-to-end against a throwaway corpus:

1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling
   in at ~50 KB/s never trips them and download_pdf() could hang as long
   as EPA kept dribbling. It now streams with an overall 180 s deadline
   and 2 attempts instead of 4. Verified live: the 33.8 MB 241-441 label
   tripped the deadline at 8.9 MB on attempt 1 and completed on the
   retry, where before it cost 600 s.

2. Make it incremental. A committed filter cache remembers not-row-crop
   verdicts (TTL 180 d, with a deterministic +/-30 d per-product jitter
   so a cold run's verdicts do not all expire in the same month and
   resurrect this bug). Products already on disk are re-downloaded only
   when EPA reports a new label acceptance date or PDF URL, so dropping
   --force costs no freshness — unlike the old skip-if-exists check,
   which never noticed a revision.

3. Parallelise. --workers (default 6) with a shared 5 req/sec ceiling,
   replacing the per-request sleeps: faster without being ruder.

Measured on 300 products: cold 4.6 min -> warm 1.0 min, 0 errors
(0 wrote / 80 unchanged / 220 filtered-cached). A steady-state monthly
refresh should land at ~15-20 min against a 170-minute job timeout,
which now fails legibly instead of letting the container disappear.

The workflow commits scrape/state so the next run starts warm, but
decides `changed` from the corpus paths alone — otherwise the cache
would fake a corpus diff every month and trigger a needless reindex
and image push.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
gitea-actions bot added 1 commit 2026-09-01 15:08:31 -04:00
claude merged commit 98842d1ed6 into main 2026-09-01 16:40:50 -04:00
claude deleted branch fix/epa-ppls-fit-in-runner-budget 2026-09-01 16:40:51 -04:00
Sign in to join this conversation.