The monthly refresh has failed every month since 2026-06-01, always at exactly 3:00:00, across two runner hosts and three act_runner versions. The job container runs as /bin/sleep 10800 — act_runner's default runner.timeout — so at 3 h it vanishes mid-step and the run dies with the misleading container "GITEA-ACTIONS-TASK-…" does not exist.
It was never going to fit: --force re-fetched all 11,415 candidate registrations monthly. Profiling the Sep 1 log (2,198 products in 2h44m before the kill):
~7,300 discarded as not-row-crop at a 1.31 s median — essentially all of it the hard-coded REQUEST_DELAY_SECONDS = 1.1 sleep. ~2.8 h/month spent deciding to throw things away.
~4,070 written at a 6.6 s median but a 15.0 s mean — single PDFs burning minutes (100-1262 = 689 s, 241-441 = 602 s).
Extrapolated full run: 15-20 h, 5-7× the cap.
Changes
Bound PDF downloads. httpx timeouts are per-read, so a body trickling at ~50 KB/s never trips them. download_pdf now streams with an overall 180 s deadline and 2 attempts (was 4). Verified live: the 33.8 MB 241-441 label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry — previously 600 s.
Incremental refresh. A committed filter cache remembers not-row-crop verdicts (TTL 180 d with a deterministic ±30 d per-product jitter, so a cold run's verdicts do not all expire in the same month and resurrect this bug). Products on disk are re-fetched only when EPA reports a new label date or PDF URL — so dropping --force costs no freshness, unlike the old skip-if-exists check which never noticed a revision.
Parallelism.--workers (default 6) behind a shared 5 req/sec ceiling, replacing the per-request sleeps.
Measured
300 products against a throwaway corpus: cold 4.6 min → warm 1.0 min, 0 errors (0 wrote / 80 unchanged / 220 filtered-cached). --force and --workers 1 both still exercised. Steady state should land ~15-20 min against the new 170-minute job timeout, which fails legibly instead of letting the container disappear.
The workflow commits scrape/state so the next run starts warm, but decides changed from the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex + image push.
The monthly refresh has failed every month since 2026-06-01, always at **exactly 3:00:00**, across two runner hosts and three act_runner versions. The job container runs as `/bin/sleep 10800` — act_runner's default `runner.timeout` — so at 3 h it vanishes mid-step and the run dies with the misleading `container "GITEA-ACTIONS-TASK-…" does not exist`.
It was never going to fit: `--force` re-fetched all **11,415** candidate registrations monthly. Profiling the Sep 1 log (2,198 products in 2h44m before the kill):
- ~7,300 discarded as not-row-crop at a 1.31 s median — essentially all of it the hard-coded `REQUEST_DELAY_SECONDS = 1.1` sleep. ~2.8 h/month spent deciding to throw things away.
- ~4,070 written at a 6.6 s median but a **15.0 s mean** — single PDFs burning minutes (`100-1262` = 689 s, `241-441` = 602 s).
- Extrapolated full run: **15-20 h**, 5-7× the cap.
### Changes
1. **Bound PDF downloads.** httpx timeouts are per-read, so a body trickling at ~50 KB/s never trips them. `download_pdf` now streams with an overall 180 s deadline and 2 attempts (was 4). Verified live: the 33.8 MB `241-441` label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry — previously 600 s.
2. **Incremental refresh.** A committed filter cache remembers not-row-crop verdicts (TTL 180 d with a deterministic ±30 d per-product jitter, so a cold run's verdicts do not all expire in the same month and resurrect this bug). Products on disk are re-fetched only when EPA reports a new label date or PDF URL — so dropping `--force` costs no freshness, unlike the old skip-if-exists check which never noticed a revision.
3. **Parallelism.** `--workers` (default 6) behind a shared 5 req/sec ceiling, replacing the per-request sleeps.
### Measured
300 products against a throwaway corpus: **cold 4.6 min → warm 1.0 min**, 0 errors (`0 wrote / 80 unchanged / 220 filtered-cached`). `--force` and `--workers 1` both still exercised. Steady state should land ~15-20 min against the new 170-minute job timeout, which fails legibly instead of letting the container disappear.
The workflow commits `scrape/state` so the next run starts warm, but decides `changed` from the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex + image push.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
The monthly refresh has failed every month since 2026-06-01, always at
exactly 3:00:00, on two runner hosts and three act_runner versions. The
job container is created as `/bin/sleep 10800` (act_runner's default
`runner.timeout`), so at 3 h it vanishes mid-step and the run dies with
the misleading `container "GITEA-ACTIONS-TASK-..." does not exist` /
`docker daemon ping ... context deadline exceeded`.
It was never going to fit. `--force` re-fetched all 11,415 candidate
registrations every month. Profiling the 2026-09-01 log (2,198 products
in 2 h 44 m before the kill):
- ~7,300 discarded as not-row-crop at a 1.31 s median, essentially all
of it the hard-coded REQUEST_DELAY_SECONDS=1.1 sleep — ~2.8 h/month
spent deciding to throw things away.
- ~4,070 written at a 6.6 s median but a 15.0 s MEAN: individual PDFs
burned minutes (100-1262 = 689 s, 241-441 = 602 s).
- Extrapolated full run: 15-20 h, 5-7x the cap.
Three changes, measured end-to-end against a throwaway corpus:
1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling
in at ~50 KB/s never trips them and download_pdf() could hang as long
as EPA kept dribbling. It now streams with an overall 180 s deadline
and 2 attempts instead of 4. Verified live: the 33.8 MB 241-441 label
tripped the deadline at 8.9 MB on attempt 1 and completed on the
retry, where before it cost 600 s.
2. Make it incremental. A committed filter cache remembers not-row-crop
verdicts (TTL 180 d, with a deterministic +/-30 d per-product jitter
so a cold run's verdicts do not all expire in the same month and
resurrect this bug). Products already on disk are re-downloaded only
when EPA reports a new label acceptance date or PDF URL, so dropping
--force costs no freshness — unlike the old skip-if-exists check,
which never noticed a revision.
3. Parallelise. --workers (default 6) with a shared 5 req/sec ceiling,
replacing the per-request sleeps: faster without being ruder.
Measured on 300 products: cold 4.6 min -> warm 1.0 min, 0 errors
(0 wrote / 80 unchanged / 220 filtered-cached). A steady-state monthly
refresh should land at ~15-20 min against a 170-minute job timeout,
which now fails legibly instead of letting the container disappear.
The workflow commits scrape/state so the next run starts warm, but
decides `changed` from the corpus paths alone — otherwise the cache
would fake a corpus diff every month and trigger a needless reindex
and image push.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The monthly refresh has failed every month since 2026-06-01, always at exactly 3:00:00, across two runner hosts and three act_runner versions. The job container runs as
/bin/sleep 10800— act_runner's defaultrunner.timeout— so at 3 h it vanishes mid-step and the run dies with the misleadingcontainer "GITEA-ACTIONS-TASK-…" does not exist.It was never going to fit:
--forcere-fetched all 11,415 candidate registrations monthly. Profiling the Sep 1 log (2,198 products in 2h44m before the kill):REQUEST_DELAY_SECONDS = 1.1sleep. ~2.8 h/month spent deciding to throw things away.100-1262= 689 s,241-441= 602 s).Changes
download_pdfnow streams with an overall 180 s deadline and 2 attempts (was 4). Verified live: the 33.8 MB241-441label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry — previously 600 s.--forcecosts no freshness, unlike the old skip-if-exists check which never noticed a revision.--workers(default 6) behind a shared 5 req/sec ceiling, replacing the per-request sleeps.Measured
300 products against a throwaway corpus: cold 4.6 min → warm 1.0 min, 0 errors (
0 wrote / 80 unchanged / 220 filtered-cached).--forceand--workers 1both still exercised. Steady state should land ~15-20 min against the new 170-minute job timeout, which fails legibly instead of letting the container disappear.The workflow commits
scrape/stateso the next run starts warm, but decideschangedfrom the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex + image push.🤖 Generated with Claude Code
https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR
The monthly refresh has failed every month since 2026-06-01, always at exactly 3:00:00, on two runner hosts and three act_runner versions. The job container is created as `/bin/sleep 10800` (act_runner's default `runner.timeout`), so at 3 h it vanishes mid-step and the run dies with the misleading `container "GITEA-ACTIONS-TASK-..." does not exist` / `docker daemon ping ... context deadline exceeded`. It was never going to fit. `--force` re-fetched all 11,415 candidate registrations every month. Profiling the 2026-09-01 log (2,198 products in 2 h 44 m before the kill): - ~7,300 discarded as not-row-crop at a 1.31 s median, essentially all of it the hard-coded REQUEST_DELAY_SECONDS=1.1 sleep — ~2.8 h/month spent deciding to throw things away. - ~4,070 written at a 6.6 s median but a 15.0 s MEAN: individual PDFs burned minutes (100-1262 = 689 s, 241-441 = 602 s). - Extrapolated full run: 15-20 h, 5-7x the cap. Three changes, measured end-to-end against a throwaway corpus: 1. Bound PDF downloads. httpx timeouts are per-read, so a body trickling in at ~50 KB/s never trips them and download_pdf() could hang as long as EPA kept dribbling. It now streams with an overall 180 s deadline and 2 attempts instead of 4. Verified live: the 33.8 MB 241-441 label tripped the deadline at 8.9 MB on attempt 1 and completed on the retry, where before it cost 600 s. 2. Make it incremental. A committed filter cache remembers not-row-crop verdicts (TTL 180 d, with a deterministic +/-30 d per-product jitter so a cold run's verdicts do not all expire in the same month and resurrect this bug). Products already on disk are re-downloaded only when EPA reports a new label acceptance date or PDF URL, so dropping --force costs no freshness — unlike the old skip-if-exists check, which never noticed a revision. 3. Parallelise. --workers (default 6) with a shared 5 req/sec ceiling, replacing the per-request sleeps: faster without being ruder. Measured on 300 products: cold 4.6 min -> warm 1.0 min, 0 errors (0 wrote / 80 unchanged / 220 filtered-cached). A steady-state monthly refresh should land at ~15-20 min against a 170-minute job timeout, which now fails legibly instead of letting the container disappear. The workflow commits scrape/state so the next run starts warm, but decides `changed` from the corpus paths alone — otherwise the cache would fake a corpus diff every month and trigger a needless reindex and image push. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01FnVuG79cYPcRLTp4pC8ujR