Files
seed-mcp/scrape/sources/becks_products.py
T
justin 75f714b454 Phase 4-5: deployable container + corpus snapshot + CI fixes
deploy/docker-compose.yml — replace <product>/<registry> placeholders
with concrete values for Drawbar's stack:
- image: git.jpaul.io/justin/seed-mcp:latest (CF tunnel for pulls; CI
  pushes via LAN 192.168.0.2:1234 to avoid 100 MB body cap)
- container_name: seed-mcp
- port 8001:8000 (8001 host-side to not collide with crop-chem-docs
  on 8000)
- PRODUCT_NAME=crop_seed, hybrid search enabled, stateless HTTP
- llama-rerank shared with crop-chem-docs (NOT redefined here —
  expected to already be in Drawbar's parent compose network)
- networks.drawbar-mcp external: true so seed-mcp joins the existing
  cross-MCP shared network

.gitignore — corpus/ is now COMMITTED, not ignored. The monthly
refresh workflow scrapes and commits corpus changes; the image-only
workflow rebuilds indexes from the committed corpus. Allowing the
corpus to flow through git means the :corpus-YYYY.MM.DD image tag
pins to a specific seed-catalog snapshot. chroma/ and bm25/ remain
ignored — those are deterministically derived from corpus.

Initial committed snapshot: 614 varieties.
- bayer_seeds: 475 (DEKALB 288 + Asgrow 102 + WestBred 85)
- golden_harvest: 139 (Syngenta corn + soy; 36 sitemap URLs
  302-redirected = discontinued)

rag/chunk.py — normalize brand and crop to uppercase/lowercase in
Chroma metadata so cross-vendor brand-filter lookups don't break on
casing inconsistency (Bayer stores "DEKALB", Golden Harvest stores
"Golden Harvest"; _build_where uppercases user-supplied brand which
matched the former but not the latter pre-fix). Sidecar JSON keeps
original casing for display.

Stub scrapers (nk, agripro, becks_pfr, becks_products) — change
return code from 2 to 0 so the monthly-refresh CI workflow doesn't
fail on deferred sources. Real implementations will return 0 on
success / 1 on failure when they ship.

Smoke-tested cross-vendor retrieval against the 614-chunk index:
- list_versions shows both vendors with correct facet counts
- broad "corn hybrid 100 RM" query returns both DEKALB and Golden
  Harvest hits in top 5
- brand='Golden Harvest' filter returns 3 GH-only varieties
- variety-code prefilter still works (E085Z5 → top hit on GH)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 13:40:05 -04:00

50 lines
1.7 KiB
Python

"""Beck's product catalog scraper (identity-only until SeedIQ XHR sniff lands).
Source: Same public Sanity GROQ API as ``becks_pfr`` (no auth).
Expected count: ~860 products (corn + soy + wheat).
Current limitation: Beck's exposes IDENTITY fields publicly (product
name, RM/MG, basic trait stack) but routes the AGRONOMIC + DISEASE
ratings through their SeedIQ application, which is gated behind a
browser session cookie. The public Sanity records do not include
ratings.
What we CAN ship without SeedIQ:
- Product identity for confirmation ("yes Beck's has hybrid X at RM 112")
- RM (corn) / MG (soy) / class (wheat)
- Trait stack
- Basic descriptive text
What needs the SeedIQ XHR endpoint (BLOCKED on user sniff):
- Disease ratings (GLS, NCLB, Goss's, etc.)
- Agronomic ratings (standability, drought, etc.)
- Regional recommendations
For now this scraper is DEFERRED. Run when:
- User captures the SeedIQ XHR URL + cookie/header pattern from
browser dev tools, OR
- We decide to ship Beck's as identity-only and let the LLM say
"Beck's has this hybrid; ask your Beck's rep for full agronomic
ratings" (less useful but avoids the empty-data UX).
Yellow verdict in sources.json reflects this — ``--all`` skips it.
TODO: implement (deferred).
"""
from __future__ import annotations
import sys
def main(argv: list[str] | None = None) -> int:
print("becks_products: deferred — SeedIQ XHR sniff required for ratings, run only if user has captured the endpoint",
file=sys.stderr)
# Return 0 so the monthly CI workflow doesn't fail when this
# source is listed but not yet implemented (and may never be,
# if SeedIQ gates persist).
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))