seed-mcp

Author	SHA1	Message	Date
claude	54094a0d43	Add university-extension variety trials: Illinois VT + Iowa ICPT + Ohio OCPT (+123 trial docs) Independent third-party performance data — land-grant programs that test every entered brand side-by-side with replication + LSD stats. This is the legitimate way to get Pioneer / DEKALB / Brevant / Channel performance the corpus can't scrape directly (data_type=trial, results[] shape; falls through the trial chunker). - illinois_vt_trials (30 docs, 1,392 rows) — U of Illinois VT. Per-region XLSX (openpyxl), corn + soy + WHEAT, 2024+2025. Rich per-site agronomic metadata; corn-following-corn vs -soybean kept distinct. - iowa_icpt_trials (24 docs, 674 rows) — Iowa State ICPT. ASP.NET GridView (viewstate postback for year/district), corn + soy by district x season. - ohio_ocpt_trials (69 docs, 4,647 rows) — OSU/CFAES OCPT. Report PDF (pdfplumber; per-site column groups split by header Yield-token count + x-coord footnote bucketing), corn + soy per site, 2024+2025. 91 distinct seed brands across the three; majors confirmed present in the independent rankings: DEKALB 395, Golden Harvest 249, Channel 241, NK 212, Xitavo 135, LG 103, Pioneer 88, Asgrow 59. (A brand only appears where it ENTERED a given program — e.g. Brevant not in Iowa, DEKALB/Channel not in Illinois — true negatives, not parse gaps.) - rag/chunk.py: gated `include_region` on _render_gh_plot_chunk; the 3 university sources route through it so the region/district is in the embedded chunk + labeled "variety trial (cross-vendor, independent third-party)". Existing plot sources (gh/lg/agrigold/proharvest) unchanged. - requirements.txt: openpyxl (Illinois XLSX; scrape-time only). - sources.json + README/CLAUDE/lessons: registered + attributed; lessons trial-data + Pioneer entries updated (Pioneer/DEKALB performance now available indirectly via these trials). Validation: all 123 chunk via rag.chunk.chunks_from_trial (0 errors), 0 out-of-range yields, 0 dup keys. Public land-grant data; attribution recorded in each tos_note. CI rebuilds the index from the committed corpus.	2026-06-10 08:35:50 -04:00
claude	0bac06b7b6	Add RobSeeCo (Rob-See-Co + Innotech): 130 corn/soy varieties from the seed-guide PDF (#18 ) Image rebuild (skip scrape) / build (push) Successful in 4m48s Details Co-authored-by: claude <claude@jpaul.io> Co-committed-by: claude <claude@jpaul.io>	2026-06-09 23:29:38 -04:00
claude	84ad2b1de6	Add 4 independent seed brands: Latham + Stine + 1st Choice + Burrus (+623 varieties) (#17 ) Image rebuild (skip scrape) / build (push) Successful in 4m44s Details Co-authored-by: claude <claude@jpaul.io> Co-committed-by: claude <claude@jpaul.io>	2026-06-04 21:58:07 -04:00
claude	22e8092faf	Add ProHarvest Seeds: 119 varieties + 161 cross-vendor plot reports (#16 ) Image rebuild (skip scrape) / build (push) Successful in 5m46s Details Co-authored-by: claude <claude@jpaul.io> Co-committed-by: claude <claude@jpaul.io>	2026-06-04 21:05:30 -04:00
justin	d65c7d0d67	README: reflect deployed state — 5,073 chunks, eval numbers, 6 tools The scaffold-era README was out of sync with the shipped product: - Vendor counts stale (recon estimates, not actual deployed counts) - Trial data sources (gh_plot_reports + agripro_trials) entirely unmentioned - Tool list listed `corpus_status` (doesn't exist) and missed both `lookup_variety` and `search_trials` - Build-phase table showed everything as "pending" / "next" but Phases 1-8 + 11 all shipped Rewrite to reflect the deployed state: - Corpus inventory: 760 variety records + 4,313 trial documents = 5,073 chunks across 6 sources - All 6 MCP tools documented with their purpose - Eval baseline table (hybrid+rerank wins 100%, P@1 90%, MRR 0.905) with the surprising findings (dense alone is noise; hybrid w/o rerank is WORSE than BM25 alone) - Deploy mechanics: Watchtower chain, 4-GPU embedder pool, shared llama-rerank sidecar with the network-attach gotcha - Status table: ✅ on the phases that shipped, deferred work list (becks_pfr, 2023 plot backfill, NK trials, Channel Seed brand) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-25 17:50:05 -04:00
justin	ac40e05734	seed-mcp scaffold: clone docs-mcp-template, customize for crop_seed PRODUCT_NAME Image rebuild (skip scrape) / build (push) Failing after 7s Details Sibling project to crop-chem-docs, same MCP-template lineage. Corpus is seed/hybrid varieties across 6 vendors instead of pesticide labels. What's customized vs. the template: - CLAUDE.md: vendor matrix, build priority, Pioneer fallback policy, canonical sidecar schema (per-crop), Golden Harvest disease-scale reversal gotcha, no-IPv6 / HTTPS-clone note - README.md: vendor coverage table, tool list, phase status - Dockerfile: PRODUCT_NAME=crop_seed default, sources.json (not bundles.json), HYBRID_SEARCH=true, OLLAMA_URL + RERANK_URL Docker DNS defaults (same llama-rerank sidecar as crop-chem-docs) - .gitea/workflows/refresh.yml: monthly cron (seed catalogs move slowly), 5 GREEN scraper steps, corpus-YYYY.MM.DD tag for Drawbar pinning, continue-on-error on GC step - .gitea/workflows/image-only.yml: paths filter + cancel-in-progress concurrency group - scripts/registry_gc.py: lifted from crop-chem-docs (correct Gitea packages API URL + UA header to bypass CF block on default Python-urllib UA) - sources.json: catalog of 6 sources + scope_filter + per-source schema notes + Pioneer-exclusion rationale - scrape/runner.py: dispatcher with --all = GREEN-only - scrape/sources/{bayer_seeds,golden_harvest,nk,agripro,becks_pfr, becks_products}.py: stub modules with implementation notes - docs_mcp/server.py: PRODUCT_NAME default → crop_seed, PRODUCT_DOCS_URL → repo URL Pioneer is intentionally NOT a source. ToS bans automation; dealer locator is login-gated. The MCP returns a curated fallback lesson directing the user to pioneer.com. Next phases: - Phase 1: implement bayer_seeds (lift-and-shift from crop-chem-docs Bayer scraper; same __NEXT_DATA__ infra) - Phase 7: curate eval/queries.jsonl - Phase 11: lessons.md with Pioneer fallback + disease-scale notes Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-25 12:28:49 -04:00

6 Commits