Trial-data scrapers: gh_plot_reports + agripro_trials + search_trials tool

This PR introduces TRIAL data — yield-performance results from real field trials — as a SEPARATE data type alongside variety identity. The two are complementary: search_docs → "What's the disease resistance of DKC62-08RIB?" (variety identity — what it IS) search_trials → "Which corn hybrid won the IA 2024 trials?" (performance data — how it PERFORMED) scrape/sources/gh_plot_reports.py — Golden Harvest plot reports - 4,618 expected (2024+2025; 2023 deferred to a backfill pass). - URL: /<crop>/plot-report/<state>/<year>/<plot_id> - Cross-vendor: each plot lists products from multiple brands (NK / DEKALB / Golden Harvest / Enogen / Pioneer / Channel) side by side at one cooperator's field — the kind of independent comparison data Bayer doesn't publish itself. - Generic per-column metrics dict (Yield/MST/Test Weight/$/Ac for corn+soy, Ton/Acre + Milk + Beef columns for silage). - Politeness: 1 req/sec, retries on 429/5xx, no redirect-follow. scrape/sources/agripro_trials.py — AgriPro regional trial PDFs - 14 unique PDFs (38 sitemap links deduped) at /trials-data - pdfplumber text extraction, region/year detection from filename - Verbatim PDF text preserved in chunk body so variety + yield number adjacency drives retrieval (AP Iliad's Aberdeen ID yield matches a query about "AP Iliad Idaho yield") rag/chunk.py — chunks_from_trial() dispatching by source - Plot reports: identity preamble + Top-5 by primary metric + full ranking table. Metric labels chosen from the data (corn/soy use "Yield", silage uses "Ton/Acre"). - AgriPro PDFs: identity preamble + verbatim trial body inline so per-location yields surface for region+variety queries. - Variety chunks get data_type="variety" metadata; trial chunks get data_type="trial". Single Chroma collection; the tool router filters by data_type rather than maintaining two collections. rag/index.py — dispatch by sidecar's data_type field rag/bm25.py — new filter columns (data_type, year, state) docs_mcp/server.py — sixth MCP tool: search_trials(crop?, state?, year?, product?, k=10) - Filters trial chunks via where={"data_type": "trial", ...} - Optional product substring post-filter for "DKC62-08RIB Iowa 2024" style searches - search_docs now defaults to data_type="variety" so trial chunks don't bleed into variety identity queries - Tool docstring routes the agent: "use lookup_variety to verify identity details on any trial winner you surface" NK trial endpoint (/NKSeeds/wsProxy.asmx/GetPlotResult) is documented as deferred — the ASMX-SOAP shape returned empty XML on initial probe. Bayer per-variety yield data is not publicly indexed at all — documented in the trial-scope note (DEKALB/Asgrow trial data flows through Channel reps, not the web). AgRevival research books exist as 10 large annual PDFs but are deferred (low ROI per parse). Initial corpus shipped in this PR: 14 AgriPro trial PDFs. The 4,618 Golden Harvest plot reports are scraping in background and will be added in a follow-up corpus-snapshot PR (~70 min ETA). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:19:03 -04:00
parent 7b3da908e0
commit c737871c4c
35 changed files with 3302 additions and 25 deletions
@@ -42,7 +42,12 @@ DEFAULT_DB_NAME = "crop_seed_docs.db"
 # Columns we expose as filterable metadata. Mirrors what
 # ``docs_mcp.server._build_where`` accepts so the same filter dict
 # works for both Chroma and BM25 without per-retriever translation.
-FILTER_COLUMNS = ("source", "vendor", "brand", "crop", "source_key", "ordinal")
+# data_type / year / state / region are trial-specific facets; variety
+# chunks leave them empty.
+FILTER_COLUMNS = (
+    "source", "vendor", "brand", "crop", "source_key",
+    "data_type", "year", "state", "ordinal",
+)


 # Allowlist tokenizer for free-text queries. FTS5's parser chokes on
@@ -131,8 +136,9 @@ class BM25Index:
            con.executescript(self._schema_sql())
            con.executemany(
                "INSERT INTO chunks_meta "
-                "(id, source, vendor, brand, crop, source_key, ordinal) "
-                "VALUES (?, ?, ?, ?, ?, ?, ?)",
+                "(id, source, vendor, brand, crop, source_key, "
+                " data_type, year, state, ordinal) "
+                "VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
                [
                    (
                        r["id"],
@@ -141,6 +147,9 @@ class BM25Index:
                        r["metadata"].get("brand") or "",
                        r["metadata"].get("crop") or "",
                        r["metadata"].get("source_key") or "",
+                        r["metadata"].get("data_type") or "variety",
+                        int(r["metadata"]["year"]) if isinstance(r["metadata"].get("year"), int) else None,
+                        r["metadata"].get("state") or "",
                        int(r["metadata"].get("ordinal") or 0),
                    )
                    for r in records
@@ -216,12 +225,18 @@ class BM25Index:
            brand      TEXT,
            crop       TEXT,
            source_key TEXT,
+            data_type  TEXT,
+            year       INTEGER,
+            state      TEXT,
            ordinal    INTEGER
        );
        CREATE INDEX idx_meta_source     ON chunks_meta(source);
        CREATE INDEX idx_meta_crop       ON chunks_meta(crop);
        CREATE INDEX idx_meta_brand      ON chunks_meta(brand);
        CREATE INDEX idx_meta_source_key ON chunks_meta(source_key);
+        CREATE INDEX idx_meta_data_type  ON chunks_meta(data_type);
+        CREATE INDEX idx_meta_year       ON chunks_meta(year);
+        CREATE INDEX idx_meta_state      ON chunks_meta(state);

        CREATE VIRTUAL TABLE chunks_fts USING fts5(
            text,