feat(scrape): vision OCR of image-only doc pages (support matrices)
Some HPE pages carry information only inside an image. The Morpheus release schedule / support-lifecycle matrix (sf000111242en_us) is the case that surfaced this: the entire "which version ships when, when a stream hits Maintenance/EOL" table is a JPEG, so html_to_md dropped every value and retrieval could never answer "when does 8.x.x reach end of life". New scrape/vision.py (opt-in via VISION_OCR) transcribes qualifying images with a local Ollama vision model and appends a labeled markdown block to the page so the text is chunked, embedded and retrieved. Model/prompt chosen against ground truth on the matrix (14x4 cells): - qwen2.5vl:7b + a bare "transcribe the table" prompt: 55-56/56, ~18s, GPU-resident, and prompt-ROBUST (accurate without hand-tuning). - gemma3:12b needs an exact "read row by row, keep cells aligned" prompt or it shifts a column by one row; qwen2.5vl:32b is no more accurate and 5x slower (CPU spill). Both documented in the code. - Note: self-consistency (2 samples must agree) only catches RANDOM flakiness — a wrong read is stable across seeds — so every block also ships a "verify against the source image" caveat. Reliability/cost: - content-hash cache in corpus/.vision-cache/ (committed) — each unique image OCR'd once ever, shared across version bundles. - VISION_MAX_NEW bounds NEW OCRs per run so the first pass can't balloon into hours; the cache fills incrementally. Deferred count is logged. - every failure path degrades to the pre-vision behavior; never blocks the scrape. Also: - add the morpheus_release_schedule bundle (sf single-doc) + 2 eval golden queries for it. - fetch_single_doc: title falls back to the bundle title (not docId) when a page has no <h1> (sf solution articles). - Pillow in requirements-vision.txt (scrape-only; kept out of the server image), installed + VISION_* wired into refresh.yml. Verified locally: matrix transcribes to the exact 14-row table; second run is a cache hit (ocr=0). Co-Authored-By: Claude Opus 4.8 <[email protected]> Claude-Session: https://claude.ai/code/session_01LFowQzJu7k97QLCRDSAeh1
This commit is contained in:
@@ -37,6 +37,18 @@ env:
|
||||
OLLAMA_URLS: http://192.168.0.2:11435,http://192.168.0.2:11436,http://192.168.0.125:11434,http://192.168.0.126:11434
|
||||
EMBED_MODEL: nomic-embed-text
|
||||
|
||||
# Vision OCR of image-only pages (support/lifecycle matrices etc. — see
|
||||
# scrape/vision.py). Uses the host's primary Ollama on :11434, the only one
|
||||
# with a vision model; the embed pool above is nomic-only. qwen2.5vl:7b
|
||||
# scored best on the release-schedule matrix (55-56/56, ~18s, GPU-resident,
|
||||
# prompt-robust). Content-hash cached in corpus/.vision-cache/; VISION_MAX_NEW
|
||||
# bounds NEW OCRs per run so the cache fills incrementally instead of a
|
||||
# multi-hour first pass. Any failure degrades to the pre-vision behavior.
|
||||
VISION_OCR: "1"
|
||||
VISION_URL: http://192.168.0.2:11434
|
||||
VISION_MODEL: qwen2.5vl:7b
|
||||
VISION_MAX_NEW: "60"
|
||||
|
||||
PRODUCT_NAME: morpheus
|
||||
|
||||
jobs:
|
||||
@@ -64,6 +76,9 @@ jobs:
|
||||
run: |
|
||||
python -m pip install -q --upgrade pip
|
||||
python -m pip install -q -r requirements.txt
|
||||
# Vision OCR deps (Pillow) — only the scrape step needs these; kept
|
||||
# out of requirements.txt so they never bloat the server image.
|
||||
python -m pip install -q -r requirements-vision.txt
|
||||
|
||||
# ---- Phase 1: scrape ---------------------------------------
|
||||
- name: Refresh bundle catalog
|
||||
|
||||
Reference in New Issue
Block a user