Commit Graph
2 Commits
Author SHA1 Message Date
claudeandClaude Opus 4.8 1d3e52a5ed feat(scrape): vision OCR of image-only doc pages (support matrices)
Some HPE pages carry information only inside an image. The Morpheus
release schedule / support-lifecycle matrix (sf000111242en_us) is the
case that surfaced this: the entire "which version ships when, when a
stream hits Maintenance/EOL" table is a JPEG, so html_to_md dropped
every value and retrieval could never answer "when does 8.x.x reach
end of life".

New scrape/vision.py (opt-in via VISION_OCR) transcribes qualifying
images with a local Ollama vision model and appends a labeled markdown
block to the page so the text is chunked, embedded and retrieved.

Model/prompt chosen against ground truth on the matrix (14x4 cells):
- qwen2.5vl:7b + a bare "transcribe the table" prompt: 55-56/56, ~18s,
  GPU-resident, and prompt-ROBUST (accurate without hand-tuning).
- gemma3:12b needs an exact "read row by row, keep cells aligned"
  prompt or it shifts a column by one row; qwen2.5vl:32b is no more
  accurate and 5x slower (CPU spill). Both documented in the code.
- Note: self-consistency (2 samples must agree) only catches RANDOM
  flakiness — a wrong read is stable across seeds — so every block also
  ships a "verify against the source image" caveat.

Reliability/cost:
- content-hash cache in corpus/.vision-cache/ (committed) — each unique
  image OCR'd once ever, shared across version bundles.
- VISION_MAX_NEW bounds NEW OCRs per run so the first pass can't balloon
  into hours; the cache fills incrementally. Deferred count is logged.
- every failure path degrades to the pre-vision behavior; never blocks
  the scrape.

Also:
- add the morpheus_release_schedule bundle (sf single-doc) + 2 eval
  golden queries for it.
- fetch_single_doc: title falls back to the bundle title (not docId)
  when a page has no <h1> (sf solution articles).
- Pillow in requirements-vision.txt (scrape-only; kept out of the
  server image), installed + VISION_* wired into refresh.yml.

Verified locally: matrix transcribes to the exact 14-row table; second
run is a cache hit (ocr=0).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01LFowQzJu7k97QLCRDSAeh1
2026-07-23 09:35:00 -04:00
justinandClaude Opus 4.7 fa448f94e1 build out morpheus-docs MCP stack, mirroring hvm-docs through Phases 1-13
Initial scaffold: the docs-mcp-template clone with all the
HVM-validated stack ported across, customized for Morpheus
Enterprise (PRODUCT_NAME=morpheus, server name morpheus-docs).

Bundles (live-discovered 2026-05-22; 1710 cataloged pages total):
* morpheus_user_manual_8_1_0  sd00007510en_us  568 pages (Feb 2026)
* morpheus_user_manual_8_1_1  sd00007621en_us  569 pages (Mar 2026)
* morpheus_user_manual_8_1_2  sd00007732en_us  569 pages (Apr 2026)
* morpheus_release_notes_8_1_0  sd00007496en_us  single-doc
* morpheus_release_notes_8_1_1  sd00007610en_us  single-doc
* morpheus_release_notes_8_1_2  sd00007733en_us  single-doc
* morpheus_quickspecs            a50009231enw     html-file (live
  curl_cffi against www.hpe.com; all 12+ Enterprise SKUs captured —
  S6E64..S6E73AAE for new/renewal/upgrade × 1/3/5-yr terms, plus
  services SKUs HA124A1#V38/V39 and H46SBA1).

No Deployment Guide or Qualification Matrix on HPE Support for
Morpheus Enterprise specifically — the only QM (sd00006551en_us)
covers HVM clusters managed by Morpheus and lives in hvm-docs.

Stack carried forward from hvm-docs:
* rag/{index,chunk,embeddings,bm25}.py — including the
  MAX_CHARS=4000 chunk-cap fix for table-dense content
* docs_mcp/{server,usage}.py — 11 MCP tools, BM25-default search,
  cross-encoder rerank, hybrid behind HYBRID_SEARCH=true,
  morpheus_api_lessons (renamed from hvm_api_lessons), env-gated
  submit_doc_bug
* docs_mcp/api_lessons.md — Morpheus-specific scaffold covering
  licensing model, HVM elevation path, REST vs Plugin API, with
  TODO markers for sections to flesh out from real ops experience
* scrape/{runner,quickspecs,changelog,bundles}.py — TOC + single-doc
  + html-file modes, curl_cffi Chrome120 for www.hpe.com edge bypass
* eval/{retrievers,run_eval}.py + queries.jsonl scaffold (4 placeholder
  queries; populate after first scrape)
* scripts/{rerank_server,usage_report,registry_gc}.py
* .gitea/workflows/{refresh,image-only}.yml — same Gitea Actions
  setup zerto-docs uses (push LAN, pull public-URL, GPU Ollama pool)
* deploy/docker-compose.yml — morpheus-docs-mcp service definition,
  shared jina-rerank sidecar, Watchtower-labeled
* Dockerfile, requirements.txt, requirements-rerank.txt

Verified locally: scrape produced 1599 .md pages (some TOC entries
are parent-only and yield no body), 6353 chunks all under the 4 KB
cap, MCP server boots and lists 11 tools cleanly.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-22 15:26:24 -04:00