Port docs-mcp-template upgrades (eval, citations, chunking, nomic prefixes) #25

Closed
opened 2026-09-29 22:22:01 -04:00 by claude · 0 comments
Contributor

This clone (seed-mcp)

  • Already on mcp 2.x (MCPServer, mcp>=2,<3). Skip step 1.
  • Eval already works (eval/queries.jsonl ~21 queries; P@1 / MRR / Passed in eval/run_eval.py). Do not replace that query schema. Add k-curve / JSONL sidecar / eval.pvalue / eval.trace beside it.
  • Dense path still uses query_texts= (docs_mcp/server.py and eval/retrievers.py) — that's the prefix change.
  • Separate outage: #21. Do not mix this PR with that fix.

Port the docs-mcp-template upgrades that landed 2026-09-29 onto this clone, then eval-gate the retrieval ones. Do not treat a green import as “search got better.”

Source of truth

All of this is already on git.jpaul.io, repo justin/docs-mcp-template, branch main (HEAD e6630a1).

Template PR What
#11 mcp 2.x: FastMCP → MCPServer, stateless_http=True on run(), mcp>=2,<3, CI import docs_mcp.server smoke
#12 Phase 7 eval: P@1 + MRR + k-curve @1/5/10/20, JSONL sidecar, eval.pvalue permutation test
#13 search_docs citations: [1] **title** — url (docs_mcp/format.py)
#14 Heading-recursive chunker, keep synthetic chunk-0 (rag/chunk.py)
#15 Eval miss dump: eval/trace.py + run_eval --trace → trace.jsonl / misses.md
#16 Nomic prefixes at embed time only (search_query: / search_document:). Stored Chroma/BM25/rerank text stays unprefixed

Read those files on main; do not copy the template’s search_docs body wholesale — this clone has product-specific filters/tools.

Code to look at (template)

  • docs_mcp/server.py — MCPServer + run() kwargs; _dense_ids uses query_embeddings= + embed_texts(..., prefix=EMBED_QUERY_PREFIX); search_docs renders via format_search_hits
  • docs_mcp/format.py + docs_mcp/test_format.py — citation helper (no mcp/chroma deps)
  • rag/chunk.py + rag/test_chunk.py — heading-recursive pack
  • rag/embeddings.py — apply_prefix / embed_texts / EMBED_QUERY_PREFIX / EMBED_DOC_PREFIX
  • rag/index.py — col.upsert(..., documents=raw_texts, embeddings=prefixed_vectors)
  • rag/test_embeddings.py — prefix unit tests
  • eval/run_eval.py, eval/pvalue.py, eval/trace.py, eval/retrievers.py — harness; retrievers reuse docs_mcp.server helpers (same rule as dashboard/CONTRACT.md)
  • eval/test_metrics.py, eval/test_trace.py
  • .gitea/workflows/image-only.yml + refresh.yml — docker run ... -c "import docs_mcp.server" before push
  • PLAN.md Phase 2 / 3 / 7 notes

Worked recipe for 2.x (if this clone is still FastMCP): justin/seed-mcp #23 and template #11. tools/list dump before/after must stay equivalent. @mcp.tool() unchanged.

What “good” looks like (do in this order)

0. Baseline. If eval/queries.jsonl exists, run the current harness and save eval/results/pre-template.md (+ sidecar if you have one). If eval is still a stub or there is no queries.jsonl, populate/fix eval before changing chunking or prefixes.

1. mcp 2.x — only if this clone still imports mcp.server.fastmcp. stateless_http=True stays required, but it is a run() kwarg now. Add the CI import smoke. Pin mcp>=2,<3. Skip this step if already on MCPServer.

2. Citations — port docs_mcp/format.py and wire search_docs to it. Safe, no reindex. Unit test from the template.

3. Eval tooling — add k-curve + JSONL sidecar + eval.pvalue + eval.trace. If this clone already has a working eval (seed-mcp does), do not replace the query schema; add pvalue/trace/k-curve beside it. Retrievers must keep using this clone’s docs_mcp.server helpers, not a copy-paste of template retrievers.py that misses product filters.

4. Chunking — port heading-recursive rag/chunk.py, keep chunk-0. --rebuild. Re-run eval. If P@1 drops, stop and dump --trace; do not ship on vibes.

5. Nomic prefixes — port embed-time prefixes. Stored documents must stay unprefixed (assert in a test or by grepping a hydrated chunk). --rebuild required. Eval + python -m eval.pvalue --a pre.jsonl --b post.jsonl. Default prefixes on for nomic; EMBED_QUERY_PREFIX= / EMBED_DOC_PREFIX= empty string disables.

Ship 2+3 even if 4/5 are a wash. 4 and 5 are the quality bets; 0–3 are the ones that make the bets reviewable.

Verify

  • python -m unittest for whatever tests you port (docs_mcp.test_format, rag.test_chunk, rag.test_embeddings, eval.test_metrics, eval.test_trace)
  • python -m docs_mcp.server --help boots on mcp 2.x
  • After rebuild: search_docs still returns the right pages on a couple known queries; citations look like [1] **...**; hydrated text has no search_document: prefix
  • Eval numbers in the PR body. Permutation significant or an honest “n too small / no movement”
  • Do not claim a retrieval win without those numbers

Out of scope

  • UltraRAG YAML client, generation server, VisRAG, MinerU, switching jina→bge
  • Replacing this clone’s product tools (diff_versions, interop, etc.)
  • Closing anything on docs-mcp-template

Use gitea-ship: branch → commit as claude → PR with Closes #<this>.

## This clone (`seed-mcp`) - **Already on mcp 2.x** (`MCPServer`, `mcp>=2,<3`). Skip step 1. - Eval **already works** (`eval/queries.jsonl` ~21 queries; P@1 / MRR / Passed in `eval/run_eval.py`). Do **not** replace that query schema. Add k-curve / JSONL sidecar / `eval.pvalue` / `eval.trace` **beside** it. - Dense path still uses `query_texts=` (`docs_mcp/server.py` and `eval/retrievers.py`) — that's the prefix change. - Separate outage: #21. Do not mix this PR with that fix. Port the docs-mcp-template upgrades that landed 2026-09-29 onto this clone, then **eval-gate** the retrieval ones. Do not treat a green import as “search got better.” ## Source of truth All of this is already on **git.jpaul.io**, repo `justin/docs-mcp-template`, branch `main` (HEAD `e6630a1`). | Template PR | What | |---|---| | [#11](https://git.jpaul.io/justin/docs-mcp-template/pulls/11) | mcp 2.x: `FastMCP` → `MCPServer`, `stateless_http=True` on `run()`, `mcp>=2,<3`, CI `import docs_mcp.server` smoke | | [#12](https://git.jpaul.io/justin/docs-mcp-template/pulls/12) | Phase 7 eval: P@1 + MRR + k-curve @1/5/10/20, JSONL sidecar, `eval.pvalue` permutation test | | [#13](https://git.jpaul.io/justin/docs-mcp-template/pulls/13) | `search_docs` citations: `[1] **title** — url` (`docs_mcp/format.py`) | | [#14](https://git.jpaul.io/justin/docs-mcp-template/pulls/14) | Heading-recursive chunker, **keep synthetic chunk-0** (`rag/chunk.py`) | | [#15](https://git.jpaul.io/justin/docs-mcp-template/pulls/15) | Eval miss dump: `eval/trace.py` + `run_eval --trace` → `trace.jsonl` / `misses.md` | | [#16](https://git.jpaul.io/justin/docs-mcp-template/pulls/16) | Nomic prefixes at **embed time only** (`search_query:` / `search_document:`). Stored Chroma/BM25/rerank text stays unprefixed | Read those files on `main`; do **not** copy the template’s `search_docs` body wholesale — this clone has product-specific filters/tools. ### Code to look at (template) - `docs_mcp/server.py` — `MCPServer` + `run()` kwargs; `_dense_ids` uses `query_embeddings=` + `embed_texts(..., prefix=EMBED_QUERY_PREFIX)`; `search_docs` renders via `format_search_hits` - `docs_mcp/format.py` + `docs_mcp/test_format.py` — citation helper (no mcp/chroma deps) - `rag/chunk.py` + `rag/test_chunk.py` — heading-recursive pack - `rag/embeddings.py` — `apply_prefix` / `embed_texts` / `EMBED_QUERY_PREFIX` / `EMBED_DOC_PREFIX` - `rag/index.py` — `col.upsert(..., documents=raw_texts, embeddings=prefixed_vectors)` - `rag/test_embeddings.py` — prefix unit tests - `eval/run_eval.py`, `eval/pvalue.py`, `eval/trace.py`, `eval/retrievers.py` — harness; retrievers **reuse** `docs_mcp.server` helpers (same rule as `dashboard/CONTRACT.md`) - `eval/test_metrics.py`, `eval/test_trace.py` - `.gitea/workflows/image-only.yml` + `refresh.yml` — `docker run ... -c "import docs_mcp.server"` **before** push - `PLAN.md` Phase 2 / 3 / 7 notes Worked recipe for 2.x (if this clone is still FastMCP): `justin/seed-mcp` #23 and template #11. `tools/list` dump before/after must stay equivalent. `@mcp.tool()` unchanged. ## What “good” looks like (do in this order) **0. Baseline.** If `eval/queries.jsonl` exists, run the **current** harness and save `eval/results/pre-template.md` (+ sidecar if you have one). If eval is still a stub or there is no `queries.jsonl`, populate/fix eval **before** changing chunking or prefixes. **1. mcp 2.x** — only if this clone still imports `mcp.server.fastmcp`. `stateless_http=True` stays required, but it is a `run()` kwarg now. Add the CI import smoke. Pin `mcp>=2,<3`. Skip this step if already on MCPServer. **2. Citations** — port `docs_mcp/format.py` and wire `search_docs` to it. Safe, no reindex. Unit test from the template. **3. Eval tooling** — add k-curve + JSONL sidecar + `eval.pvalue` + `eval.trace`. If this clone already has a working eval (seed-mcp does), **do not replace the query schema**; add pvalue/trace/k-curve beside it. Retrievers must keep using this clone’s `docs_mcp.server` helpers, not a copy-paste of template `retrievers.py` that misses product filters. **4. Chunking** — port heading-recursive `rag/chunk.py`, keep chunk-0. `--rebuild`. Re-run eval. If P@1 drops, stop and dump `--trace`; do not ship on vibes. **5. Nomic prefixes** — port embed-time prefixes. Stored documents must stay unprefixed (assert in a test or by grepping a hydrated chunk). `--rebuild` **required**. Eval + `python -m eval.pvalue --a pre.jsonl --b post.jsonl`. Default prefixes on for nomic; `EMBED_QUERY_PREFIX=` / `EMBED_DOC_PREFIX=` empty string disables. Ship 2+3 even if 4/5 are a wash. 4 and 5 are the quality bets; 0–3 are the ones that make the bets reviewable. ## Verify - `python -m unittest` for whatever tests you port (`docs_mcp.test_format`, `rag.test_chunk`, `rag.test_embeddings`, `eval.test_metrics`, `eval.test_trace`) - `python -m docs_mcp.server --help` boots on mcp 2.x - After rebuild: `search_docs` still returns the right pages on a couple known queries; citations look like `[1] **...**`; hydrated text has **no** `search_document:` prefix - Eval numbers in the PR body. Permutation `significant` or an honest “n too small / no movement” - Do not claim a retrieval win without those numbers ## Out of scope - UltraRAG YAML client, generation server, VisRAG, MinerU, switching jina→bge - Replacing this clone’s product tools (`diff_versions`, interop, etc.) - Closing anything on `docs-mcp-template` Use `gitea-ship`: branch → commit as `claude` → PR with `Closes #<this>`.
claude added the ai-readyfeatureP2 labels 2026-09-29 22:22:01 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: justin/seed-mcp#25