drawbar-backend-seed-mcp-1 on trashpanda is in a crash loop — RestartCount=11685 as of 2026-09-09 20:00Z, exit 1 every ~10 s, Restarting (1) in docker ps:
File "/app/docs_mcp/server.py", line 28, in <module>
from mcp.server.fastmcp import FastMCP
ModuleNotFoundError: No module named 'mcp.server.fastmcp'. This is mcp 2.x,
where FastMCP was renamed to MCPServer (from mcp.server.mcpserver import MCPServer) ...
Root cause
requirements.txt line 2: mcp[fastmcp]>=1.0.0 — no upper bound. The refresh.yml corpus-refresh workflow rebuilt the image on 2026-09-01 11:25 EDT, pip resolved mcp 2.1.1, and 2.x removed mcp.server.fastmcp. The workflow pushes :latest, which is the Watchtower target; Watchtower pulled it at 16:21Z the same day and the service has been down since — 8 days unnoticed, because nothing alerts on a restarting container and the advisor just loses one tool source.
The image is git.jpaul.io/justin/seed-mcp:latest; chem-mcp is safe only because it's pinned to corpus-2026.05.24 — the next corpus refresh of any docs-MCP built from the same template hits this.
Fix (scoped — ai-ready)
Minimum, ship today:
requirements.txt: mcp[fastmcp]>=1.0.0,<2 (and check hvm-docs / morpheus-docs / opsramp / zerto-docs / crop-chem-docs + the docs-mcp-template — same line, same trap).
Rebuild + push; Watchtower picks it up. Confirm docker ps shows Up … (healthy) and the advisor's seed tools register.
Proper follow-up (separate PR): migrate to mcp.server.mcpserver.MCPServer per the 2.x rename and drop the pin.
Structural follow-up worth a line in the template's CLAUDE.md: a corpus refresh should not be able to change the runtime dependency set. Either pin the whole requirements file (pip-compile) or build the runtime layer from a pinned base and only swap the corpus.
## What's happening
`drawbar-backend-seed-mcp-1` on trashpanda is in a crash loop — **`RestartCount=11685`** as of 2026-09-09 20:00Z, exit 1 every ~10 s, `Restarting (1)` in `docker ps`:
```
File "/app/docs_mcp/server.py", line 28, in <module>
from mcp.server.fastmcp import FastMCP
ModuleNotFoundError: No module named 'mcp.server.fastmcp'. This is mcp 2.x,
where FastMCP was renamed to MCPServer (from mcp.server.mcpserver import MCPServer) ...
```
## Root cause
`requirements.txt` line 2: `mcp[fastmcp]>=1.0.0` — **no upper bound**. The `refresh.yml` corpus-refresh workflow rebuilt the image on **2026-09-01 11:25 EDT**, pip resolved **`mcp 2.1.1`**, and 2.x removed `mcp.server.fastmcp`. The workflow pushes `:latest`, which is the Watchtower target; Watchtower pulled it at 16:21Z the same day and the service has been down since — **8 days** unnoticed, because nothing alerts on a restarting container and the advisor just loses one tool source.
The image is `git.jpaul.io/justin/seed-mcp:latest`; `chem-mcp` is safe only because it's pinned to `corpus-2026.05.24` — the next corpus refresh of *any* docs-MCP built from the same template hits this.
## Fix (scoped — `ai-ready`)
Minimum, ship today:
1. `requirements.txt`: `mcp[fastmcp]>=1.0.0,<2` (and check `hvm-docs` / `morpheus-docs` / `opsramp` / `zerto-docs` / `crop-chem-docs` + the **docs-mcp-template** — same line, same trap).
2. Rebuild + push; Watchtower picks it up. Confirm `docker ps` shows `Up … (healthy)` and the advisor's seed tools register.
Proper follow-up (separate PR): migrate to `mcp.server.mcpserver.MCPServer` per the 2.x rename and drop the pin.
Structural follow-up worth a line in the template's CLAUDE.md: a corpus refresh should not be able to change the *runtime* dependency set. Either pin the whole requirements file (`pip-compile`) or build the runtime layer from a pinned base and only swap the corpus.
## Verify
```
docker inspect --format '{{.RestartCount}} {{.State.Status}}' drawbar-backend-seed-mcp-1 # stops climbing; running
docker logs --tail 5 drawbar-backend-seed-mcp-1 # no ModuleNotFoundError
docker run --rm --entrypoint python git.jpaul.io/justin/seed-mcp:latest -c "import importlib.metadata as m;print(m.version('mcp'))" # 1.x
```
Found during a `docker ps` on trashpanda while shipping Drawbar/planning#113.
Now visible without anyone looking: the new lab-observability hub on trashpanda (Loki + Alloy, deployed 2026-09-10 00:23Z) counted 10 ModuleNotFoundError tracebacks from drawbar-backend-seed-mcp-1 in its first 5 minutes, and the Grafana rule Repeated Python traceback in a container went to Alerting for this container within 3 minutes of the stack coming up. Query for the record: sum(count_over_time({container="drawbar-backend-seed-mcp-1"} |~ ModuleNotFoundError [5m])). The fix in this issue is unchanged; this is the detection side that was missing for 8 days.
Now visible without anyone looking: the new lab-observability hub on trashpanda (Loki + Alloy, deployed 2026-09-10 00:23Z) counted **10 `ModuleNotFoundError` tracebacks from `drawbar-backend-seed-mcp-1` in its first 5 minutes**, and the Grafana rule *Repeated Python traceback in a container* went to **Alerting** for this container within 3 minutes of the stack coming up. Query for the record: `sum(count_over_time({container="drawbar-backend-seed-mcp-1"} |~ `ModuleNotFoundError` [5m]))`. The fix in this issue is unchanged; this is the detection side that was missing for 8 days.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What's happening
drawbar-backend-seed-mcp-1on trashpanda is in a crash loop —RestartCount=11685as of 2026-09-09 20:00Z, exit 1 every ~10 s,Restarting (1)indocker ps:Root cause
requirements.txtline 2:mcp[fastmcp]>=1.0.0— no upper bound. Therefresh.ymlcorpus-refresh workflow rebuilt the image on 2026-09-01 11:25 EDT, pip resolvedmcp 2.1.1, and 2.x removedmcp.server.fastmcp. The workflow pushes:latest, which is the Watchtower target; Watchtower pulled it at 16:21Z the same day and the service has been down since — 8 days unnoticed, because nothing alerts on a restarting container and the advisor just loses one tool source.The image is
git.jpaul.io/justin/seed-mcp:latest;chem-mcpis safe only because it's pinned tocorpus-2026.05.24— the next corpus refresh of any docs-MCP built from the same template hits this.Fix (scoped —
ai-ready)Minimum, ship today:
requirements.txt:mcp[fastmcp]>=1.0.0,<2(and checkhvm-docs/morpheus-docs/opsramp/zerto-docs/crop-chem-docs+ the docs-mcp-template — same line, same trap).docker psshowsUp … (healthy)and the advisor's seed tools register.Proper follow-up (separate PR): migrate to
mcp.server.mcpserver.MCPServerper the 2.x rename and drop the pin.Structural follow-up worth a line in the template's CLAUDE.md: a corpus refresh should not be able to change the runtime dependency set. Either pin the whole requirements file (
pip-compile) or build the runtime layer from a pinned base and only swap the corpus.Verify
Found during a
docker pson trashpanda while shipping Drawbar/planning#113.Now visible without anyone looking: the new lab-observability hub on trashpanda (Loki + Alloy, deployed 2026-09-10 00:23Z) counted 10
ModuleNotFoundErrortracebacks fromdrawbar-backend-seed-mcp-1in its first 5 minutes, and the Grafana rule Repeated Python traceback in a container went to Alerting for this container within 3 minutes of the stack coming up. Query for the record:sum(count_over_time({container="drawbar-backend-seed-mcp-1"} |~ModuleNotFoundError[5m])). The fix in this issue is unchanged; this is the detection side that was missing for 8 days.