# Troubleshooting ## Goal Turn any red signal from the deployment — a status code, a journal gap, a stuck health check — into a cause, a fix, and a re-check, without guessing: every diagnostic below names the real command, field, or file that decides. ## Scope Component `law-mcp-server` over HTTP (crate `law-mcp`), plus its image build, proxy edge, and answers service (`LAW_ANSWERS_URL`). All statuses, codes, and commands are real; all tokens, hostnames, and payloads are synthetic created-examples. For sizing causes, see [Capacity and scaling](/operate/capacity-scaling/); for what "checked" means, see [Verification and support matrix](/operate/verification-support-matrix/). ## Applies to | Branch | Coverage in this article | |---|---| | MCP (`law-mcp-server` over HTTP) | full symptom → fix tables | | `law serve` | five serve failure scenarios below (journal write, lost response, crash tail, cold workers, full disk), plus the per-procedure tables in [Deploy a private HTTP service](/operate/private-http-service/), [Upgrades and compatibility](/operate/upgrades-compatibility/), [Rollback](/operate/rollback/) | ## Prerequisites - Shell access to the host or the container runtime; `curl` and `jq` (operator tools, not shipped in the `scratch` image). - The instance's `X-Law-Call-Id` for the failing call (every response carries one — real, `src/http.rs`) and its journal stream. - The image digest and profile the instance should run (see [Upgrades and compatibility](/operate/upgrades-compatibility/) for pinning). ## Steps Triage in this order; each area below follows symptom → diagnostics → cause → fix → check. ### Auth (401, startup bind refusal) - Symptom: `401` with `WWW-Authenticate: Bearer resource="law-mcp"`, or the process exits `2` at startup refusing a non-loopback bind. - Diagnostics: repeat the call with `-v` and confirm the `Authorization: Bearer ` header is byte-exact; read stderr for the bind-refusal line naming the host. `-v` prints the live secret: never paste that output into a ticket, chat, or CI log — compare locally, then discard the terminal scrollback; for a shareable trace, re-run with a revoked synthetic token instead. - Cause: wrong or missing token (the server compares the full `Bearer ` string — real, `src/http.rs` `guard()`), or a public bind with neither `LAW_MCP_TOKEN` nor `LAW_MCP_PUBLIC=1` (refused by config validation). - Fix: set the token via `--token` / `LAW_MCP_TOKEN` (real), or set `LAW_MCP_PUBLIC=1` only as a deliberate, recorded decision. - Check: the same call returns `200`; startup prints the `law-mcp-http: listening …` line to stderr. ### Proxy and edge (502 / 504 / wrong client address) - Symptom: edge `502`/`504`, or journal `client` fields that all read as one address. - Diagnostics: `curl` the backend `GET /healthz` directly (bypassing the edge); compare with the edge response; check edge `proxy_read_timeout` against `LAW_MCP_CALL_TIMEOUT`. - Cause: backend down (see Memory, Filesystem); timeout inversion — the edge gave up first (the reference-setting is call limit 110 s below proxy 120 s, `src/main.rs`); or the edge does not forward the client address (the server takes the first `X-Forwarded-For` entry — real, `src/http.rs` `context()`). - Fix: revive the backend; restore the timeout ordering; forward `X-Forwarded-For` and `X-Request-Id` at the edge (operator policy; header trust holds by layout — one public path through the edge). - Check: direct and edge `/healthz` both `200` with `"status":"ok"`; journal `client` varies per caller. ### Slice pins (`--check-pin` exit 1, image refuses profile) - Symptom: `image_tree.py --check-pin ` exits `1`, or the build step fails before `docker build`. - Diagnostics: run the check; it prints the path-by-path diff against the slice pin file for the profile (real, schema `arxo.mcp-slice-pin/1`). - Cause: tree contents drifted from the pin; profile `all` requested (refused: images are built for a slice — real, `image_tree.py`); missing CLIR for a prefix; pin schema mismatch. - Fix: regenerate the tree from the same commit and re-pin only after reviewing the diff file; never hand-edit the tree into silence. Never build an image for `all`. - Check: `--check-pin` exits `0`; the image build repeats the same two commands plus `--manifest`. ### Missing resources (startup exit 2, empty answers) - Symptom: process exits `2` with `law-mcp: …` on stderr, or tools answer empty on a fresh slice. - Diagnostics: read the full stderr line (all startup failures exit `2` — real); list the profile files (`*.json`) for the requested name (real examples: `kz`, `phys`, `chem`, `code-civil`, `de-bgb`); verify the slice carries its canon and `calendarPackage` per the profile JSON. - Cause: unknown profile name, `--root` pointing away from the tree (`find_root` needs the checkout layout), or a slice whose canon the question does not touch. - Fix: correct `--root`/`--profile`; rebuild the slice for the needed profile; re-ask against the profile whose canon covers the question. - Check: startup `listening` line names the intended profile; `/healthz` `"profile"` matches it; the CI-style `law_search` smoke call returns non-error text. ### Incompatible versions (400 protocol version, stale client) - Symptom: `400` naming the protocol version and the known list, or `initialize` negotiating down unexpectedly. - Diagnostics: read the `known =` list in the `400` body (real, `src/http.rs` `guard()`); compare with the client's `mcp-protocol-version` header. - Cause: client sends a revision outside `2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` (real, `SUPPORTED_PROTOCOL_VERSIONS`) on a non-`initialize` call. - Fix: upgrade or pin the client to a listed revision; `initialize` alone is exempt by design. - Check: the same call returns `200`; `/healthz` is unaffected (it never negotiates versions). ### Filesystem and permissions (artifact failures, read-only tree) - Symptom: `explain` of a saved answer fails with `ANSWER_RESOURCE_UNAVAILABLE`, or the container cannot start. - Diagnostics: confirm `LAW_MCP_ARTIFACT_DIR`, else `$HOME/.cache/law-dsl/mcp-answers/` (real, `src/execution.rs`), exists and is readable by `10001:0` (real image user, `Dockerfile`); check `HOME` is set when the variable is not (an unset `HOME` is exactly the coded failure above). The server only *reads* this directory — nothing in it writes there (see the [effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix)). - Cause: missing artifact file, wrong ownership, missing `HOME` in a minimal runtime, or nobody staged the artifact (no save step ran out-of-band). - Fix: stage the artifact file where the reading instance sees it (shared mount or copy step) and export the dir explicitly; `chown` readable to the image UID (operator policy); do not run the image as root to dodge ownership (the `USER` line is the isolation boundary — see [Runtime and filesystem isolation](/operate/runtime-filesystem-isolation/)). - Check: save-then-explain round-trips on the same instance. ### Timeouts and overruns (-32001, rising `overrunCalls`) - Symptom: JSON-RPC error `-32001` with `data.code` `CALL_TIMEOUT` and `limitSeconds`, and/or `/healthz` `overrunCalls` above `0`. - Diagnostics: read `overrunCalls` before, during, and after load; find the matching journal record (the abandoned call is journaled; its late completion arrives as a separate record — real, `src/http.rs` `dispatch()`). - Cause: the call exceeded `LAW_MCP_CALL_TIMEOUT` (default: no limit — real); workers are never cancelled (implementation-limit, no cancellation points in `law-eval`). - Fix: set or raise the limit only for question classes that need it; shed load at the edge; warm lazy indexes (see [Capacity and scaling](/operate/capacity-scaling/) step 2). - Check: `overrunCalls` returns to `0` after the drain; no new `-32001` at the planned concurrency. ### Memory and full disk (OOM kills, stalled writes) - Symptom: container restarts with OOM status, or the instance stops answering while the process lives. - Diagnostics: runtime OOM counters and RSS history; disk usage of the artifact directory, container log store (the journal fallback emits each call's JSON record to stderr — real, `src/journal.rs`; one primary line per served call, plus a possible late `OVERRUN_FINISHED` for the same call id), and journald quota on systemd hosts. - Cause: no in-code memory cap (implementation-limit) met more concurrent or overrun calls than the box holds; or a full disk under artifacts/logs. - Fix: lower edge concurrency, add identical instances per [Capacity and scaling](/operate/capacity-scaling/), rotate/prune logs and artifacts (operator policy); raise the container limit only from measured RSS, never blindly. - Check: `/healthz` stays `200` through a repeated load run; RSS flat across runs; artifact round-trip works. ### Journal gaps (missing records, `MCP_TRUNCATED`) - Symptom: calls with no journal record, or records with `MCP_TRUNCATED=oversize`. - Diagnostics: check for `/run/systemd/journal/socket` on the host; read container stderr; compare record count against served calls minus `/healthz` hits (health checks are never journaled — real). - Cause: no socket where the instance runs (macOS, tests, containers) — stderr JSON lines are the designed fallback, not a failure (real, `src/journal.rs`); oversize datagrams are resent without bodies with `MCP_TRUNCATED` (real); or recording was disabled with `LAW_MCP_LOG=0` (real default: on). - Fix: collect stderr as the journal where there is no socket; raise `LAW_MCP_LOG_MAX_BODY` from the 64 KiB-per-side default (real; `0` removes the cap) only within the ~212 KiB Linux datagram budget (real, `src/journal.rs`); re-enable `LAW_MCP_LOG` unless silence is deliberate and recorded. - Check: every served non-health call id has its primary record (match `OVERRUN_FINISHED` extras by call id, don't just count lines); missing ids are investigated as record loss, not assumed absent calls; no new `MCP_TRUNCATED` at production body sizes. ### Replay and publication (answers service, prepare refusals) - Symptom: `law_prepare_answer` / `law_publish_answer` report unavailable or refuse; publication calls hang ~30 s. - Diagnostics: confirm `LAW_ANSWERS_URL` and `LAW_ANSWERS_TOKEN` are both set and non-empty (either missing yields "not configured" — real, `src/publication.rs`); time the call (service client timeout is 30 s — implementation-limit, `src/service_client.rs`; responses cap at 32 MiB — implementation-limit, `src/publication.rs`). - Cause: unwired or unreachable answers service; or a replay mismatch — any divergence of inputs, rights, resources, code, or result rejects the whole document (interface contract of the prepare call). - Fix: wire and reach the service; for replay refusals, re-run with byte-identical inputs against the identical slice and compare field by field — do not edit the frozen inputs into agreement. - Check: prepare succeeds on the unchanged inputs; publish returns its URL; revoke closes it. ### Serve: journal write or sync failure (`SERVE_JOURNAL_FAILED`, then `failed`) - Symptom: one `POST /v1/ask` answers 500 `SERVE_JOURNAL_FAILED` (`segment N: record/fsync: `); `/readyz` flips to 503 `{"status":"failed", "reason":"decision journal broken: …"}`; every later ask answers 503 `SERVE_NOT_READY`. `/healthz` stays 200 — the process lives, it just does not serve (real, `service.rs` `failure()` + `readiness()`). - Diagnostics: read the `readyz` reason (it names the segment and the OS error); check server stderr, `df`/`dmesg`, mount flags (read-only?), and directory ownership on the journal dir. - Cause: the record write or `sync_data` failed and the journal is marked broken for the rest of the process lifetime (real, `journal.rs` `append`): a decision is never served without its fsync'd record (C2). Retrying the call cannot succeed — the same append path fails again. - Fix: fence ingress, fix the underlying cause (space, mount, permissions), then restart the process (only a restart clears the broken flag — a re-open rebuilds the index and cuts any torn tail with a log line). Verify `/readyz` 200 with the expected pin, spot-replay one pre-failure decision id, then reopen ingress. - Check: `/readyz` 200; a control ask answers 200 with a new fsync'd record; no new `SERVE_JOURNAL_FAILED`. ### Serve: HTTP response lost after the decision was saved (idempotency-key rule) - Symptom: client timed out or dropped the connection; the decision may already be recorded — re-asking naively would mint a second decision for the same case. - Diagnostics: if the first send carried an `Idempotency-Key`, nothing is lost: the key claims the recorded bytes. - Cause: transport failure after commit, not a server error. - Fix: resend the byte-identical body with the SAME `Idempotency-Key` (1–255 printable ASCII, no spaces — real, `service.rs`). The server replays the recorded bytes with the same `decisionId` (real, `Claim::Replay`). Never retry a lost response with a fresh key (second decision) or a changed body under the old key (422 `SERVE_IDEMPOTENCY_CONFLICT`). A resend while the first call still runs answers 409 `SERVE_IDEMPOTENCY_IN_PROGRESS` — wait, then resend. Rule of the road: generate the key client-side BEFORE the first send; keys need a journal (a key at `--no-journal` is 400 `SERVE_IDEMPOTENCY_KEY_INVALID`). - Check: the resend returns the same `decisionId` and `resultHash` as the journaled record; exactly one record carries that key. ### Serve: crash stop during a write (torn tail check and recovery) - Symptom: after a SIGKILL/crash/power cut, the last segment may end mid-line; restart behavior decides whether anything was lost. - Diagnostics: BEFORE restarting, inventory the copy with the read-only kit helper (see [Backup and restore](/operate/backup-restore/)): `kit/bin/arxo-list-decisions `. `TORN_TAIL` on the last segment is the expected crash shape (unconfirmed bytes, never a decision); any `REFUSAL` names file and offset and fails the drill. - Cause: the killed call left either a complete fsync'd record or none (real) — only the unconfirmed tail is ever partial. - Fix: snapshot the segments if the tail bytes are evidence (restart truncates them and logs `cut torn tail of N bytes from segment M`), then restart normally, confirm `/readyz` 200 with the same pin and `programHash`, and re-ask the interrupted case explicitly (new decision, new record). - Check: `/readyz` 200; `arxo-list-decisions` on the live dir lists every confirmed record with no `TORN_TAIL` and no `REFUSAL`. ### Serve: workers unavailable or world not warm (`warming`, `NOT_READY`) - Symptom: `/healthz` 200 but `/v1/ask` answers 503 `SERVE_NOT_READY`, and `/readyz` answers 503 `{"status":"warming", warmed, workers}`. - Diagnostics: compare `warmed` vs `workers` in both endpoints (real fields); read server stderr for worker launch errors. - Cause: workers still warming (normal after start or a `--watch-world` flip), or a worker that never warms. This is NOT a token, pin, or journal problem — those have their own codes. - Fix: wait for `/readyz` 200 before opening ingress (the only correct ingress condition — never `/healthz`); if warming never completes, inspect the worker launch path and the world bytes, then restart. A `failed` status instead of `warming` means warmup itself errored — the reason names it. - Check: `/readyz` 200 with the expected pin; a control ask answers 200. ### Serve: disk full (what blocks, what is already safe, recovery) - Symptom: asks start failing with 500 `SERVE_JOURNAL_FAILED` naming an `ENOSPC` write/sync error; `/readyz` flips to `failed` as in the journal-failure scenario above. - Diagnostics: `df` on the journal filesystem; segment sizes; stderr for the first failing segment and offset. - Cause: no space (or quota/inodes) for the next record or the next segment file. Already safe: every record fsync'd before the failure is a complete, replayable decision — the failing call has NO record and its decision was never served (C2). Blocked: all further asks (503 `SERVE_NOT_READY`), segment rotation, and any restart-time manifest write until space returns. - Fix: fence ingress; free space or grow the volume (do NOT delete or hand-edit journal segments — evidence); restart; verify `/readyz` 200 + pin; spot-replay; reopen ingress. Set a disk alert below the failure point (operator policy). - Check: same as the journal-failure check, plus free-space headroom confirmed on the journal filesystem. ## Expected result Every red signal maps to exactly one area above, the fix lands, and the area's check is green before traffic resumes. ## Result check - `GET /healthz`: `200`, `"status":"ok"`, `"oracleGuard":"absent"`, `overrunCalls` `0` after drain. - One CI-style smoke `law_search` through the edge returns non-error text. - The incident's `X-Law-Call-Id` resolves to a complete journal record. ## Failures and diagnostics - Two areas match at once (e.g. `504` + rising `overrunCalls`): fix the lower layer first (backend capacity), then re-test the upper (edge timeouts) — never both at once. - A fix that cannot produce its check is not a fix: roll it back and re-triage. - Anything outside these areas (kernel, storage array, cloud fabric) is host territory: record the boundary and hand over with the call IDs and digests attached. ## Support boundaries - Supported: the diagnostics and checks above against the matrix in [Verification and support matrix](/operate/verification-support-matrix/). "Fixed" here means only "the area's check is green on the pinned slice". - Implementation limits surfaced here, not worked around: 1 MiB request cap, no worker cancellation, no in-code memory cap, 30 s / 32 MiB answers-service caps. - Out of scope: debugging client applications beyond their request bytes, and any host or network fabric below the container. ## Next step Append the incident (signal, cause, fix, digest, call IDs) to the decision journal per [Logs, audit, decision journals](/operate/logs-audit-decision-journals/); if the cause was capacity, re-run [Capacity planning and scaling](/operate/capacity-scaling/).