docs← Back to article

Markdown for LLMs

Troubleshooting

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Troubleshooting

## Goal

Turn any red signal from the deployment — a status code, a journal
gap, a stuck health check — into a cause, a fix, and a re-check,
without guessing: every diagnostic below names the real command,
field, or file that decides.

## Scope

Component `law-mcp-server` over HTTP (crate `law-mcp`), plus its
image build, proxy edge, and answers service (`LAW_ANSWERS_URL`).
All statuses, codes, and commands are real; all tokens, hostnames,
and payloads are synthetic created-examples. For sizing causes, see
[Capacity and scaling](/operate/capacity-scaling/); for what
"checked" means, see
[Verification and support matrix](/operate/verification-support-matrix/).

## Applies to

| Branch | Coverage in this article |
|---|---|
| MCP (`law-mcp-server` over HTTP) | full symptom → fix tables |
| `law serve` | five serve failure scenarios below (journal write, lost response, crash tail, cold workers, full disk), plus the per-procedure tables in [Deploy a private HTTP service](/operate/private-http-service/), [Upgrades and compatibility](/operate/upgrades-compatibility/), [Rollback](/operate/rollback/) |

## Prerequisites

- Shell access to the host or the container runtime; `curl` and `jq`
  (operator tools, not shipped in the `scratch` image).
- The instance's `X-Law-Call-Id` for the failing call (every response
  carries one — real, `src/http.rs`) and its journal stream.
- The image digest and profile the instance should run
  (see [Upgrades and compatibility](/operate/upgrades-compatibility/) for pinning).

## Steps

Triage in this order; each area below follows
symptom → diagnostics → cause → fix → check.

### Auth (401, startup bind refusal)

- Symptom: `401` with `WWW-Authenticate: Bearer resource="law-mcp"`,
  or the process exits `2` at startup refusing a non-loopback bind.
- Diagnostics: repeat the call with `-v` and confirm the
  `Authorization: Bearer <token>` header is byte-exact; read stderr
  for the bind-refusal line naming the host. `-v` prints the live
  secret: never paste that output into a ticket, chat, or CI log —
  compare locally, then discard the terminal scrollback; for a
  shareable trace, re-run with a revoked synthetic token instead.
- Cause: wrong or missing token (the server compares the full
  `Bearer <token>` string — real, `src/http.rs` `guard()`), or a
  public bind with neither `LAW_MCP_TOKEN` nor `LAW_MCP_PUBLIC=1`
  (refused by config validation).
- Fix: set the token via `--token` / `LAW_MCP_TOKEN` (real), or set
  `LAW_MCP_PUBLIC=1` only as a deliberate, recorded decision.
- Check: the same call returns `200`; startup prints the
  `law-mcp-http: listening …` line to stderr.

### Proxy and edge (502 / 504 / wrong client address)

- Symptom: edge `502`/`504`, or journal `client` fields that all read
  as one address.
- Diagnostics: `curl` the backend `GET /healthz` directly (bypassing
  the edge); compare with the edge response; check edge
  `proxy_read_timeout` against `LAW_MCP_CALL_TIMEOUT`.
- Cause: backend down (see Memory, Filesystem); timeout inversion —
  the edge gave up first (the reference-setting is call limit 110 s
  below proxy 120 s, `src/main.rs`); or the edge does not forward
  the client address (the server takes the first `X-Forwarded-For`
  entry — real, `src/http.rs` `context()`).
- Fix: revive the backend; restore the timeout ordering; forward
  `X-Forwarded-For` and `X-Request-Id` at the edge (operator policy;
  header trust holds by layout — one public path through the edge).
- Check: direct and edge `/healthz` both `200` with `"status":"ok"`;
  journal `client` varies per caller.

### Slice pins (`--check-pin` exit 1, image refuses profile)

- Symptom: `image_tree.py <profile> --check-pin <dir>` exits `1`, or
  the build step fails before `docker build`.
- Diagnostics: run the check; it prints the path-by-path diff
  against the slice pin file for the profile (real, schema
  `arxo.mcp-slice-pin/1`).
- Cause: tree contents drifted from the pin; profile `all` requested
  (refused: images are built for a slice — real, `image_tree.py`);
  missing CLIR for a prefix; pin schema mismatch.
- Fix: regenerate the tree from the same commit and re-pin only
  after reviewing the diff file; never hand-edit the tree into
  silence. Never build an image for `all`.
- Check: `--check-pin` exits `0`; the image build repeats the
  same two commands plus `--manifest`.

### Missing resources (startup exit 2, empty answers)

- Symptom: process exits `2` with `law-mcp: …` on stderr, or tools
  answer empty on a fresh slice.
- Diagnostics: read the full stderr line (all startup failures exit
  `2` — real); list the profile files (`*.json`) for
  the requested name (real examples: `kz`, `phys`, `chem`,
  `code-civil`, `de-bgb`); verify the slice carries its canon and
  `calendarPackage` per the profile JSON.
- Cause: unknown profile name, `--root` pointing away from the tree
  (`find_root` needs the checkout layout), or a slice whose canon
  the question does not touch.
- Fix: correct `--root`/`--profile`; rebuild the slice for the
  needed profile; re-ask against the profile whose canon covers the
  question.
- Check: startup `listening` line names the intended profile;
  `/healthz` `"profile"` matches it; the CI-style `law_search`
  smoke call returns non-error text.

### Incompatible versions (400 protocol version, stale client)

- Symptom: `400` naming the protocol version and the known list, or
  `initialize` negotiating down unexpectedly.
- Diagnostics: read the `known =` list in the `400` body (real,
  `src/http.rs` `guard()`); compare with the client's
  `mcp-protocol-version` header.
- Cause: client sends a revision outside `2025-11-25`,
  `2025-06-18`, `2025-03-26`, `2024-11-05` (real,
  `SUPPORTED_PROTOCOL_VERSIONS`) on a non-`initialize` call.
- Fix: upgrade or pin the client to a listed revision; `initialize`
  alone is exempt by design.
- Check: the same call returns `200`; `/healthz` is unaffected
  (it never negotiates versions).

### Filesystem and permissions (artifact failures, read-only tree)

- Symptom: `explain` of a saved answer fails with
  `ANSWER_RESOURCE_UNAVAILABLE`, or the container cannot start.
- Diagnostics: confirm `LAW_MCP_ARTIFACT_DIR`, else
  `$HOME/.cache/law-dsl/mcp-answers/<profile>` (real,
  `src/execution.rs`), exists and is readable by `10001:0` (real
  image user, `Dockerfile`); check `HOME` is set when the variable
  is not (an unset `HOME` is exactly the coded failure above).
  The server only *reads* this directory — nothing in it writes
  there (see the
  [effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix)).
- Cause: missing artifact file, wrong ownership, missing `HOME` in
  a minimal runtime, or nobody staged the artifact (no save step
  ran out-of-band).
- Fix: stage the artifact file where the reading instance sees it
  (shared mount or copy step) and export the dir explicitly;
  `chown` readable to the image UID (operator policy); do not run
  the image as root to dodge ownership (the `USER` line is the
  isolation boundary — see [Runtime and filesystem isolation](/operate/runtime-filesystem-isolation/)).
- Check: save-then-explain round-trips on the same instance.

### Timeouts and overruns (-32001, rising `overrunCalls`)

- Symptom: JSON-RPC error `-32001` with `data.code`
  `CALL_TIMEOUT` and `limitSeconds`, and/or `/healthz`
  `overrunCalls` above `0`.
- Diagnostics: read `overrunCalls` before, during, and after load;
  find the matching journal record (the abandoned call is journaled;
  its late completion arrives as a separate record — real,
  `src/http.rs` `dispatch()`).
- Cause: the call exceeded `LAW_MCP_CALL_TIMEOUT` (default: no
  limit — real); workers are never cancelled
  (implementation-limit, no cancellation points in `law-eval`).
- Fix: set or raise the limit only for question classes that need
  it; shed load at the edge; warm lazy indexes (see
  [Capacity and scaling](/operate/capacity-scaling/) step 2).
- Check: `overrunCalls` returns to `0` after the drain; no new
  `-32001` at the planned concurrency.

### Memory and full disk (OOM kills, stalled writes)

- Symptom: container restarts with OOM status, or the instance stops
  answering while the process lives.
- Diagnostics: runtime OOM counters and RSS history; disk usage of
  the artifact directory, container log store (the journal fallback
  emits each call's JSON record to stderr — real,
  `src/journal.rs`; one primary line per served call, plus a
  possible late `OVERRUN_FINISHED` for the same call id), and
  journald quota on systemd hosts.
- Cause: no in-code memory cap (implementation-limit) met more
  concurrent or overrun calls than the box holds; or a full disk
  under artifacts/logs.
- Fix: lower edge concurrency, add identical instances per
  [Capacity and scaling](/operate/capacity-scaling/), rotate/prune logs and artifacts
  (operator policy); raise the container limit only from measured
  RSS, never blindly.
- Check: `/healthz` stays `200` through a repeated load run; RSS
  flat across runs; artifact round-trip works.

### Journal gaps (missing records, `MCP_TRUNCATED`)

- Symptom: calls with no journal record, or records with
  `MCP_TRUNCATED=oversize`.
- Diagnostics: check for `/run/systemd/journal/socket` on the host;
  read container stderr; compare record count against served calls
  minus `/healthz` hits (health checks are never journaled — real).
- Cause: no socket where the instance runs (macOS, tests,
  containers) — stderr JSON lines are the designed fallback, not a
  failure (real, `src/journal.rs`); oversize datagrams are resent
  without bodies with `MCP_TRUNCATED` (real); or recording was
  disabled with `LAW_MCP_LOG=0` (real default: on).
- Fix: collect stderr as the journal where there is no socket;
  raise `LAW_MCP_LOG_MAX_BODY` from the 64 KiB-per-side default
  (real; `0` removes the cap) only within the ~212 KiB Linux
  datagram budget (real, `src/journal.rs`); re-enable `LAW_MCP_LOG`
  unless silence is deliberate and recorded.
- Check: every served non-health call id has its primary record
  (match `OVERRUN_FINISHED` extras by call id, don't just count
  lines); missing ids are investigated as record loss, not
  assumed absent calls; no new `MCP_TRUNCATED` at production
  body sizes.

### Replay and publication (answers service, prepare refusals)

- Symptom: `law_prepare_answer` / `law_publish_answer` report
  unavailable or refuse; publication calls hang ~30 s.
- Diagnostics: confirm `LAW_ANSWERS_URL` and `LAW_ANSWERS_TOKEN`
  are both set and non-empty (either missing yields
  "not configured" — real, `src/publication.rs`); time the call
  (service client timeout is 30 s — implementation-limit,
  `src/service_client.rs`; responses cap at 32 MiB —
  implementation-limit, `src/publication.rs`).
- Cause: unwired or unreachable answers service; or a replay
  mismatch — any divergence of inputs, rights, resources, code, or
  result rejects the whole document (interface contract of the
  prepare call).
- Fix: wire and reach the service; for replay refusals, re-run with
  byte-identical inputs against the identical slice and compare
  field by field — do not edit the frozen inputs into agreement.
- Check: prepare succeeds on the unchanged inputs; publish returns
  its URL; revoke closes it.

### Serve: journal write or sync failure (`SERVE_JOURNAL_FAILED`, then `failed`)

- Symptom: one `POST /v1/ask` answers 500
  `SERVE_JOURNAL_FAILED` (`segment N: record/fsync: <os
  error>`); `/readyz` flips to 503 `{"status":"failed",
  "reason":"decision journal broken: …"}`; every later ask
  answers 503 `SERVE_NOT_READY`. `/healthz` stays 200 — the
  process lives, it just does not serve (real, `service.rs`
  `failure()` + `readiness()`).
- Diagnostics: read the `readyz` reason (it names the segment
  and the OS error); check server stderr, `df`/`dmesg`, mount
  flags (read-only?), and directory ownership on the journal
  dir.
- Cause: the record write or `sync_data` failed and the journal
  is marked broken for the rest of the process lifetime (real,
  `journal.rs` `append`): a decision is never served without
  its fsync'd record (C2). Retrying the call cannot succeed —
  the same append path fails again.
- Fix: fence ingress, fix the underlying cause (space, mount,
  permissions), then restart the process (only a restart
  clears the broken flag — a re-open rebuilds the index and
  cuts any torn tail with a log line). Verify `/readyz` 200
  with the expected pin, spot-replay one pre-failure decision
  id, then reopen ingress.
- Check: `/readyz` 200; a control ask answers 200 with a new
  fsync'd record; no new `SERVE_JOURNAL_FAILED`.

### Serve: HTTP response lost after the decision was saved (idempotency-key rule)

- Symptom: client timed out or dropped the connection; the
  decision may already be recorded — re-asking naively would
  mint a second decision for the same case.
- Diagnostics: if the first send carried an `Idempotency-Key`,
  nothing is lost: the key claims the recorded bytes.
- Cause: transport failure after commit, not a server error.
- Fix: resend the byte-identical body with the SAME
  `Idempotency-Key` (1–255 printable ASCII, no spaces — real,
  `service.rs`). The server replays the recorded bytes with
  the same `decisionId` (real, `Claim::Replay`). Never retry a
  lost response with a fresh key (second decision) or a
  changed body under the old key (422
  `SERVE_IDEMPOTENCY_CONFLICT`). A resend while the first call
  still runs answers 409 `SERVE_IDEMPOTENCY_IN_PROGRESS` —
  wait, then resend. Rule of the road: generate the key
  client-side BEFORE the first send; keys need a journal (a
  key at `--no-journal` is 400 `SERVE_IDEMPOTENCY_KEY_INVALID`).
- Check: the resend returns the same `decisionId` and
  `resultHash` as the journaled record; exactly one record
  carries that key.

### Serve: crash stop during a write (torn tail check and recovery)

- Symptom: after a SIGKILL/crash/power cut, the last segment
  may end mid-line; restart behavior decides whether anything
  was lost.
- Diagnostics: BEFORE restarting, inventory the copy with the
  read-only kit helper (see [Backup and restore](/operate/backup-restore/)):
  `kit/bin/arxo-list-decisions <journal-copy>`. `TORN_TAIL`
  on the last segment is the expected crash shape (unconfirmed
  bytes, never a decision); any `REFUSAL` names file and
  offset and fails the drill.
- Cause: the killed call left either a complete fsync'd record
  or none (real) — only the unconfirmed tail is ever partial.
- Fix: snapshot the segments if the tail bytes are evidence
  (restart truncates them and logs `cut torn tail of N bytes
  from segment M`), then restart normally, confirm `/readyz`
  200 with the same pin and `programHash`, and re-ask the
  interrupted case explicitly (new decision, new record).
- Check: `/readyz` 200; `arxo-list-decisions` on the live dir
  lists every confirmed record with no `TORN_TAIL` and no
  `REFUSAL`.

### Serve: workers unavailable or world not warm (`warming`, `NOT_READY`)

- Symptom: `/healthz` 200 but `/v1/ask` answers 503
  `SERVE_NOT_READY`, and `/readyz` answers 503
  `{"status":"warming", warmed, workers}`.
- Diagnostics: compare `warmed` vs `workers` in both endpoints
  (real fields); read server stderr for worker launch errors.
- Cause: workers still warming (normal after start or a
  `--watch-world` flip), or a worker that never warms. This is
  NOT a token, pin, or journal problem — those have their own
  codes.
- Fix: wait for `/readyz` 200 before opening ingress (the only
  correct ingress condition — never `/healthz`); if warming
  never completes, inspect the worker launch path and the
  world bytes, then restart. A `failed` status instead of
  `warming` means warmup itself errored — the reason names it.
- Check: `/readyz` 200 with the expected pin; a control ask
  answers 200.

### Serve: disk full (what blocks, what is already safe, recovery)

- Symptom: asks start failing with 500
  `SERVE_JOURNAL_FAILED` naming an `ENOSPC` write/sync error;
  `/readyz` flips to `failed` as in the journal-failure
  scenario above.
- Diagnostics: `df` on the journal filesystem; segment sizes;
  stderr for the first failing segment and offset.
- Cause: no space (or quota/inodes) for the next record or the
  next segment file. Already safe: every record fsync'd before
  the failure is a complete, replayable decision — the failing
  call has NO record and its decision was never served (C2).
  Blocked: all further asks (503 `SERVE_NOT_READY`), segment
  rotation, and any restart-time manifest write until space
  returns.
- Fix: fence ingress; free space or grow the volume (do NOT
  delete or hand-edit journal segments — evidence); restart;
  verify `/readyz` 200 + pin; spot-replay; reopen ingress.
  Set a disk alert below the failure point (operator policy).
- Check: same as the journal-failure check, plus free-space
  headroom confirmed on the journal filesystem.

## Expected result

Every red signal maps to exactly one area above, the fix lands, and
the area's check is green before traffic resumes.

## Result check

- `GET /healthz`: `200`, `"status":"ok"`,
  `"oracleGuard":"absent"`, `overrunCalls` `0` after drain.
- One CI-style smoke `law_search` through the edge returns
  non-error text.
- The incident's `X-Law-Call-Id` resolves to a complete journal
  record.

## Failures and diagnostics

- Two areas match at once (e.g. `504` + rising `overrunCalls`):
  fix the lower layer first (backend capacity), then re-test the
  upper (edge timeouts) — never both at once.
- A fix that cannot produce its check is not a fix: roll it back
  and re-triage.
- Anything outside these areas (kernel, storage array, cloud
  fabric) is host territory: record the boundary and hand over with
  the call IDs and digests attached.

## Support boundaries

- Supported: the diagnostics and checks above against the matrix in
  [Verification and support matrix](/operate/verification-support-matrix/). "Fixed" here means only "the
  area's check is green on the pinned slice".
- Implementation limits surfaced here, not worked around: 1 MiB
  request cap, no worker cancellation, no in-code memory cap, 30 s
  / 32 MiB answers-service caps.
- Out of scope: debugging client applications beyond their request
  bytes, and any host or network fabric below the container.

## Next step

Append the incident (signal, cause, fix, digest, call IDs) to the
decision journal per [Logs, audit, decision journals](/operate/logs-audit-decision-journals/); if the
cause was capacity, re-run [Capacity planning and scaling](/operate/capacity-scaling/).