Skip to content
docs
Arxo ↗

Troubleshooting

For LLMs10 sections

Turn any red signal from the deployment — a status code, a journal gap, a stuck health check — into a cause, a fix, and a re-check, without guessing: every diagnostic below names the real command, field, or file that decides.

Component law-mcp-server over HTTP (crate law-mcp), plus its image build, proxy edge, and answers service (LAW_ANSWERS_URL). All statuses, codes, and commands are real; all tokens, hostnames, and payloads are synthetic created-examples. For sizing causes, see Capacity and scaling; for what “checked” means, see Verification and support matrix.

BranchCoverage in this article
MCP (law-mcp-server over HTTP)full symptom → fix tables
law servefive serve failure scenarios below (journal write, lost response, crash tail, cold workers, full disk), plus the per-procedure tables in Deploy a private HTTP service, Upgrades and compatibility, Rollback
  • Shell access to the host or the container runtime; curl and jq (operator tools, not shipped in the scratch image).
  • The instance’s X-Law-Call-Id for the failing call (every response carries one — real, src/http.rs) and its journal stream.
  • The image digest and profile the instance should run (see Upgrades and compatibility for pinning).

Triage in this order; each area below follows symptom → diagnostics → cause → fix → check.

  • Symptom: 401 with WWW-Authenticate: Bearer resource="law-mcp", or the process exits 2 at startup refusing a non-loopback bind.
  • Diagnostics: repeat the call with -v and confirm the Authorization: Bearer <token> header is byte-exact; read stderr for the bind-refusal line naming the host. -v prints the live secret: never paste that output into a ticket, chat, or CI log — compare locally, then discard the terminal scrollback; for a shareable trace, re-run with a revoked synthetic token instead.
  • Cause: wrong or missing token (the server compares the full Bearer <token> string — real, src/http.rs guard()), or a public bind with neither LAW_MCP_TOKEN nor LAW_MCP_PUBLIC=1 (refused by config validation).
  • Fix: set the token via --token / LAW_MCP_TOKEN (real), or set LAW_MCP_PUBLIC=1 only as a deliberate, recorded decision.
  • Check: the same call returns 200; startup prints the law-mcp-http: listening … line to stderr.

Proxy and edge (502 / 504 / wrong client address)

Section titled “Proxy and edge (502 / 504 / wrong client address)”
  • Symptom: edge 502/504, or journal client fields that all read as one address.
  • Diagnostics: curl the backend GET /healthz directly (bypassing the edge); compare with the edge response; check edge proxy_read_timeout against LAW_MCP_CALL_TIMEOUT.
  • Cause: backend down (see Memory, Filesystem); timeout inversion — the edge gave up first (the reference-setting is call limit 110 s below proxy 120 s, src/main.rs); or the edge does not forward the client address (the server takes the first X-Forwarded-For entry — real, src/http.rs context()).
  • Fix: revive the backend; restore the timeout ordering; forward X-Forwarded-For and X-Request-Id at the edge (operator policy; header trust holds by layout — one public path through the edge).
  • Check: direct and edge /healthz both 200 with "status":"ok"; journal client varies per caller.

Slice pins (--check-pin exit 1, image refuses profile)

Section titled “Slice pins (--check-pin exit 1, image refuses profile)”
  • Symptom: image_tree.py <profile> --check-pin <dir> exits 1, or the build step fails before docker build.
  • Diagnostics: run the check; it prints the path-by-path diff against the slice pin file for the profile (real, schema arxo.mcp-slice-pin/1).
  • Cause: tree contents drifted from the pin; profile all requested (refused: images are built for a slice — real, image_tree.py); missing CLIR for a prefix; pin schema mismatch.
  • Fix: regenerate the tree from the same commit and re-pin only after reviewing the diff file; never hand-edit the tree into silence. Never build an image for all.
  • Check: --check-pin exits 0; the image build repeats the same two commands plus --manifest.

Missing resources (startup exit 2, empty answers)

Section titled “Missing resources (startup exit 2, empty answers)”
  • Symptom: process exits 2 with law-mcp: … on stderr, or tools answer empty on a fresh slice.
  • Diagnostics: read the full stderr line (all startup failures exit 2 — real); list the profile files (*.json) for the requested name (real examples: kz, phys, chem, code-civil, de-bgb); verify the slice carries its canon and calendarPackage per the profile JSON.
  • Cause: unknown profile name, --root pointing away from the tree (find_root needs the checkout layout), or a slice whose canon the question does not touch.
  • Fix: correct --root/--profile; rebuild the slice for the needed profile; re-ask against the profile whose canon covers the question.
  • Check: startup listening line names the intended profile; /healthz "profile" matches it; the CI-style law_search smoke call returns non-error text.

Incompatible versions (400 protocol version, stale client)

Section titled “Incompatible versions (400 protocol version, stale client)”
  • Symptom: 400 naming the protocol version and the known list, or initialize negotiating down unexpectedly.
  • Diagnostics: read the known = list in the 400 body (real, src/http.rs guard()); compare with the client’s mcp-protocol-version header.
  • Cause: client sends a revision outside 2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05 (real, SUPPORTED_PROTOCOL_VERSIONS) on a non-initialize call.
  • Fix: upgrade or pin the client to a listed revision; initialize alone is exempt by design.
  • Check: the same call returns 200; /healthz is unaffected (it never negotiates versions).

Filesystem and permissions (artifact failures, read-only tree)

Section titled “Filesystem and permissions (artifact failures, read-only tree)”
  • Symptom: explain of a saved answer fails with ANSWER_RESOURCE_UNAVAILABLE, or the container cannot start.
  • Diagnostics: confirm LAW_MCP_ARTIFACT_DIR, else $HOME/.cache/law-dsl/mcp-answers/<profile> (real, src/execution.rs), exists and is readable by 10001:0 (real image user, Dockerfile); check HOME is set when the variable is not (an unset HOME is exactly the coded failure above). The server only reads this directory — nothing in it writes there (see the effects matrix).
  • Cause: missing artifact file, wrong ownership, missing HOME in a minimal runtime, or nobody staged the artifact (no save step ran out-of-band).
  • Fix: stage the artifact file where the reading instance sees it (shared mount or copy step) and export the dir explicitly; chown readable to the image UID (operator policy); do not run the image as root to dodge ownership (the USER line is the isolation boundary — see Runtime and filesystem isolation).
  • Check: save-then-explain round-trips on the same instance.

Timeouts and overruns (-32001, rising overrunCalls)

Section titled “Timeouts and overruns (-32001, rising overrunCalls)”
  • Symptom: JSON-RPC error -32001 with data.code CALL_TIMEOUT and limitSeconds, and/or /healthz overrunCalls above 0.
  • Diagnostics: read overrunCalls before, during, and after load; find the matching journal record (the abandoned call is journaled; its late completion arrives as a separate record — real, src/http.rs dispatch()).
  • Cause: the call exceeded LAW_MCP_CALL_TIMEOUT (default: no limit — real); workers are never cancelled (implementation-limit, no cancellation points in law-eval).
  • Fix: set or raise the limit only for question classes that need it; shed load at the edge; warm lazy indexes (see Capacity and scaling step 2).
  • Check: overrunCalls returns to 0 after the drain; no new -32001 at the planned concurrency.

Memory and full disk (OOM kills, stalled writes)

Section titled “Memory and full disk (OOM kills, stalled writes)”
  • Symptom: container restarts with OOM status, or the instance stops answering while the process lives.
  • Diagnostics: runtime OOM counters and RSS history; disk usage of the artifact directory, container log store (the journal fallback emits each call’s JSON record to stderr — real, src/journal.rs; one primary line per served call, plus a possible late OVERRUN_FINISHED for the same call id), and journald quota on systemd hosts.
  • Cause: no in-code memory cap (implementation-limit) met more concurrent or overrun calls than the box holds; or a full disk under artifacts/logs.
  • Fix: lower edge concurrency, add identical instances per Capacity and scaling, rotate/prune logs and artifacts (operator policy); raise the container limit only from measured RSS, never blindly.
  • Check: /healthz stays 200 through a repeated load run; RSS flat across runs; artifact round-trip works.

Journal gaps (missing records, MCP_TRUNCATED)

Section titled “Journal gaps (missing records, MCP_TRUNCATED)”
  • Symptom: calls with no journal record, or records with MCP_TRUNCATED=oversize.
  • Diagnostics: check for /run/systemd/journal/socket on the host; read container stderr; compare record count against served calls minus /healthz hits (health checks are never journaled — real).
  • Cause: no socket where the instance runs (macOS, tests, containers) — stderr JSON lines are the designed fallback, not a failure (real, src/journal.rs); oversize datagrams are resent without bodies with MCP_TRUNCATED (real); or recording was disabled with LAW_MCP_LOG=0 (real default: on).
  • Fix: collect stderr as the journal where there is no socket; raise LAW_MCP_LOG_MAX_BODY from the 64 KiB-per-side default (real; 0 removes the cap) only within the ~212 KiB Linux datagram budget (real, src/journal.rs); re-enable LAW_MCP_LOG unless silence is deliberate and recorded.
  • Check: every served non-health call id has its primary record (match OVERRUN_FINISHED extras by call id, don’t just count lines); missing ids are investigated as record loss, not assumed absent calls; no new MCP_TRUNCATED at production body sizes.

Replay and publication (answers service, prepare refusals)

Section titled “Replay and publication (answers service, prepare refusals)”
  • Symptom: law_prepare_answer / law_publish_answer report unavailable or refuse; publication calls hang ~30 s.
  • Diagnostics: confirm LAW_ANSWERS_URL and LAW_ANSWERS_TOKEN are both set and non-empty (either missing yields “not configured” — real, src/publication.rs); time the call (service client timeout is 30 s — implementation-limit, src/service_client.rs; responses cap at 32 MiB — implementation-limit, src/publication.rs).
  • Cause: unwired or unreachable answers service; or a replay mismatch — any divergence of inputs, rights, resources, code, or result rejects the whole document (interface contract of the prepare call).
  • Fix: wire and reach the service; for replay refusals, re-run with byte-identical inputs against the identical slice and compare field by field — do not edit the frozen inputs into agreement.
  • Check: prepare succeeds on the unchanged inputs; publish returns its URL; revoke closes it.

Serve: journal write or sync failure (SERVE_JOURNAL_FAILED, then failed)

Section titled “Serve: journal write or sync failure (SERVE_JOURNAL_FAILED, then failed)”
  • Symptom: one POST /v1/ask answers 500 SERVE_JOURNAL_FAILED (segment N: record/fsync: <os error>); /readyz flips to 503 {"status":"failed", "reason":"decision journal broken: …"}; every later ask answers 503 SERVE_NOT_READY. /healthz stays 200 — the process lives, it just does not serve (real, service.rs failure() + readiness()).
  • Diagnostics: read the readyz reason (it names the segment and the OS error); check server stderr, df/dmesg, mount flags (read-only?), and directory ownership on the journal dir.
  • Cause: the record write or sync_data failed and the journal is marked broken for the rest of the process lifetime (real, journal.rs append): a decision is never served without its fsync’d record (C2). Retrying the call cannot succeed — the same append path fails again.
  • Fix: fence ingress, fix the underlying cause (space, mount, permissions), then restart the process (only a restart clears the broken flag — a re-open rebuilds the index and cuts any torn tail with a log line). Verify /readyz 200 with the expected pin, spot-replay one pre-failure decision id, then reopen ingress.
  • Check: /readyz 200; a control ask answers 200 with a new fsync’d record; no new SERVE_JOURNAL_FAILED.

Serve: HTTP response lost after the decision was saved (idempotency-key rule)

Section titled “Serve: HTTP response lost after the decision was saved (idempotency-key rule)”
  • Symptom: client timed out or dropped the connection; the decision may already be recorded — re-asking naively would mint a second decision for the same case.
  • Diagnostics: if the first send carried an Idempotency-Key, nothing is lost: the key claims the recorded bytes.
  • Cause: transport failure after commit, not a server error.
  • Fix: resend the byte-identical body with the SAME Idempotency-Key (1–255 printable ASCII, no spaces — real, service.rs). The server replays the recorded bytes with the same decisionId (real, Claim::Replay). Never retry a lost response with a fresh key (second decision) or a changed body under the old key (422 SERVE_IDEMPOTENCY_CONFLICT). A resend while the first call still runs answers 409 SERVE_IDEMPOTENCY_IN_PROGRESS — wait, then resend. Rule of the road: generate the key client-side BEFORE the first send; keys need a journal (a key at --no-journal is 400 SERVE_IDEMPOTENCY_KEY_INVALID).
  • Check: the resend returns the same decisionId and resultHash as the journaled record; exactly one record carries that key.

Serve: crash stop during a write (torn tail check and recovery)

Section titled “Serve: crash stop during a write (torn tail check and recovery)”
  • Symptom: after a SIGKILL/crash/power cut, the last segment may end mid-line; restart behavior decides whether anything was lost.
  • Diagnostics: BEFORE restarting, inventory the copy with the read-only kit helper (see Backup and restore): kit/bin/arxo-list-decisions <journal-copy>. TORN_TAIL on the last segment is the expected crash shape (unconfirmed bytes, never a decision); any REFUSAL names file and offset and fails the drill.
  • Cause: the killed call left either a complete fsync’d record or none (real) — only the unconfirmed tail is ever partial.
  • Fix: snapshot the segments if the tail bytes are evidence (restart truncates them and logs cut torn tail of N bytes from segment M), then restart normally, confirm /readyz 200 with the same pin and programHash, and re-ask the interrupted case explicitly (new decision, new record).
  • Check: /readyz 200; arxo-list-decisions on the live dir lists every confirmed record with no TORN_TAIL and no REFUSAL.

Serve: workers unavailable or world not warm (warming, NOT_READY)

Section titled “Serve: workers unavailable or world not warm (warming, NOT_READY)”
  • Symptom: /healthz 200 but /v1/ask answers 503 SERVE_NOT_READY, and /readyz answers 503 {"status":"warming", warmed, workers}.
  • Diagnostics: compare warmed vs workers in both endpoints (real fields); read server stderr for worker launch errors.
  • Cause: workers still warming (normal after start or a --watch-world flip), or a worker that never warms. This is NOT a token, pin, or journal problem — those have their own codes.
  • Fix: wait for /readyz 200 before opening ingress (the only correct ingress condition — never /healthz); if warming never completes, inspect the worker launch path and the world bytes, then restart. A failed status instead of warming means warmup itself errored — the reason names it.
  • Check: /readyz 200 with the expected pin; a control ask answers 200.

Serve: disk full (what blocks, what is already safe, recovery)

Section titled “Serve: disk full (what blocks, what is already safe, recovery)”
  • Symptom: asks start failing with 500 SERVE_JOURNAL_FAILED naming an ENOSPC write/sync error; /readyz flips to failed as in the journal-failure scenario above.
  • Diagnostics: df on the journal filesystem; segment sizes; stderr for the first failing segment and offset.
  • Cause: no space (or quota/inodes) for the next record or the next segment file. Already safe: every record fsync’d before the failure is a complete, replayable decision — the failing call has NO record and its decision was never served (C2). Blocked: all further asks (503 SERVE_NOT_READY), segment rotation, and any restart-time manifest write until space returns.
  • Fix: fence ingress; free space or grow the volume (do NOT delete or hand-edit journal segments — evidence); restart; verify /readyz 200 + pin; spot-replay; reopen ingress. Set a disk alert below the failure point (operator policy).
  • Check: same as the journal-failure check, plus free-space headroom confirmed on the journal filesystem.

Every red signal maps to exactly one area above, the fix lands, and the area’s check is green before traffic resumes.

  • GET /healthz: 200, "status":"ok", "oracleGuard":"absent", overrunCalls 0 after drain.
  • One CI-style smoke law_search through the edge returns non-error text.
  • The incident’s X-Law-Call-Id resolves to a complete journal record.
  • Two areas match at once (e.g. 504 + rising overrunCalls): fix the lower layer first (backend capacity), then re-test the upper (edge timeouts) — never both at once.
  • A fix that cannot produce its check is not a fix: roll it back and re-triage.
  • Anything outside these areas (kernel, storage array, cloud fabric) is host territory: record the boundary and hand over with the call IDs and digests attached.
  • Supported: the diagnostics and checks above against the matrix in Verification and support matrix. “Fixed” here means only “the area’s check is green on the pinned slice”.
  • Implementation limits surfaced here, not worked around: 1 MiB request cap, no worker cancellation, no in-code memory cap, 30 s / 32 MiB answers-service caps.
  • Out of scope: debugging client applications beyond their request bytes, and any host or network fabric below the container.

Append the incident (signal, cause, fix, digest, call IDs) to the decision journal per Logs, audit, decision journals; if the cause was capacity, re-run Capacity planning and scaling.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.