Skip to content
docs
Arxo ↗

Health and observability

For LLMs10 sections

Tell from the outside whether the server is alive, ready, and answering correctly — using the real health endpoint, the real logs and spans, and an operator-run canary for answer correctness.

Component law-mcp over --http. Scenario: a supervisor and a monitoring stack watching a deployed instance. All hostnames and hashes below are synthetic examples.

BranchCoverage in this article
MCP (law-mcp-server --http)/healthz, --healthcheck, canary, OTLP
law servestep 7: /healthz (liveness) vs /readyz (readiness gate) + endpoint access table in Deploy a private HTTP service; serve branch of Verification and support matrix
  • HTTP mode on (--http, real) with a known --port / LAW_MCP_PORT (default 8722, real) and profile (--profile / LAW_MCP_PROFILE, default all, real).
  • The law-mcp-server binary available for the --healthcheck probe, or any HTTP client for GET /healthz (both real).
  1. Poll liveness with the real endpoint: GET /healthz answers 200 with a JSON object (real, http.rs route): status:"ok", pid, profile (server profile name), verificationPending:0, verified:0, engineOwnsIr:true, oracleGuard:"absent", driftScope:"semanticHash", overrunCalls (detached workers still running). The probe is never journaled (real): poll every few seconds without log spam.
  2. Run the packaged probe where a binary check fits. Real (main.rs): law-mcp-server --healthcheck [--port PORT] connects to 127.0.0.1:{port} with a 4 s timeout, sends GET /healthz, and exits 0 only on HTTP/1.1 200 plus a "status":"ok" body. Anything else exits 2 with healthcheck: unexpected /healthz response (or the connect/read error). There is no separate readiness endpoint: serving /healthz with the profile loaded IS ready (real behavior by construction — the profile loads before listen).
  3. Watch saturation, not just aliveness. Real signals: overrunCalls rising (workers outliving the call timeout), journal MCP_MS percentiles climbing per MCP_TOOL, CALL_TIMEOUT / INTERNAL codes appearing in MCP_CODE, OTLP drop counter lines on stderr (queue 256 full, real). None of these has a threshold in code — thresholds are operator policy (created-example alert: overrunCalls > 0 for 5 minutes).
  4. Collect spans when call timing matters. Real (otlp.rs): with LAW_MCP_OTLP_URL set, each call emits one trace — root span mcp.<tool> plus stage spans — over OTLP/HTTP JSON on a background thread (3 s POST timeout, never blocks the call). Trace id = the 32-hex journal call id, so one id joins the journald record, the span, and the X-Law-Call-Id header; an inbound W3C traceparent parents the call under the client’s trace. Spans carry names, outcomes, durations, and reproducibility hashes only — no call bodies (real). Without the URL, no thread and no socket is opened. There is no Prometheus endpoint in code.
  5. Record versions and artifact ids with every release. Real sources: serverInfo.version (0.1.0, server.rs), the profile name in /healthz, the binary hash from the install record (article 13), and per-answer programHash/caseHash/resultHash/codeHash in responses and journal fields. Pin these four hashes in the change record; they are what law_explain replays against.
  6. Run a canary for correctness (operator-policy-example; no canary exists in code). Created-example: every 5 minutes ask a pinned question over a pinned case (e.g. the 14-day BGB period ending 20 March) and compare resultHash to the recorded value. Alert on any mismatch or on a non-OK outcome. Keep the canary case immutable and out of the served tenant roots so tenant edits cannot break it.
  7. For law serve, keep liveness and readiness apart (real, service.rs). GET /healthz is liveness: 200 status:"ok" with pid, pin, programHash, workers, warmed, queue and call counters — it answers 200 even while workers warm and even after a failure. GET /readyz is the readiness gate: 200 {"status":"ready", pin, programHash, workers} only when every worker is warm AND no failure is recorded; 503 {"status":"warming", warmed, workers} while workers warm; 503 {"status":"failed", reason} after a warmup failure or once the decision journal breaks (reason: "decision journal broken: …", real). POST /v1/ask enforces the same gate and answers 503 SERVE_NOT_READY while /readyz is not 200 — the balancer opens ingress on /readyz 200 only, never on /healthz. None of the three observability routes (/healthz, /readyz, /metrics) takes a token; /v1/world and /v1/ask do (full table in Deploy a private HTTP service).
  • Liveness: GET /healthz → 200 + status:"ok"; the packaged probe exits 0.
  • Readiness: same check, because the listener starts only after the profile loads; a process that answers /healthz serves /mcp.
  • law serve readiness is separate: /healthz 200 proves the process lives; only /readyz 200 proves it serves. Ingress, supervisors, and the rollback/upgrade gates all key on /readyz, and a failed reason names the cause (warmup or broken journal) without further probing.
  • Correctness: the canary’s resultHash equals the pinned value on every run; drift means the release, profile, or calendar input changed.
  • curl -s http://127.0.0.1:8722/healthz (created-example host/port matching the real defaults) prints the real field set above with overrunCalls:0 on an idle server.
  • law-mcp-server --healthcheck --port 8722 exits 0 while the server runs, non-zero within ~4 s when it does not (real timeout).
  • Stop the OTLP receiver: calls still succeed with unchanged answer bytes; only the stderr drop counter moves (real non-blocking export).
  • GET /healthz refused or slow while /mcp works: compare direct backend access against access through the edge, and check backend load (overrunCalls, worker saturation). A shared accept loop does not guarantee equal per-request latency — suspect the edge or the prober only after the direct-vs-edge comparison clears the backend.
  • overrunCalls stuck above zero: detached workers still computing (normal briefly) or wedged (restart the stateless process to reclaim; see article 11).
  • Canary hash mismatch with OK outcomes: inputs moved — compare programHash (canon changed), caseHash (case changed), codeHash (engine changed) to isolate which (real hash roles).
  • Infra errors vs content states (real split): HTTP statuses (401/403/404/405/411/413), JSON-RPC errors (-32001 timeout, -32603 internal), and SERVICE_* codes are transport/infra; evaluation statuses (established, not established, conflicts, whyNot) inside a 200/OK answer are content. Alert on the first; route the second to the case owner.
  • “Monitored” in this article means only: the named endpoint, fields, spans, and hashes exist and behave as described. Dashboards, retention, alert routes, and SLO numbers are the operator’s stack.
  • The health endpoint reports process state, not answer correctness — only the canary (operator-run) checks that the engine still computes the expected bytes.
  • verificationPending/verified are constant 0 and engineOwnsIr/oracleGuard/driftScope are fixed strings in this implementation (real, http.rs): treat them as compatibility fields, not as live verification telemetry.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.