# Health and observability ## Goal Tell from the outside whether the server is alive, ready, and answering correctly — using the real health endpoint, the real logs and spans, and an operator-run canary for answer correctness. ## Scope Component `law-mcp` over `--http`. Scenario: a supervisor and a monitoring stack watching a deployed instance. All hostnames and hashes below are synthetic examples. ## Applies to | Branch | Coverage in this article | |---|---| | MCP (`law-mcp-server --http`) | `/healthz`, `--healthcheck`, canary, OTLP | | `law serve` | step 7: `/healthz` (liveness) vs `/readyz` (readiness gate) + endpoint access table in [Deploy a private HTTP service](/operate/private-http-service/); serve branch of [Verification and support matrix](/operate/verification-support-matrix/) | ## Prerequisites - HTTP mode on (`--http`, real) with a known `--port` / `LAW_MCP_PORT` (default `8722`, real) and profile (`--profile` / `LAW_MCP_PROFILE`, default `all`, real). - The `law-mcp-server` binary available for the `--healthcheck` probe, or any HTTP client for `GET /healthz` (both real). ## Steps 1. Poll liveness with the real endpoint: `GET /healthz` answers `200` with a JSON object (real, `http.rs` `route`): `status:"ok"`, `pid`, `profile` (server profile name), `verificationPending:0`, `verified:0`, `engineOwnsIr:true`, `oracleGuard:"absent"`, `driftScope:"semanticHash"`, `overrunCalls` (detached workers still running). The probe is never journaled (real): poll every few seconds without log spam. 2. Run the packaged probe where a binary check fits. Real (`main.rs`): `law-mcp-server --healthcheck [--port PORT]` connects to `127.0.0.1:{port}` with a 4 s timeout, sends `GET /healthz`, and exits `0` only on `HTTP/1.1 200` plus a `"status":"ok"` body. Anything else exits `2` with `healthcheck: unexpected /healthz response` (or the connect/read error). There is no separate readiness endpoint: serving `/healthz` with the profile loaded IS ready (real behavior by construction — the profile loads before `listen`). 3. Watch saturation, not just aliveness. Real signals: `overrunCalls` rising (workers outliving the call timeout), journal `MCP_MS` percentiles climbing per `MCP_TOOL`, `CALL_TIMEOUT` / `INTERNAL` codes appearing in `MCP_CODE`, OTLP drop counter lines on stderr (queue 256 full, real). None of these has a threshold in code — thresholds are operator policy (created-example alert: `overrunCalls > 0` for 5 minutes). 4. Collect spans when call timing matters. Real (`otlp.rs`): with `LAW_MCP_OTLP_URL` set, each call emits one trace — root span `mcp.` plus stage spans — over OTLP/HTTP JSON on a background thread (3 s POST timeout, never blocks the call). Trace id = the 32-hex journal call id, so one id joins the journald record, the span, and the `X-Law-Call-Id` header; an inbound W3C `traceparent` parents the call under the client's trace. Spans carry names, outcomes, durations, and reproducibility hashes only — no call bodies (real). Without the URL, no thread and no socket is opened. There is no Prometheus endpoint in code. 5. Record versions and artifact ids with every release. Real sources: `serverInfo.version` (`0.1.0`, `server.rs`), the profile name in `/healthz`, the binary hash from the install record (article 13), and per-answer `programHash`/`caseHash`/`resultHash`/`codeHash` in responses and journal fields. Pin these four hashes in the change record; they are what `law_explain` replays against. 6. Run a canary for correctness (operator-policy-example; no canary exists in code). Created-example: every 5 minutes ask a pinned question over a pinned case (e.g. the 14-day BGB period ending 20 March) and compare `resultHash` to the recorded value. Alert on any mismatch or on a non-`OK` outcome. Keep the canary case immutable and out of the served tenant roots so tenant edits cannot break it. 7. For `law serve`, keep liveness and readiness apart (real, `service.rs`). `GET /healthz` is liveness: 200 `status:"ok"` with `pid`, pin, `programHash`, `workers`, `warmed`, queue and call counters — it answers 200 even while workers warm and even after a failure. `GET /readyz` is the readiness gate: 200 `{"status":"ready", pin, programHash, workers}` only when every worker is warm AND no failure is recorded; 503 `{"status":"warming", warmed, workers}` while workers warm; 503 `{"status":"failed", reason}` after a warmup failure or once the decision journal breaks (`reason: "decision journal broken: …"`, real). `POST /v1/ask` enforces the same gate and answers 503 `SERVE_NOT_READY` while `/readyz` is not 200 — the balancer opens ingress on `/readyz` 200 only, never on `/healthz`. None of the three observability routes (`/healthz`, `/readyz`, `/metrics`) takes a token; `/v1/world` and `/v1/ask` do (full table in [Deploy a private HTTP service](/operate/private-http-service/)). ## Expected result - Liveness: `GET /healthz` → `200` + `status:"ok"`; the packaged probe exits `0`. - Readiness: same check, because the listener starts only after the profile loads; a process that answers `/healthz` serves `/mcp`. - `law serve` readiness is separate: `/healthz` 200 proves the process lives; only `/readyz` 200 proves it serves. Ingress, supervisors, and the rollback/upgrade gates all key on `/readyz`, and a `failed` reason names the cause (warmup or broken journal) without further probing. - Correctness: the canary's `resultHash` equals the pinned value on every run; drift means the release, profile, or calendar input changed. ## Result check - `curl -s http://127.0.0.1:8722/healthz` (created-example host/port matching the real defaults) prints the real field set above with `overrunCalls:0` on an idle server. - `law-mcp-server --healthcheck --port 8722` exits `0` while the server runs, non-zero within ~4 s when it does not (real timeout). - Stop the OTLP receiver: calls still succeed with unchanged answer bytes; only the stderr drop counter moves (real non-blocking export). ## Failures and diagnostics - `GET /healthz` refused or slow while `/mcp` works: compare direct backend access against access through the edge, and check backend load (`overrunCalls`, worker saturation). A shared accept loop does not guarantee equal per-request latency — suspect the edge or the prober only after the direct-vs-edge comparison clears the backend. - `overrunCalls` stuck above zero: detached workers still computing (normal briefly) or wedged (restart the stateless process to reclaim; see article 11). - Canary hash mismatch with `OK` outcomes: inputs moved — compare `programHash` (canon changed), `caseHash` (case changed), `codeHash` (engine changed) to isolate which (real hash roles). - Infra errors vs content states (real split): HTTP statuses (`401`/`403`/`404`/`405`/`411`/`413`), JSON-RPC errors (`-32001` timeout, `-32603` internal), and `SERVICE_*` codes are transport/infra; evaluation statuses (`established`, `not established`, conflicts, `whyNot`) inside a `200`/`OK` answer are content. Alert on the first; route the second to the case owner. ## Support boundaries - "Monitored" in this article means only: the named endpoint, fields, spans, and hashes exist and behave as described. Dashboards, retention, alert routes, and SLO numbers are the operator's stack. - The health endpoint reports process state, not answer correctness — only the canary (operator-run) checks that the engine still computes the expected bytes. - `verificationPending`/`verified` are constant `0` and `engineOwnsIr`/`oracleGuard`/`driftScope` are fixed strings in this implementation (real, `http.rs`): treat them as compatibility fields, not as live verification telemetry. ## Next step - Article 11 ([Resource limits and cancellation](/operate/resource-limits-cancellation/)) for what `overrunCalls` and `CALL_TIMEOUT` imply for capacity; article 09 ([Logs, audit trail, and decision journals](/operate/logs-audit-decision-journals/)) for the journal fields behind every code named here.