Markdown for LLMs
Health and observability
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Health and observability
## Goal
Tell from the outside whether the server is alive, ready, and
answering correctly — using the real health endpoint, the real logs
and spans, and an operator-run canary for answer correctness.
## Scope
Component `law-mcp` over `--http`. Scenario: a supervisor and a
monitoring stack watching a deployed instance. All hostnames and
hashes below are synthetic examples.
## Applies to
| Branch | Coverage in this article |
|---|---|
| MCP (`law-mcp-server --http`) | `/healthz`, `--healthcheck`, canary, OTLP |
| `law serve` | step 7: `/healthz` (liveness) vs `/readyz` (readiness gate) + endpoint access table in [Deploy a private HTTP service](/operate/private-http-service/); serve branch of [Verification and support matrix](/operate/verification-support-matrix/) |
## Prerequisites
- HTTP mode on (`--http`, real) with a known `--port` / `LAW_MCP_PORT`
(default `8722`, real) and profile (`--profile` / `LAW_MCP_PROFILE`,
default `all`, real).
- The `law-mcp-server` binary available for the `--healthcheck` probe, or
any HTTP client for `GET /healthz` (both real).
## Steps
1. Poll liveness with the real endpoint: `GET /healthz` answers
`200` with a JSON object (real, `http.rs` `route`):
`status:"ok"`, `pid`, `profile` (server profile name),
`verificationPending:0`, `verified:0`, `engineOwnsIr:true`,
`oracleGuard:"absent"`, `driftScope:"semanticHash"`,
`overrunCalls` (detached workers still running). The probe is
never journaled (real): poll every few seconds without log spam.
2. Run the packaged probe where a binary check fits. Real
(`main.rs`): `law-mcp-server --healthcheck [--port PORT]` connects to
`127.0.0.1:{port}` with a 4 s timeout, sends
`GET /healthz`, and exits `0` only on `HTTP/1.1 200` plus a
`"status":"ok"` body. Anything else exits `2` with
`healthcheck: unexpected /healthz response` (or the connect/read
error). There is no separate readiness endpoint: serving
`/healthz` with the profile loaded IS ready (real behavior by
construction — the profile loads before `listen`).
3. Watch saturation, not just aliveness. Real signals: `overrunCalls`
rising (workers outliving the call timeout), journal `MCP_MS`
percentiles climbing per `MCP_TOOL`, `CALL_TIMEOUT` / `INTERNAL`
codes appearing in `MCP_CODE`, OTLP drop counter lines on stderr
(queue 256 full, real). None of these has a threshold in code —
thresholds are operator policy (created-example alert:
`overrunCalls > 0` for 5 minutes).
4. Collect spans when call timing matters. Real (`otlp.rs`): with
`LAW_MCP_OTLP_URL` set, each call emits one trace — root span
`mcp.<tool>` plus stage spans — over OTLP/HTTP JSON on a
background thread (3 s POST timeout, never blocks the call).
Trace id = the 32-hex journal call id, so one id joins the
journald record, the span, and the `X-Law-Call-Id` header; an
inbound W3C `traceparent` parents the call under the client's
trace. Spans carry names, outcomes, durations, and
reproducibility hashes only — no call bodies (real). Without the
URL, no thread and no socket is opened. There is no Prometheus
endpoint in code.
5. Record versions and artifact ids with every release. Real sources:
`serverInfo.version` (`0.1.0`, `server.rs`), the profile name in
`/healthz`, the binary hash from the install record (article 13),
and per-answer `programHash`/`caseHash`/`resultHash`/`codeHash`
in responses and journal fields. Pin these four hashes in the
change record; they are what `law_explain` replays against.
6. Run a canary for correctness (operator-policy-example; no canary
exists in code). Created-example: every 5 minutes ask a pinned
question over a pinned case (e.g. the 14-day BGB period ending
20 March) and compare `resultHash` to the recorded value. Alert
on any mismatch or on a non-`OK` outcome. Keep the canary case
immutable and out of the served tenant roots so tenant edits
cannot break it.
7. For `law serve`, keep liveness and readiness apart (real,
`service.rs`). `GET /healthz` is liveness: 200 `status:"ok"`
with `pid`, pin, `programHash`, `workers`, `warmed`, queue
and call counters — it answers 200 even while workers warm
and even after a failure. `GET /readyz` is the readiness
gate: 200 `{"status":"ready", pin, programHash, workers}`
only when every worker is warm AND no failure is recorded;
503 `{"status":"warming", warmed, workers}` while workers
warm; 503 `{"status":"failed", reason}` after a warmup
failure or once the decision journal breaks
(`reason: "decision journal broken: …"`, real). `POST
/v1/ask` enforces the same gate and answers 503
`SERVE_NOT_READY` while `/readyz` is not 200 — the balancer
opens ingress on `/readyz` 200 only, never on `/healthz`.
None of the three observability routes (`/healthz`,
`/readyz`, `/metrics`) takes a token; `/v1/world` and
`/v1/ask` do (full table in [Deploy a private HTTP
service](/operate/private-http-service/)).
## Expected result
- Liveness: `GET /healthz` → `200` + `status:"ok"`; the packaged
probe exits `0`.
- Readiness: same check, because the listener starts only after the
profile loads; a process that answers `/healthz` serves `/mcp`.
- `law serve` readiness is separate: `/healthz` 200 proves the
process lives; only `/readyz` 200 proves it serves. Ingress,
supervisors, and the rollback/upgrade gates all key on
`/readyz`, and a `failed` reason names the cause (warmup or
broken journal) without further probing.
- Correctness: the canary's `resultHash` equals the pinned value on
every run; drift means the release, profile, or calendar input
changed.
## Result check
- `curl -s http://127.0.0.1:8722/healthz` (created-example host/port
matching the real defaults) prints the real field set above with
`overrunCalls:0` on an idle server.
- `law-mcp-server --healthcheck --port 8722` exits `0` while the server
runs, non-zero within ~4 s when it does not (real timeout).
- Stop the OTLP receiver: calls still succeed with unchanged answer
bytes; only the stderr drop counter moves (real non-blocking
export).
## Failures and diagnostics
- `GET /healthz` refused or slow while `/mcp` works: compare direct
backend access against access through the edge, and check backend
load (`overrunCalls`, worker saturation). A shared accept loop
does not guarantee equal per-request latency — suspect the edge
or the prober only after the direct-vs-edge comparison clears
the backend.
- `overrunCalls` stuck above zero: detached workers still computing
(normal briefly) or wedged (restart the stateless process to
reclaim; see article 11).
- Canary hash mismatch with `OK` outcomes: inputs moved — compare
`programHash` (canon changed), `caseHash` (case changed),
`codeHash` (engine changed) to isolate which (real hash roles).
- Infra errors vs content states (real split): HTTP statuses
(`401`/`403`/`404`/`405`/`411`/`413`), JSON-RPC errors
(`-32001` timeout, `-32603` internal), and `SERVICE_*` codes are
transport/infra; evaluation statuses (`established`,
`not established`, conflicts, `whyNot`) inside a `200`/`OK`
answer are content. Alert on the first; route the second to the
case owner.
## Support boundaries
- "Monitored" in this article means only: the named endpoint,
fields, spans, and hashes exist and behave as described.
Dashboards, retention, alert routes, and SLO numbers are the
operator's stack.
- The health endpoint reports process state, not answer
correctness — only the canary (operator-run) checks that the
engine still computes the expected bytes.
- `verificationPending`/`verified` are constant `0` and
`engineOwnsIr`/`oracleGuard`/`driftScope` are fixed strings in
this implementation (real, `http.rs`): treat them as
compatibility fields, not as live verification telemetry.
## Next step
- Article 11 ([Resource limits and cancellation](/operate/resource-limits-cancellation/)) for what `overrunCalls` and `CALL_TIMEOUT` imply for
capacity; article 09 ([Logs, audit trail, and decision journals](/operate/logs-audit-decision-journals/)) for the journal fields behind every code
named here.