docs← Back to article

Markdown for LLMs

Health and observability

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Health and observability

## Goal

Tell from the outside whether the server is alive, ready, and
answering correctly — using the real health endpoint, the real logs
and spans, and an operator-run canary for answer correctness.

## Scope

Component `law-mcp` over `--http`. Scenario: a supervisor and a
monitoring stack watching a deployed instance. All hostnames and
hashes below are synthetic examples.

## Applies to

| Branch | Coverage in this article |
|---|---|
| MCP (`law-mcp-server --http`) | `/healthz`, `--healthcheck`, canary, OTLP |
| `law serve` | step 7: `/healthz` (liveness) vs `/readyz` (readiness gate) + endpoint access table in [Deploy a private HTTP service](/operate/private-http-service/); serve branch of [Verification and support matrix](/operate/verification-support-matrix/) |

## Prerequisites

- HTTP mode on (`--http`, real) with a known `--port` / `LAW_MCP_PORT`
  (default `8722`, real) and profile (`--profile` / `LAW_MCP_PROFILE`,
  default `all`, real).
- The `law-mcp-server` binary available for the `--healthcheck` probe, or
  any HTTP client for `GET /healthz` (both real).

## Steps

1. Poll liveness with the real endpoint: `GET /healthz` answers
   `200` with a JSON object (real, `http.rs` `route`):
   `status:"ok"`, `pid`, `profile` (server profile name),
   `verificationPending:0`, `verified:0`, `engineOwnsIr:true`,
   `oracleGuard:"absent"`, `driftScope:"semanticHash"`,
   `overrunCalls` (detached workers still running). The probe is
   never journaled (real): poll every few seconds without log spam.
2. Run the packaged probe where a binary check fits. Real
   (`main.rs`): `law-mcp-server --healthcheck [--port PORT]` connects to
   `127.0.0.1:{port}` with a 4 s timeout, sends
   `GET /healthz`, and exits `0` only on `HTTP/1.1 200` plus a
   `"status":"ok"` body. Anything else exits `2` with
   `healthcheck: unexpected /healthz response` (or the connect/read
   error). There is no separate readiness endpoint: serving
   `/healthz` with the profile loaded IS ready (real behavior by
   construction — the profile loads before `listen`).
3. Watch saturation, not just aliveness. Real signals: `overrunCalls`
   rising (workers outliving the call timeout), journal `MCP_MS`
   percentiles climbing per `MCP_TOOL`, `CALL_TIMEOUT` / `INTERNAL`
   codes appearing in `MCP_CODE`, OTLP drop counter lines on stderr
   (queue 256 full, real). None of these has a threshold in code —
   thresholds are operator policy (created-example alert:
   `overrunCalls > 0` for 5 minutes).
4. Collect spans when call timing matters. Real (`otlp.rs`): with
   `LAW_MCP_OTLP_URL` set, each call emits one trace — root span
   `mcp.<tool>` plus stage spans — over OTLP/HTTP JSON on a
   background thread (3 s POST timeout, never blocks the call).
   Trace id = the 32-hex journal call id, so one id joins the
   journald record, the span, and the `X-Law-Call-Id` header; an
   inbound W3C `traceparent` parents the call under the client's
   trace. Spans carry names, outcomes, durations, and
   reproducibility hashes only — no call bodies (real). Without the
   URL, no thread and no socket is opened. There is no Prometheus
   endpoint in code.
5. Record versions and artifact ids with every release. Real sources:
   `serverInfo.version` (`0.1.0`, `server.rs`), the profile name in
   `/healthz`, the binary hash from the install record (article 13),
   and per-answer `programHash`/`caseHash`/`resultHash`/`codeHash`
   in responses and journal fields. Pin these four hashes in the
   change record; they are what `law_explain` replays against.
6. Run a canary for correctness (operator-policy-example; no canary
   exists in code). Created-example: every 5 minutes ask a pinned
   question over a pinned case (e.g. the 14-day BGB period ending
   20 March) and compare `resultHash` to the recorded value. Alert
   on any mismatch or on a non-`OK` outcome. Keep the canary case
   immutable and out of the served tenant roots so tenant edits
   cannot break it.
7. For `law serve`, keep liveness and readiness apart (real,
   `service.rs`). `GET /healthz` is liveness: 200 `status:"ok"`
   with `pid`, pin, `programHash`, `workers`, `warmed`, queue
   and call counters — it answers 200 even while workers warm
   and even after a failure. `GET /readyz` is the readiness
   gate: 200 `{"status":"ready", pin, programHash, workers}`
   only when every worker is warm AND no failure is recorded;
   503 `{"status":"warming", warmed, workers}` while workers
   warm; 503 `{"status":"failed", reason}` after a warmup
   failure or once the decision journal breaks
   (`reason: "decision journal broken: …"`, real). `POST
   /v1/ask` enforces the same gate and answers 503
   `SERVE_NOT_READY` while `/readyz` is not 200 — the balancer
   opens ingress on `/readyz` 200 only, never on `/healthz`.
   None of the three observability routes (`/healthz`,
   `/readyz`, `/metrics`) takes a token; `/v1/world` and
   `/v1/ask` do (full table in [Deploy a private HTTP
   service](/operate/private-http-service/)).

## Expected result

- Liveness: `GET /healthz` → `200` + `status:"ok"`; the packaged
  probe exits `0`.
- Readiness: same check, because the listener starts only after the
  profile loads; a process that answers `/healthz` serves `/mcp`.
- `law serve` readiness is separate: `/healthz` 200 proves the
  process lives; only `/readyz` 200 proves it serves. Ingress,
  supervisors, and the rollback/upgrade gates all key on
  `/readyz`, and a `failed` reason names the cause (warmup or
  broken journal) without further probing.
- Correctness: the canary's `resultHash` equals the pinned value on
  every run; drift means the release, profile, or calendar input
  changed.

## Result check

- `curl -s http://127.0.0.1:8722/healthz` (created-example host/port
  matching the real defaults) prints the real field set above with
  `overrunCalls:0` on an idle server.
- `law-mcp-server --healthcheck --port 8722` exits `0` while the server
  runs, non-zero within ~4 s when it does not (real timeout).
- Stop the OTLP receiver: calls still succeed with unchanged answer
  bytes; only the stderr drop counter moves (real non-blocking
  export).

## Failures and diagnostics

- `GET /healthz` refused or slow while `/mcp` works: compare direct
  backend access against access through the edge, and check backend
  load (`overrunCalls`, worker saturation). A shared accept loop
  does not guarantee equal per-request latency — suspect the edge
  or the prober only after the direct-vs-edge comparison clears
  the backend.
- `overrunCalls` stuck above zero: detached workers still computing
  (normal briefly) or wedged (restart the stateless process to
  reclaim; see article 11).
- Canary hash mismatch with `OK` outcomes: inputs moved — compare
  `programHash` (canon changed), `caseHash` (case changed),
  `codeHash` (engine changed) to isolate which (real hash roles).
- Infra errors vs content states (real split): HTTP statuses
  (`401`/`403`/`404`/`405`/`411`/`413`), JSON-RPC errors
  (`-32001` timeout, `-32603` internal), and `SERVICE_*` codes are
  transport/infra; evaluation statuses (`established`,
  `not established`, conflicts, `whyNot`) inside a `200`/`OK`
  answer are content. Alert on the first; route the second to the
  case owner.

## Support boundaries

- "Monitored" in this article means only: the named endpoint,
  fields, spans, and hashes exist and behave as described.
  Dashboards, retention, alert routes, and SLO numbers are the
  operator's stack.
- The health endpoint reports process state, not answer
  correctness — only the canary (operator-run) checks that the
  engine still computes the expected bytes.
- `verificationPending`/`verified` are constant `0` and
  `engineOwnsIr`/`oracleGuard`/`driftScope` are fixed strings in
  this implementation (real, `http.rs`): treat them as
  compatibility fields, not as live verification telemetry.

## Next step

- Article 11 ([Resource limits and cancellation](/operate/resource-limits-cancellation/)) for what `overrunCalls` and `CALL_TIMEOUT` imply for
  capacity; article 09 ([Logs, audit trail, and decision journals](/operate/logs-audit-decision-journals/)) for the journal fields behind every code
  named here.