Health and observability
For LLMs10 sections
Tell from the outside whether the server is alive, ready, and answering correctly — using the real health endpoint, the real logs and spans, and an operator-run canary for answer correctness.
Component law-mcp over --http. Scenario: a supervisor and a
monitoring stack watching a deployed instance. All hostnames and
hashes below are synthetic examples.
Applies to
Section titled “Applies to”| Branch | Coverage in this article |
|---|---|
MCP (law-mcp-server --http) | /healthz, --healthcheck, canary, OTLP |
law serve | step 7: /healthz (liveness) vs /readyz (readiness gate) + endpoint access table in Deploy a private HTTP service; serve branch of Verification and support matrix |
Prerequisites
Section titled “Prerequisites”- HTTP mode on (
--http, real) with a known--port/LAW_MCP_PORT(default8722, real) and profile (--profile/LAW_MCP_PROFILE, defaultall, real). - The
law-mcp-serverbinary available for the--healthcheckprobe, or any HTTP client forGET /healthz(both real).
- Poll liveness with the real endpoint:
GET /healthzanswers200with a JSON object (real,http.rsroute):status:"ok",pid,profile(server profile name),verificationPending:0,verified:0,engineOwnsIr:true,oracleGuard:"absent",driftScope:"semanticHash",overrunCalls(detached workers still running). The probe is never journaled (real): poll every few seconds without log spam. - Run the packaged probe where a binary check fits. Real
(
main.rs):law-mcp-server --healthcheck [--port PORT]connects to127.0.0.1:{port}with a 4 s timeout, sendsGET /healthz, and exits0only onHTTP/1.1 200plus a"status":"ok"body. Anything else exits2withhealthcheck: unexpected /healthz response(or the connect/read error). There is no separate readiness endpoint: serving/healthzwith the profile loaded IS ready (real behavior by construction — the profile loads beforelisten). - Watch saturation, not just aliveness. Real signals:
overrunCallsrising (workers outliving the call timeout), journalMCP_MSpercentiles climbing perMCP_TOOL,CALL_TIMEOUT/INTERNALcodes appearing inMCP_CODE, OTLP drop counter lines on stderr (queue 256 full, real). None of these has a threshold in code — thresholds are operator policy (created-example alert:overrunCalls > 0for 5 minutes). - Collect spans when call timing matters. Real (
otlp.rs): withLAW_MCP_OTLP_URLset, each call emits one trace — root spanmcp.<tool>plus stage spans — over OTLP/HTTP JSON on a background thread (3 s POST timeout, never blocks the call). Trace id = the 32-hex journal call id, so one id joins the journald record, the span, and theX-Law-Call-Idheader; an inbound W3Ctraceparentparents the call under the client’s trace. Spans carry names, outcomes, durations, and reproducibility hashes only — no call bodies (real). Without the URL, no thread and no socket is opened. There is no Prometheus endpoint in code. - Record versions and artifact ids with every release. Real sources:
serverInfo.version(0.1.0,server.rs), the profile name in/healthz, the binary hash from the install record (article 13), and per-answerprogramHash/caseHash/resultHash/codeHashin responses and journal fields. Pin these four hashes in the change record; they are whatlaw_explainreplays against. - Run a canary for correctness (operator-policy-example; no canary
exists in code). Created-example: every 5 minutes ask a pinned
question over a pinned case (e.g. the 14-day BGB period ending
20 March) and compare
resultHashto the recorded value. Alert on any mismatch or on a non-OKoutcome. Keep the canary case immutable and out of the served tenant roots so tenant edits cannot break it. - For
law serve, keep liveness and readiness apart (real,service.rs).GET /healthzis liveness: 200status:"ok"withpid, pin,programHash,workers,warmed, queue and call counters — it answers 200 even while workers warm and even after a failure.GET /readyzis the readiness gate: 200{"status":"ready", pin, programHash, workers}only when every worker is warm AND no failure is recorded; 503{"status":"warming", warmed, workers}while workers warm; 503{"status":"failed", reason}after a warmup failure or once the decision journal breaks (reason: "decision journal broken: …", real).POST /v1/askenforces the same gate and answers 503SERVE_NOT_READYwhile/readyzis not 200 — the balancer opens ingress on/readyz200 only, never on/healthz. None of the three observability routes (/healthz,/readyz,/metrics) takes a token;/v1/worldand/v1/askdo (full table in Deploy a private HTTP service).
Expected result
Section titled “Expected result”- Liveness:
GET /healthz→200+status:"ok"; the packaged probe exits0. - Readiness: same check, because the listener starts only after the
profile loads; a process that answers
/healthzserves/mcp. law servereadiness is separate:/healthz200 proves the process lives; only/readyz200 proves it serves. Ingress, supervisors, and the rollback/upgrade gates all key on/readyz, and afailedreason names the cause (warmup or broken journal) without further probing.- Correctness: the canary’s
resultHashequals the pinned value on every run; drift means the release, profile, or calendar input changed.
Result check
Section titled “Result check”curl -s http://127.0.0.1:8722/healthz(created-example host/port matching the real defaults) prints the real field set above withoverrunCalls:0on an idle server.law-mcp-server --healthcheck --port 8722exits0while the server runs, non-zero within ~4 s when it does not (real timeout).- Stop the OTLP receiver: calls still succeed with unchanged answer bytes; only the stderr drop counter moves (real non-blocking export).
Failures and diagnostics
Section titled “Failures and diagnostics”GET /healthzrefused or slow while/mcpworks: compare direct backend access against access through the edge, and check backend load (overrunCalls, worker saturation). A shared accept loop does not guarantee equal per-request latency — suspect the edge or the prober only after the direct-vs-edge comparison clears the backend.overrunCallsstuck above zero: detached workers still computing (normal briefly) or wedged (restart the stateless process to reclaim; see article 11).- Canary hash mismatch with
OKoutcomes: inputs moved — compareprogramHash(canon changed),caseHash(case changed),codeHash(engine changed) to isolate which (real hash roles). - Infra errors vs content states (real split): HTTP statuses
(
401/403/404/405/411/413), JSON-RPC errors (-32001timeout,-32603internal), andSERVICE_*codes are transport/infra; evaluation statuses (established,not established, conflicts,whyNot) inside a200/OKanswer are content. Alert on the first; route the second to the case owner.
Support boundaries
Section titled “Support boundaries”- “Monitored” in this article means only: the named endpoint, fields, spans, and hashes exist and behave as described. Dashboards, retention, alert routes, and SLO numbers are the operator’s stack.
- The health endpoint reports process state, not answer correctness — only the canary (operator-run) checks that the engine still computes the expected bytes.
verificationPending/verifiedare constant0andengineOwnsIr/oracleGuard/driftScopeare fixed strings in this implementation (real,http.rs): treat them as compatibility fields, not as live verification telemetry.
Next step
Section titled “Next step”- Article 11 (Resource limits and cancellation) for what
overrunCallsandCALL_TIMEOUTimply for capacity; article 09 (Logs, audit trail, and decision journals) for the journal fields behind every code named here.
Documentation for Arxo. Writings — blog.arxo.io.
Anonymous visit counts on stats.arxo.io, no cookies.