# Capacity planning and scaling ## Goal Size a `law-mcp-server --http` deployment from measurements of the served endpoint under a realistic question mix — and know exactly when to add instances, what must stay identical across them, and what must stay per-instance. ## Scope Component `law-mcp-server` (crate `law-mcp`), HTTP mode behind an operator edge, as set up in [Deploy a private MCP server](/operate/private-mcp-server/). Server identity `serverInfo.version "0.1.0"` (real); image `linux/amd64`, builder `rust:1.94.1-alpine` (real build facts). Every endpoint, flag, and variable below is real; every load number, threshold, and replica count is a synthetic created-example marked as such. ## Applies to | Branch | Coverage in this article | |---|---| | MCP (`law-mcp-server --http`) | full sizing recipe: measure, order timeouts, scale on identical slices | | `law serve` | no dedicated sizing recipe in this section yet — start from the serve caps and workers in [Deploy a private HTTP service](/operate/private-http-service/) and measure the served endpoint the same way | ## Prerequisites - A running instance with `GET /healthz` returning `"status":"ok"`. - The section's journals article ([Logs, audit and decision journals](/operate/logs-audit-decision-journals/)) and limits article ([Resource limits and cancellation](/operate/resource-limits-cancellation/)) read: this article measures, those two explain what the numbers mean. - A question mix representative of production (tool names real, e.g. `law_search`, `law_ask`; arguments synthetic). ## Steps 1. Baseline one call through the full path (real). Time a `POST /mcp` `tools/call` through the edge, not an in-process evaluation: transport framing, the journal record (written *before* the answer is sent, `src/http.rs` `record()`), and the `X-Law-Call-Id` round trip are all on the served path and invisible to a library-call timer. 2. Warm the lazy indexes (real mechanism). The tool catalog and the task-guide index load once per process on first access, not at handshake (`OnceLock`, `src/server.rs`); the first search-class call pays a tree walk. Send one call per tool family you serve before measuring, and allow container cold start past the `HEALTHCHECK --start-period=20s` window (reference-setting, `Dockerfile`). 3. Apply load in the shape of production (operator-policy-example shape, created-example numbers). Drive concurrent `POST /mcp` calls — e.g. 8 at once for 5 minutes, mixed `law_search` and `law_ask` — and record served latency (p50/p95 at the edge), HTTP statuses, `/healthz` `overrunCalls`, and journal completeness. There is no in-code connection cap (thread per accepted connection, `src/http.rs` `serve()`), so the ceiling you find is the machine's, not a configured pool's: implement no retry storm against it. 4. Size memory from the run (real mechanism, created-example numbers). The server sets no heap or cgroup limit itself (implementation-limit: no memory governor in the crate). Measure steady RSS plus growth per concurrent call, noting that a call abandoned at `LAW_MCP_CALL_TIMEOUT` keeps its worker thread until the computation finishes — there are no cancellation points in `law-eval` (`src/http.rs` `dispatch()`), so overrun threads hold memory after the client already got `-32001`. Set the container memory limit from that measurement plus headroom (operator-policy-example). 5. Order the two timeouts before scaling out (reference-setting). Keep `LAW_MCP_CALL_TIMEOUT` below the edge `proxy_read_timeout`: the code comment names 110 s against nginx 120 s (`src/main.rs`) so an overrun fails with a JSON-RPC response, not a broken connection. The default is no limit (real default); any other pair is operator policy, but the ordering is load-bearing. 6. Scale out only on identical slices (real conditions). Stateless transport is necessary but not sufficient for "any instance can serve any call" — affinity is decided by the [effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix): pure evaluation plus the two remote legs are affinity-free; case writes need a shared tree (or sticky routing); `law://answer/` reads need a shared artifact dir (or the owning instance). Identical instances additionally require: same image digest (the `mcp-image` workflow prints one `digest` per build for pinning); same profile and tree (`image_tree.py --check-pin ` exit `0` for each); same `LAW_MCP_CALL_TIMEOUT`; same answers-service wiring (`LAW_ANSWERS_URL`/`LAW_ANSWERS_TOKEN`) or all unwired. 7. Keep per-instance what identifies the instance (real + operator policy). Journals: one journald unit/namespace per instance (operator-policy-example; the record carries no instance field of its own, so the unit name is the attribution). Answer artifacts: `LAW_MCP_ARTIFACT_DIR`, else `$HOME/.cache/law-dsl/mcp-answers/` (real, `src/execution.rs`) — an `explain` against a saved `answerResource` reads the artifact from this directory, so the call must reach an instance that sees the file: route to the owning instance, or share one directory across instances (operator policy; whoever stages artifacts there needs write; the server itself only reads; a shared artifact directory is outside the checked matrix in [Verification and support matrix](/operate/verification-support-matrix/)). ## Expected result - Served p95 within the operator's budget at the planned concurrency, with `overrunCalls` returning to `0` after the drain. - No `-32001` (`CALL_TIMEOUT`), no `413` (request over the `MAX_BODY` 1 MiB implementation-limit), no `MCP_TRUNCATED=oversize` journal records. - Every instance answers `/healthz` with its own `pid`, the same `profile`, and `"oracleGuard":"absent"`. ## Result check - `GET /healthz` during and after load: `200`, `"status":"ok"`, `overrunCalls` back to `0`. - Journal completeness: every served call id carries its primary record (`/healthz` is never journaled — real); correlate by `MCP_CALL_ID`, not by counting lines — a late-finishing call adds an `OVERRUN_FINISHED` record for the same id, and a journal-write failure can drop a record while the call succeeds (see [Data flow](/operate/data-flow-storage-retention/) and [Logs, audit, decision journals](/operate/logs-audit-decision-journals/)). No `MCP_TRUNCATED` fields at production body sizes. - Memory: RSS flat across repeated identical runs; growth that tracks `overrunCalls` is abandoned workers, not a leak in the accounting — confirm by draining load and re-reading `/healthz`. ## Failures and diagnostics - `overrunCalls` climbs and never drains: calls exceed the call timeout while workers pile up — lower concurrency at the edge, raise the timeout only if the questions legitimately need it, and re-measure memory (each overrun thread still counts). - Edge `504` while the server journal shows `200`: timeout inversion — the edge gave up before the server answered; restore the call-limit-below-proxy ordering (step 5). - First-call latency spike after every restart: cold lazy indexes — extend warmup (step 2), do not size steady-state capacity from cold numbers. - `413` under load: some client sends bodies over 1 MiB (implementation-limit `MAX_BODY`, `src/http.rs`) — fix the client or split the call; the cap is not tunable by flag or variable. ## Support boundaries - Supported: single-instance sizing and identical-slice scale-out per the conditions in step 6; the checks in `Result check`. - Implementation limits: no in-code connection cap, memory cap, or per-user quota; one `LAW_MCP_TOKEN` per process, so tenant separation means separate instances/tokens/edge routes (operator policy) — the client label in the journal is the edge-supplied `X-Forwarded-For` string, trusted by layout only (`src/http.rs` `context()`). - "Isolated" in this article means only: separate processes with separate tokens and journals. It does not mean the server authenticates end users — it does not. - Out of scope: shared artifact directories, multi-profile routing inside one process, and any autoscaler wiring (all operator policy, none checked by this section). ## Next step Record the measured triple (concurrency, p95, RSS) with the image digest and profile in the decision journal, then run the matrix in [Verification and support matrix](/operate/verification-support-matrix/) against the scaled deployment; on any red cell, triage with [Troubleshooting](/operate/troubleshooting/).