docs← Back to article

Markdown for LLMs

Capacity planning and scaling

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Capacity planning and scaling

## Goal

Size a `law-mcp-server --http` deployment from measurements of the
served endpoint under a realistic question mix — and know exactly when
to add instances, what must stay identical across them, and what must
stay per-instance.

## Scope

Component `law-mcp-server` (crate `law-mcp`), HTTP mode behind an
operator edge, as set up in
[Deploy a private MCP server](/operate/private-mcp-server/). Server
identity `serverInfo.version "0.1.0"` (real); image `linux/amd64`,
builder `rust:1.94.1-alpine` (real build facts). Every endpoint,
flag, and variable below is real; every load number, threshold,
and replica count is a synthetic created-example marked as such.

## Applies to

| Branch | Coverage in this article |
|---|---|
| MCP (`law-mcp-server --http`) | full sizing recipe: measure, order timeouts, scale on identical slices |
| `law serve` | no dedicated sizing recipe in this section yet — start from the serve caps and workers in [Deploy a private HTTP service](/operate/private-http-service/) and measure the served endpoint the same way |

## Prerequisites

- A running instance with `GET /healthz` returning `"status":"ok"`.
- The section's journals article ([Logs, audit and decision journals](/operate/logs-audit-decision-journals/))
  and limits article ([Resource limits and cancellation](/operate/resource-limits-cancellation/)) read: this
  article measures, those two explain what the numbers mean.
- A question mix representative of production (tool names real, e.g.
  `law_search`, `law_ask`; arguments synthetic).

## Steps

1. Baseline one call through the full path (real). Time a
   `POST /mcp` `tools/call` through the edge, not an in-process
   evaluation: transport framing, the journal record (written
   *before* the answer is sent, `src/http.rs` `record()`), and the
   `X-Law-Call-Id` round trip are all on the served path and invisible
   to a library-call timer.
2. Warm the lazy indexes (real mechanism). The tool catalog and the
   task-guide index load once per process on first access, not at
   handshake (`OnceLock`, `src/server.rs`); the first search-class
   call pays a tree walk. Send one call per tool family you serve
   before measuring, and allow container cold start past the
   `HEALTHCHECK --start-period=20s` window (reference-setting,
   `Dockerfile`).
3. Apply load in the shape of production (operator-policy-example
   shape, created-example numbers). Drive concurrent `POST /mcp`
   calls — e.g. 8 at once for 5 minutes, mixed `law_search` and
   `law_ask` — and record served latency (p50/p95 at the edge),
   HTTP statuses, `/healthz` `overrunCalls`, and journal completeness.
   There is no in-code connection cap (thread per accepted connection,
   `src/http.rs` `serve()`), so the ceiling you find is the machine's,
   not a configured pool's: implement no retry storm against it.
4. Size memory from the run (real mechanism, created-example numbers).
   The server sets no heap or cgroup limit itself
   (implementation-limit: no memory governor in the crate). Measure
   steady RSS plus growth per concurrent call, noting that a call
   abandoned at `LAW_MCP_CALL_TIMEOUT` keeps its worker thread until
   the computation finishes — there are no cancellation points in
   `law-eval` (`src/http.rs` `dispatch()`), so overrun threads hold
   memory after the client already got `-32001`. Set the container
   memory limit from that measurement plus headroom
   (operator-policy-example).
5. Order the two timeouts before scaling out (reference-setting). Keep
   `LAW_MCP_CALL_TIMEOUT` below the edge `proxy_read_timeout`: the
   code comment names 110 s against nginx 120 s (`src/main.rs`) so an
   overrun fails with a JSON-RPC response, not a broken connection.
   The default is no limit (real default); any other pair is operator
   policy, but the ordering is load-bearing.
6. Scale out only on identical slices (real conditions). Stateless
   transport is necessary but not sufficient for "any instance can
   serve any call" — affinity is decided by the
   [effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix):
   pure evaluation plus the two remote legs are affinity-free; case
   writes need a shared tree (or sticky routing); `law://answer/`
   reads need a shared artifact dir (or the owning instance).
   Identical instances additionally require: same image digest (the
   `mcp-image` workflow prints one `digest` per build for pinning);
   same profile and tree
   (`image_tree.py <profile> --check-pin <dir>` exit `0` for each);
   same `LAW_MCP_CALL_TIMEOUT`; same answers-service wiring
   (`LAW_ANSWERS_URL`/`LAW_ANSWERS_TOKEN`) or all unwired.
7. Keep per-instance what identifies the instance (real + operator
   policy). Journals: one journald unit/namespace per instance
   (operator-policy-example; the record carries no instance field of
   its own, so the unit name is the attribution). Answer artifacts:
   `LAW_MCP_ARTIFACT_DIR`, else `$HOME/.cache/law-dsl/mcp-answers/<profile>`
   (real, `src/execution.rs`) — an `explain` against a saved
   `answerResource` reads the artifact from this directory, so the
   call must reach an instance that sees the file: route to the
   owning instance, or share one directory across instances
   (operator policy; whoever stages artifacts there needs write;
   the server itself only reads; a shared artifact directory is
   outside the checked matrix in
   [Verification and support matrix](/operate/verification-support-matrix/)).

## Expected result

- Served p95 within the operator's budget at the planned concurrency,
  with `overrunCalls` returning to `0` after the drain.
- No `-32001` (`CALL_TIMEOUT`), no `413` (request over the
  `MAX_BODY` 1 MiB implementation-limit), no `MCP_TRUNCATED=oversize`
  journal records.
- Every instance answers `/healthz` with its own `pid`, the same
  `profile`, and `"oracleGuard":"absent"`.

## Result check

- `GET /healthz` during and after load: `200`,
  `"status":"ok"`, `overrunCalls` back to `0`.
- Journal completeness: every served call id carries its primary
  record (`/healthz` is never journaled — real); correlate by
  `MCP_CALL_ID`, not by counting lines — a late-finishing call
  adds an `OVERRUN_FINISHED` record for the same id, and a
  journal-write failure can drop a record while the call succeeds
  (see [Data flow](/operate/data-flow-storage-retention/) and
  [Logs, audit, decision journals](/operate/logs-audit-decision-journals/)). No
  `MCP_TRUNCATED` fields at production body sizes.
- Memory: RSS flat across repeated identical runs; growth that
  tracks `overrunCalls` is abandoned workers, not a leak in the
  accounting — confirm by draining load and re-reading `/healthz`.

## Failures and diagnostics

- `overrunCalls` climbs and never drains: calls exceed the call
  timeout while workers pile up — lower concurrency at the edge,
  raise the timeout only if the questions legitimately need it, and
  re-measure memory (each overrun thread still counts).
- Edge `504` while the server journal shows `200`: timeout inversion
  — the edge gave up before the server answered; restore the
  call-limit-below-proxy ordering (step 5).
- First-call latency spike after every restart: cold lazy indexes —
  extend warmup (step 2), do not size steady-state capacity from
  cold numbers.
- `413` under load: some client sends bodies over 1 MiB
  (implementation-limit `MAX_BODY`, `src/http.rs`) — fix the client
  or split the call; the cap is not tunable by flag or variable.

## Support boundaries

- Supported: single-instance sizing and identical-slice scale-out per
  the conditions in step 6; the checks in `Result check`.
- Implementation limits: no in-code connection cap, memory cap, or
  per-user quota; one `LAW_MCP_TOKEN` per process, so tenant
  separation means separate instances/tokens/edge routes (operator
  policy) — the client label in the journal is the edge-supplied
  `X-Forwarded-For` string, trusted by layout only (`src/http.rs`
  `context()`).
- "Isolated" in this article means only: separate processes with
  separate tokens and journals. It does not mean the server
  authenticates end users — it does not.
- Out of scope: shared artifact directories, multi-profile routing
  inside one process, and any autoscaler wiring (all operator policy,
  none checked by this section).

## Next step

Record the measured triple (concurrency, p95, RSS) with the image
digest and profile in the decision journal, then run the matrix in
[Verification and support matrix](/operate/verification-support-matrix/) against the scaled deployment; on
any red cell, triage with [Troubleshooting](/operate/troubleshooting/).