Capacity planning and scaling
Size a law-mcp-server --http deployment from measurements of the
served endpoint under a realistic question mix — and know exactly when
to add instances, what must stay identical across them, and what must
stay per-instance.
Component law-mcp-server (crate law-mcp), HTTP mode behind an
operator edge, as set up in
Deploy a private MCP server. Server
identity serverInfo.version "0.1.0" (real); image linux/amd64,
builder rust:1.94.1-alpine (real build facts). Every endpoint,
flag, and variable below is real; every load number, threshold,
and replica count is a synthetic created-example marked as such.
Applies to
Section titled “Applies to”| Branch | Coverage in this article |
|---|---|
MCP (law-mcp-server --http) | full sizing recipe: measure, order timeouts, scale on identical slices |
law serve | no dedicated sizing recipe in this section yet — start from the serve caps and workers in Deploy a private HTTP service and measure the served endpoint the same way |
Prerequisites
Section titled “Prerequisites”- A running instance with
GET /healthzreturning"status":"ok". - The section’s journals article (Logs, audit and decision journals) and limits article (Resource limits and cancellation) read: this article measures, those two explain what the numbers mean.
- A question mix representative of production (tool names real, e.g.
law_search,law_ask; arguments synthetic).
- Baseline one call through the full path (real). Time a
POST /mcptools/callthrough the edge, not an in-process evaluation: transport framing, the journal record (written before the answer is sent,src/http.rsrecord()), and theX-Law-Call-Idround trip are all on the served path and invisible to a library-call timer. - Warm the lazy indexes (real mechanism). The tool catalog and the
task-guide index load once per process on first access, not at
handshake (
OnceLock,src/server.rs); the first search-class call pays a tree walk. Send one call per tool family you serve before measuring, and allow container cold start past theHEALTHCHECK --start-period=20swindow (reference-setting,Dockerfile). - Apply load in the shape of production (operator-policy-example
shape, created-example numbers). Drive concurrent
POST /mcpcalls — e.g. 8 at once for 5 minutes, mixedlaw_searchandlaw_ask— and record served latency (p50/p95 at the edge), HTTP statuses,/healthzoverrunCalls, and journal completeness. There is no in-code connection cap (thread per accepted connection,src/http.rsserve()), so the ceiling you find is the machine’s, not a configured pool’s: implement no retry storm against it. - Size memory from the run (real mechanism, created-example numbers).
The server sets no heap or cgroup limit itself
(implementation-limit: no memory governor in the crate). Measure
steady RSS plus growth per concurrent call, noting that a call
abandoned at
LAW_MCP_CALL_TIMEOUTkeeps its worker thread until the computation finishes — there are no cancellation points inlaw-eval(src/http.rsdispatch()), so overrun threads hold memory after the client already got-32001. Set the container memory limit from that measurement plus headroom (operator-policy-example). - Order the two timeouts before scaling out (reference-setting). Keep
LAW_MCP_CALL_TIMEOUTbelow the edgeproxy_read_timeout: the code comment names 110 s against nginx 120 s (src/main.rs) so an overrun fails with a JSON-RPC response, not a broken connection. The default is no limit (real default); any other pair is operator policy, but the ordering is load-bearing. - Scale out only on identical slices (real conditions). Stateless
transport is necessary but not sufficient for “any instance can
serve any call” — affinity is decided by the
effects matrix:
pure evaluation plus the two remote legs are affinity-free; case
writes need a shared tree (or sticky routing);
law://answer/reads need a shared artifact dir (or the owning instance). Identical instances additionally require: same image digest (themcp-imageworkflow prints onedigestper build for pinning); same profile and tree (image_tree.py <profile> --check-pin <dir>exit0for each); sameLAW_MCP_CALL_TIMEOUT; same answers-service wiring (LAW_ANSWERS_URL/LAW_ANSWERS_TOKEN) or all unwired. - Keep per-instance what identifies the instance (real + operator
policy). Journals: one journald unit/namespace per instance
(operator-policy-example; the record carries no instance field of
its own, so the unit name is the attribution). Answer artifacts:
LAW_MCP_ARTIFACT_DIR, else$HOME/.cache/law-dsl/mcp-answers/<profile>(real,src/execution.rs) — anexplainagainst a savedanswerResourcereads the artifact from this directory, so the call must reach an instance that sees the file: route to the owning instance, or share one directory across instances (operator policy; whoever stages artifacts there needs write; the server itself only reads; a shared artifact directory is outside the checked matrix in Verification and support matrix).
Expected result
Section titled “Expected result”- Served p95 within the operator’s budget at the planned concurrency,
with
overrunCallsreturning to0after the drain. - No
-32001(CALL_TIMEOUT), no413(request over theMAX_BODY1 MiB implementation-limit), noMCP_TRUNCATED=oversizejournal records. - Every instance answers
/healthzwith its ownpid, the sameprofile, and"oracleGuard":"absent".
Result check
Section titled “Result check”GET /healthzduring and after load:200,"status":"ok",overrunCallsback to0.- Journal completeness: every served call id carries its primary
record (
/healthzis never journaled — real); correlate byMCP_CALL_ID, not by counting lines — a late-finishing call adds anOVERRUN_FINISHEDrecord for the same id, and a journal-write failure can drop a record while the call succeeds (see Data flow and Logs, audit, decision journals). NoMCP_TRUNCATEDfields at production body sizes. - Memory: RSS flat across repeated identical runs; growth that
tracks
overrunCallsis abandoned workers, not a leak in the accounting — confirm by draining load and re-reading/healthz.
Failures and diagnostics
Section titled “Failures and diagnostics”overrunCallsclimbs and never drains: calls exceed the call timeout while workers pile up — lower concurrency at the edge, raise the timeout only if the questions legitimately need it, and re-measure memory (each overrun thread still counts).- Edge
504while the server journal shows200: timeout inversion — the edge gave up before the server answered; restore the call-limit-below-proxy ordering (step 5). - First-call latency spike after every restart: cold lazy indexes — extend warmup (step 2), do not size steady-state capacity from cold numbers.
413under load: some client sends bodies over 1 MiB (implementation-limitMAX_BODY,src/http.rs) — fix the client or split the call; the cap is not tunable by flag or variable.
Support boundaries
Section titled “Support boundaries”- Supported: single-instance sizing and identical-slice scale-out per
the conditions in step 6; the checks in
Result check. - Implementation limits: no in-code connection cap, memory cap, or
per-user quota; one
LAW_MCP_TOKENper process, so tenant separation means separate instances/tokens/edge routes (operator policy) — the client label in the journal is the edge-suppliedX-Forwarded-Forstring, trusted by layout only (src/http.rscontext()). - “Isolated” in this article means only: separate processes with separate tokens and journals. It does not mean the server authenticates end users — it does not.
- Out of scope: shared artifact directories, multi-profile routing inside one process, and any autoscaler wiring (all operator policy, none checked by this section).
Next step
Section titled “Next step”Record the measured triple (concurrency, p95, RSS) with the image digest and profile in the decision journal, then run the matrix in Verification and support matrix against the scaled deployment; on any red cell, triage with Troubleshooting.
Documentation for Arxo. Writings — blog.arxo.io.
Anonymous visit counts on stats.arxo.io, no cookies.