Skip to content
docs
Arxo ↗

Capacity planning and scaling

For LLMs10 sections

Size a law-mcp-server --http deployment from measurements of the served endpoint under a realistic question mix — and know exactly when to add instances, what must stay identical across them, and what must stay per-instance.

Component law-mcp-server (crate law-mcp), HTTP mode behind an operator edge, as set up in Deploy a private MCP server. Server identity serverInfo.version "0.1.0" (real); image linux/amd64, builder rust:1.94.1-alpine (real build facts). Every endpoint, flag, and variable below is real; every load number, threshold, and replica count is a synthetic created-example marked as such.

BranchCoverage in this article
MCP (law-mcp-server --http)full sizing recipe: measure, order timeouts, scale on identical slices
law serveno dedicated sizing recipe in this section yet — start from the serve caps and workers in Deploy a private HTTP service and measure the served endpoint the same way
  • A running instance with GET /healthz returning "status":"ok".
  • The section’s journals article (Logs, audit and decision journals) and limits article (Resource limits and cancellation) read: this article measures, those two explain what the numbers mean.
  • A question mix representative of production (tool names real, e.g. law_search, law_ask; arguments synthetic).
  1. Baseline one call through the full path (real). Time a POST /mcp tools/call through the edge, not an in-process evaluation: transport framing, the journal record (written before the answer is sent, src/http.rs record()), and the X-Law-Call-Id round trip are all on the served path and invisible to a library-call timer.
  2. Warm the lazy indexes (real mechanism). The tool catalog and the task-guide index load once per process on first access, not at handshake (OnceLock, src/server.rs); the first search-class call pays a tree walk. Send one call per tool family you serve before measuring, and allow container cold start past the HEALTHCHECK --start-period=20s window (reference-setting, Dockerfile).
  3. Apply load in the shape of production (operator-policy-example shape, created-example numbers). Drive concurrent POST /mcp calls — e.g. 8 at once for 5 minutes, mixed law_search and law_ask — and record served latency (p50/p95 at the edge), HTTP statuses, /healthz overrunCalls, and journal completeness. There is no in-code connection cap (thread per accepted connection, src/http.rs serve()), so the ceiling you find is the machine’s, not a configured pool’s: implement no retry storm against it.
  4. Size memory from the run (real mechanism, created-example numbers). The server sets no heap or cgroup limit itself (implementation-limit: no memory governor in the crate). Measure steady RSS plus growth per concurrent call, noting that a call abandoned at LAW_MCP_CALL_TIMEOUT keeps its worker thread until the computation finishes — there are no cancellation points in law-eval (src/http.rs dispatch()), so overrun threads hold memory after the client already got -32001. Set the container memory limit from that measurement plus headroom (operator-policy-example).
  5. Order the two timeouts before scaling out (reference-setting). Keep LAW_MCP_CALL_TIMEOUT below the edge proxy_read_timeout: the code comment names 110 s against nginx 120 s (src/main.rs) so an overrun fails with a JSON-RPC response, not a broken connection. The default is no limit (real default); any other pair is operator policy, but the ordering is load-bearing.
  6. Scale out only on identical slices (real conditions). Stateless transport is necessary but not sufficient for “any instance can serve any call” — affinity is decided by the effects matrix: pure evaluation plus the two remote legs are affinity-free; case writes need a shared tree (or sticky routing); law://answer/ reads need a shared artifact dir (or the owning instance). Identical instances additionally require: same image digest (the mcp-image workflow prints one digest per build for pinning); same profile and tree (image_tree.py <profile> --check-pin <dir> exit 0 for each); same LAW_MCP_CALL_TIMEOUT; same answers-service wiring (LAW_ANSWERS_URL/LAW_ANSWERS_TOKEN) or all unwired.
  7. Keep per-instance what identifies the instance (real + operator policy). Journals: one journald unit/namespace per instance (operator-policy-example; the record carries no instance field of its own, so the unit name is the attribution). Answer artifacts: LAW_MCP_ARTIFACT_DIR, else $HOME/.cache/law-dsl/mcp-answers/<profile> (real, src/execution.rs) — an explain against a saved answerResource reads the artifact from this directory, so the call must reach an instance that sees the file: route to the owning instance, or share one directory across instances (operator policy; whoever stages artifacts there needs write; the server itself only reads; a shared artifact directory is outside the checked matrix in Verification and support matrix).
  • Served p95 within the operator’s budget at the planned concurrency, with overrunCalls returning to 0 after the drain.
  • No -32001 (CALL_TIMEOUT), no 413 (request over the MAX_BODY 1 MiB implementation-limit), no MCP_TRUNCATED=oversize journal records.
  • Every instance answers /healthz with its own pid, the same profile, and "oracleGuard":"absent".
  • GET /healthz during and after load: 200, "status":"ok", overrunCalls back to 0.
  • Journal completeness: every served call id carries its primary record (/healthz is never journaled — real); correlate by MCP_CALL_ID, not by counting lines — a late-finishing call adds an OVERRUN_FINISHED record for the same id, and a journal-write failure can drop a record while the call succeeds (see Data flow and Logs, audit, decision journals). No MCP_TRUNCATED fields at production body sizes.
  • Memory: RSS flat across repeated identical runs; growth that tracks overrunCalls is abandoned workers, not a leak in the accounting — confirm by draining load and re-reading /healthz.
  • overrunCalls climbs and never drains: calls exceed the call timeout while workers pile up — lower concurrency at the edge, raise the timeout only if the questions legitimately need it, and re-measure memory (each overrun thread still counts).
  • Edge 504 while the server journal shows 200: timeout inversion — the edge gave up before the server answered; restore the call-limit-below-proxy ordering (step 5).
  • First-call latency spike after every restart: cold lazy indexes — extend warmup (step 2), do not size steady-state capacity from cold numbers.
  • 413 under load: some client sends bodies over 1 MiB (implementation-limit MAX_BODY, src/http.rs) — fix the client or split the call; the cap is not tunable by flag or variable.
  • Supported: single-instance sizing and identical-slice scale-out per the conditions in step 6; the checks in Result check.
  • Implementation limits: no in-code connection cap, memory cap, or per-user quota; one LAW_MCP_TOKEN per process, so tenant separation means separate instances/tokens/edge routes (operator policy) — the client label in the journal is the edge-supplied X-Forwarded-For string, trusted by layout only (src/http.rs context()).
  • “Isolated” in this article means only: separate processes with separate tokens and journals. It does not mean the server authenticates end users — it does not.
  • Out of scope: shared artifact directories, multi-profile routing inside one process, and any autoscaler wiring (all operator policy, none checked by this section).

Record the measured triple (concurrency, p95, RSS) with the image digest and profile in the decision journal, then run the matrix in Verification and support matrix against the scaled deployment; on any red cell, triage with Troubleshooting.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.