Markdown for LLMs
Resource limits and cancellation
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Resource limits and cancellation ## Goal Name every size, count, and time limit the serving path enforces, what fires when each trips, and what "cancel" really stops — so an overload degrades into refusals, not into stuck processes. ## Scope Component `law-mcp` over `--http` (stdio has no transport limits), plus the per-tool caps shared by both transports. Scenario: hostile or pathological load against a deployed server. All numbers below are real code constants unless marked otherwise. ## Applies to | Branch | Coverage in this article | |---|---| | MCP (`law-mcp-server --http`) | full limit table, timeout-vs-abandonment semantics | | `law serve` | caps only (1 MiB body, 16 MiB response, queue 64, 10 000 assertions) in [Deploy a private HTTP service](/operate/private-http-service/) and [Configuration reference](/operate/configuration-reference/); serve has no graceful drain | ## Prerequisites - Know the three limit layers: edge/reverse proxy (operator's), HTTP transport (`http.rs`, `main.rs`), engine/tool caps (`law-eval`, `reference.rs`, `execution.rs`, `ask.rs`). - Default posture: no call timeout (`LAW_MCP_CALL_TIMEOUT` unset means none, real), 1 MiB request cap always on (real), per-tool caps always on (real). ## Steps 1. Size the transport. Real transport limits (`http.rs`, `main.rs`): `MAX_BODY` = 1 MiB request body → `411` when empty-length, `413` when over; `LAW_MCP_CALL_TIMEOUT` seconds (fractions allowed, `0`/empty disables) → JSON-RPC `-32001` with `data.code=CALL_TIMEOUT` and `limitSeconds` (real). Deployment reference-setting from the code comment (real comment, operator choice): `110` s under an edge `proxy_read_timeout` of `120` s, so a pathological call fails with a response, not a broken connection. 2. Understand timeout semantics. Real mechanism (`http.rs` `dispatch`): with a limit set, the computation runs on a worker thread (`law-mcp-call`); `law-eval` has no cancellation points, so on expiry the client gets the technical refusal while the worker keeps running detached. Its result is discarded; when it finishes, a journal record with `MCP_CODE=OVERRUN_FINISHED` lands; `/healthz` field `overrunCalls` counts still-running detached workers (real). Timeout bounds the client's wait, not the CPU spent. 3. Count the tool caps. Real per-tool limits: argumentation `maxArguments` 64 / `maxExtensions` 8 / `maxUndecided` 16 defaults, raisable per call via `limits` (`law-eval/src/ argumentation.rs`); structural query `rowLimit` default 20, at most 5000, else the call is refused (`reference.rs`); inspect node page 24 000 bytes, over-budget nodes answer `RESPONSE_TOO_LARGE` with `budgetBytes` (`execution.rs` `PAGE_BYTES`); neural appendix accepts at most 8 candidates (`reference/search.rs`). There is no global CPU/RAM quota and no per-tool wall clock inside the engine. 4. Bound the outbound legs. Real (`service_client.rs`, `publication.rs`, `otlp.rs`): service POSTs time out at 30 s with zero redirects; default response cap 4 MiB, Lens answers cap 32 MiB; OTLP spans queue 256 then drop (counter to stderr), one OTLP POST times out at 3 s on a background thread that never blocks a call. 5. Note what is NOT limited. Real implementation limits by absence: no cap on concurrent connections (one thread per connection, `http.rs` accept loop), no request queue with backpressure, no response-size cap on the MCP reply path itself, no memory or CPU cgroup. Concurrency control lives at the edge (operator-policy-example: proxy `limit_conn`, `client_max_body_size 1m` to match `MAX_BODY`, `proxy_read_timeout 120s`). 6. Plan overload behavior. Shed order under pressure (reference behavior assembled from the mechanisms above): over-size bodies → `413` before any work; unknown routes/methods → `404`/`405`; timed-out calls → `-32001` to the client while workers drain; OTLP drops spans rather than slowing calls; the journal never blocks a call (datagram → trimmed resend → stderr). What the server cannot do is reclaim a runaway worker early — size the host so the worst case fits, and restart the process to reclaim (stateless: no sessions, `DELETE` is refused with `405`, real). ## Expected result - Any single request is answered, refused with a code, or timed out with `-32001`; no request hangs past `LAW_MCP_CALL_TIMEOUT` plus edge timeouts from the client's viewpoint. - `overrunCalls` returns to `0` after detached workers finish. It growing without bound is *compatible with* a too-short timeout but also with overload and heavy or wedged computations — so cap concurrency and characterize the queries first (see [Capacity and scaling](/operate/capacity-scaling/)), and only then consider raising the limit while watching memory. - No call path allocates without bound on untrusted size fields: every list the client sizes (`rowLimit`, candidates, node pages) has a hard ceiling (real). ## Result check - `POST /mcp` with a 2 MiB body → `413` (real). - `LAW_MCP_CALL_TIMEOUT=0.001` (created-example) on any real call → `-32001`/`CALL_TIMEOUT`, then an `OVERRUN_FINISHED` journal line and `overrunCalls` back to `0`. - `law_query` with `rowLimit:5001` → refusal naming the row limit (real); `rowLimit:5000` → served page with `next` when truncated. ## Failures and diagnostics - `CALL_TIMEOUT` on healthy calls: limit below normal latency (check `MCP_MS` in the journal), or CPU starvation from too many concurrent workers — add the edge concurrency cap. - `Cannot spawn worker` path: thread creation failure surfaces as `Overrun` (real, `run_with_deadline`): the host is out of threads or memory — an infra error, not a content state. - Handler panic → `-32603` with `MCP_CODE=INTERNAL` (real); the connection thread logs and the process keeps serving. ## Support boundaries - "Supported" load in this article means only: refusal codes and counters behave as named. Throughput numbers, sizing tables, and autoscaling are the operator's measurements — the repo ships no load profile. - Cancellation is abandonment, not preemption: CPU/RAM spent by a detached worker is spent. True termination is process restart, which frees the process's resources; "no sessions" does not settle the fate of an in-flight case/draft mutation — retry terms for effecting operations come from each operation's own contract ([effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix)), not from the restart. - Fairness between concurrent calls is the OS scheduler's; the server has no priority queue. ## Next step - Article 12 ([Health and observability](/operate/health-observability/)) for watching `overrunCalls` and call latency from the outside; article 09 ([Logs, audit trail, and decision journals](/operate/logs-audit-decision-journals/)) for the journal fields each outcome emits.