Resource limits and cancellation
For LLMs10 sections
Name every size, count, and time limit the serving path enforces, what fires when each trips, and what “cancel” really stops — so an overload degrades into refusals, not into stuck processes.
Component law-mcp over --http (stdio has no transport limits),
plus the per-tool caps shared by both transports. Scenario: hostile
or pathological load against a deployed server. All numbers below are
real code constants unless marked otherwise.
Applies to
Section titled “Applies to”| Branch | Coverage in this article |
|---|---|
MCP (law-mcp-server --http) | full limit table, timeout-vs-abandonment semantics |
law serve | caps only (1 MiB body, 16 MiB response, queue 64, 10 000 assertions) in Deploy a private HTTP service and Configuration reference; serve has no graceful drain |
Prerequisites
Section titled “Prerequisites”- Know the three limit layers: edge/reverse proxy (operator’s),
HTTP transport (
http.rs,main.rs), engine/tool caps (law-eval,reference.rs,execution.rs,ask.rs). - Default posture: no call timeout (
LAW_MCP_CALL_TIMEOUTunset means none, real), 1 MiB request cap always on (real), per-tool caps always on (real).
- Size the transport. Real transport limits (
http.rs,main.rs):MAX_BODY= 1 MiB request body →411when empty-length,413when over;LAW_MCP_CALL_TIMEOUTseconds (fractions allowed,0/empty disables) → JSON-RPC-32001withdata.code=CALL_TIMEOUTandlimitSeconds(real). Deployment reference-setting from the code comment (real comment, operator choice):110s under an edgeproxy_read_timeoutof120s, so a pathological call fails with a response, not a broken connection. - Understand timeout semantics. Real mechanism (
http.rsdispatch): with a limit set, the computation runs on a worker thread (law-mcp-call);law-evalhas no cancellation points, so on expiry the client gets the technical refusal while the worker keeps running detached. Its result is discarded; when it finishes, a journal record withMCP_CODE=OVERRUN_FINISHEDlands;/healthzfieldoverrunCallscounts still-running detached workers (real). Timeout bounds the client’s wait, not the CPU spent. - Count the tool caps. Real per-tool limits: argumentation
maxArguments64 /maxExtensions8 /maxUndecided16 defaults, raisable per call vialimits(law-eval/src/ argumentation.rs); structural queryrowLimitdefault 20, at most 5000, else the call is refused (reference.rs); inspect node page 24 000 bytes, over-budget nodes answerRESPONSE_TOO_LARGEwithbudgetBytes(execution.rsPAGE_BYTES); neural appendix accepts at most 8 candidates (reference/search.rs). There is no global CPU/RAM quota and no per-tool wall clock inside the engine. - Bound the outbound legs. Real (
service_client.rs,publication.rs,otlp.rs): service POSTs time out at 30 s with zero redirects; default response cap 4 MiB, Lens answers cap 32 MiB; OTLP spans queue 256 then drop (counter to stderr), one OTLP POST times out at 3 s on a background thread that never blocks a call. - Note what is NOT limited. Real implementation limits by absence:
no cap on concurrent connections (one thread per connection,
http.rsaccept loop), no request queue with backpressure, no response-size cap on the MCP reply path itself, no memory or CPU cgroup. Concurrency control lives at the edge (operator-policy-example: proxylimit_conn,client_max_body_size 1mto matchMAX_BODY,proxy_read_timeout 120s). - Plan overload behavior. Shed order under pressure (reference
behavior assembled from the mechanisms above): over-size bodies
→
413before any work; unknown routes/methods →404/405; timed-out calls →-32001to the client while workers drain; OTLP drops spans rather than slowing calls; the journal never blocks a call (datagram → trimmed resend → stderr). What the server cannot do is reclaim a runaway worker early — size the host so the worst case fits, and restart the process to reclaim (stateless: no sessions,DELETEis refused with405, real).
Expected result
Section titled “Expected result”- Any single request is answered, refused with a code, or timed out
with
-32001; no request hangs pastLAW_MCP_CALL_TIMEOUTplus edge timeouts from the client’s viewpoint. overrunCallsreturns to0after detached workers finish. It growing without bound is compatible with a too-short timeout but also with overload and heavy or wedged computations — so cap concurrency and characterize the queries first (see Capacity and scaling), and only then consider raising the limit while watching memory.- No call path allocates without bound on untrusted size fields:
every list the client sizes (
rowLimit, candidates, node pages) has a hard ceiling (real).
Result check
Section titled “Result check”POST /mcpwith a 2 MiB body →413(real).LAW_MCP_CALL_TIMEOUT=0.001(created-example) on any real call →-32001/CALL_TIMEOUT, then anOVERRUN_FINISHEDjournal line andoverrunCallsback to0.law_querywithrowLimit:5001→ refusal naming the row limit (real);rowLimit:5000→ served page withnextwhen truncated.
Failures and diagnostics
Section titled “Failures and diagnostics”CALL_TIMEOUTon healthy calls: limit below normal latency (checkMCP_MSin the journal), or CPU starvation from too many concurrent workers — add the edge concurrency cap.Cannot spawn workerpath: thread creation failure surfaces asOverrun(real,run_with_deadline): the host is out of threads or memory — an infra error, not a content state.- Handler panic →
-32603withMCP_CODE=INTERNAL(real); the connection thread logs and the process keeps serving.
Support boundaries
Section titled “Support boundaries”- “Supported” load in this article means only: refusal codes and counters behave as named. Throughput numbers, sizing tables, and autoscaling are the operator’s measurements — the repo ships no load profile.
- Cancellation is abandonment, not preemption: CPU/RAM spent by a detached worker is spent. True termination is process restart, which frees the process’s resources; “no sessions” does not settle the fate of an in-flight case/draft mutation — retry terms for effecting operations come from each operation’s own contract (effects matrix), not from the restart.
- Fairness between concurrent calls is the OS scheduler’s; the server has no priority queue.
Next step
Section titled “Next step”- Article 12 (Health and observability) for watching
overrunCallsand call latency from the outside; article 09 (Logs, audit trail, and decision journals) for the journal fields each outcome emits.
Documentation for Arxo. Writings — blog.arxo.io.
Anonymous visit counts on stats.arxo.io, no cookies.