docs← Back to article

Markdown for LLMs

Resource limits and cancellation

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Resource limits and cancellation

## Goal

Name every size, count, and time limit the serving path enforces,
what fires when each trips, and what "cancel" really stops — so an
overload degrades into refusals, not into stuck processes.

## Scope

Component `law-mcp` over `--http` (stdio has no transport limits),
plus the per-tool caps shared by both transports. Scenario: hostile
or pathological load against a deployed server. All numbers below are
real code constants unless marked otherwise.

## Applies to

| Branch | Coverage in this article |
|---|---|
| MCP (`law-mcp-server --http`) | full limit table, timeout-vs-abandonment semantics |
| `law serve` | caps only (1 MiB body, 16 MiB response, queue 64, 10 000 assertions) in [Deploy a private HTTP service](/operate/private-http-service/) and [Configuration reference](/operate/configuration-reference/); serve has no graceful drain |

## Prerequisites

- Know the three limit layers: edge/reverse proxy (operator's),
  HTTP transport (`http.rs`, `main.rs`), engine/tool caps
  (`law-eval`, `reference.rs`, `execution.rs`, `ask.rs`).
- Default posture: no call timeout (`LAW_MCP_CALL_TIMEOUT` unset
  means none, real), 1 MiB request cap always on (real),
  per-tool caps always on (real).

## Steps

1. Size the transport. Real transport limits (`http.rs`, `main.rs`):
   `MAX_BODY` = 1 MiB request body → `411` when empty-length,
   `413` when over; `LAW_MCP_CALL_TIMEOUT` seconds (fractions
   allowed, `0`/empty disables) → JSON-RPC `-32001` with
   `data.code=CALL_TIMEOUT` and `limitSeconds` (real). Deployment
   reference-setting from the code comment (real comment, operator
   choice): `110` s under an edge `proxy_read_timeout` of `120` s,
   so a pathological call fails with a response, not a broken
   connection.
2. Understand timeout semantics. Real mechanism (`http.rs`
   `dispatch`): with a limit set, the computation runs on a worker
   thread (`law-mcp-call`); `law-eval` has no cancellation points,
   so on expiry the client gets the technical refusal while the
   worker keeps running detached. Its result is discarded; when it
   finishes, a journal record with `MCP_CODE=OVERRUN_FINISHED`
   lands; `/healthz` field `overrunCalls` counts still-running
   detached workers (real). Timeout bounds the client's wait, not
   the CPU spent.
3. Count the tool caps. Real per-tool limits: argumentation
   `maxArguments` 64 / `maxExtensions` 8 / `maxUndecided` 16
   defaults, raisable per call via `limits` (`law-eval/src/
   argumentation.rs`); structural query `rowLimit` default 20, at
   most 5000, else the call is refused (`reference.rs`); inspect
   node page 24 000 bytes, over-budget nodes answer
   `RESPONSE_TOO_LARGE` with `budgetBytes` (`execution.rs`
   `PAGE_BYTES`); neural appendix accepts at most 8 candidates
   (`reference/search.rs`). There is no global CPU/RAM quota and no
   per-tool wall clock inside the engine.
4. Bound the outbound legs. Real (`service_client.rs`,
   `publication.rs`, `otlp.rs`): service POSTs time out at 30 s
   with zero redirects; default response cap 4 MiB, Lens answers
   cap 32 MiB; OTLP spans queue 256 then drop (counter to stderr),
   one OTLP POST times out at 3 s on a background thread that never
   blocks a call.
5. Note what is NOT limited. Real implementation limits by absence:
   no cap on concurrent connections (one thread per connection,
   `http.rs` accept loop), no request queue with backpressure, no
   response-size cap on the MCP reply path itself, no memory or CPU
   cgroup. Concurrency control lives at the edge
   (operator-policy-example: proxy `limit_conn`, `client_max_body_size
   1m` to match `MAX_BODY`, `proxy_read_timeout 120s`).
6. Plan overload behavior. Shed order under pressure (reference
   behavior assembled from the mechanisms above): over-size bodies
   → `413` before any work; unknown routes/methods → `404`/`405`;
   timed-out calls → `-32001` to the client while workers drain;
   OTLP drops spans rather than slowing calls; the journal never
   blocks a call (datagram → trimmed resend → stderr). What the
   server cannot do is reclaim a runaway worker early — size the
   host so the worst case fits, and restart the process to reclaim
   (stateless: no sessions, `DELETE` is refused with `405`, real).

## Expected result

- Any single request is answered, refused with a code, or timed out
  with `-32001`; no request hangs past `LAW_MCP_CALL_TIMEOUT` plus
  edge timeouts from the client's viewpoint.
- `overrunCalls` returns to `0` after detached workers finish. It
  growing without bound is *compatible with* a too-short timeout
  but also with overload and heavy or wedged computations — so cap
  concurrency and characterize the queries first (see
  [Capacity and scaling](/operate/capacity-scaling/)), and only then
  consider raising the limit while watching memory.
- No call path allocates without bound on untrusted size fields:
  every list the client sizes (`rowLimit`, candidates, node pages)
  has a hard ceiling (real).

## Result check

- `POST /mcp` with a 2 MiB body → `413` (real).
- `LAW_MCP_CALL_TIMEOUT=0.001` (created-example) on any real call →
  `-32001`/`CALL_TIMEOUT`, then an `OVERRUN_FINISHED` journal line
  and `overrunCalls` back to `0`.
- `law_query` with `rowLimit:5001` → refusal naming the row limit
  (real); `rowLimit:5000` → served page with `next` when truncated.

## Failures and diagnostics

- `CALL_TIMEOUT` on healthy calls: limit below normal latency (check
  `MCP_MS` in the journal), or CPU starvation from too many
  concurrent workers — add the edge concurrency cap.
- `Cannot spawn worker` path: thread creation failure surfaces as
  `Overrun` (real, `run_with_deadline`): the host is out of threads
  or memory — an infra error, not a content state.
- Handler panic → `-32603` with `MCP_CODE=INTERNAL` (real); the
  connection thread logs and the process keeps serving.

## Support boundaries

- "Supported" load in this article means only: refusal codes and
  counters behave as named. Throughput numbers, sizing tables, and
  autoscaling are the operator's measurements — the repo ships no
  load profile.
- Cancellation is abandonment, not preemption: CPU/RAM spent by a
  detached worker is spent. True termination is process restart,
  which frees the process's resources; "no sessions" does not
  settle the fate of an in-flight case/draft mutation — retry
  terms for effecting operations come from each operation's own
  contract ([effects matrix](/operate/tool-profiles-input-policy/#effects-the-authoritative-matrix)),
  not from the restart.
- Fairness between concurrent calls is the OS scheduler's; the
  server has no priority queue.

## Next step

- Article 12 ([Health and observability](/operate/health-observability/)) for watching `overrunCalls` and call latency from the
  outside; article 09 ([Logs, audit trail, and decision journals](/operate/logs-audit-decision-journals/)) for the journal fields each outcome emits.