Skip to content
docs
Arxo ↗

Testing

For LLMs6 sections

This page describes the suite that ships with the deadline example: what it checks, in which group each check belongs, how to run all three suites, and how to add a regression when behavior changes.

Output
docs/build/examples/deadline-app

Every check below belongs to exactly one group. The group decides what a failure means — and what may change it.

GroupQuestion it answersChanged by
Subject expectationsIs the legal answer right?Only by a corrected hand computation, with the reason stated
Contract checksDo the layers trade the shapes they promise?Only by a deliberate interface change, with callers updated
Structural regressionsDid the model’s report change shape?By any canon refactor, even one that keeps every subject answer
ReproducibilityDo the same inputs give the same bytes?Only by a new build, re-recorded from an inspected run

The suite is 63 TypeScript checks plus 10 Python checks. Per file:

Output
test/input-schema.test.ts ..... 13 validation + adapter mapping
test/evaluate.test.ts ......... 13 scenarios + synthetic rendering
test/missing.test.ts .......... 14 draft flows + synthetic blockers
test/capture-replay.test.ts ... 10 round-trip + tamper matrix
test/server.test.ts ............ 9 endpoints over handleRequest
test/batch.test.ts ............. 2 fixture run + engine failure
test/contract.test.ts .......... 2 pinned spec + unknown spec
python/test_deadline.py ........ 10 the Python mirror on the engine

These say what the answer must be for reasons outside the engine. Each value below was counted by hand from the pleaded provisions before the engine ever ran:

  • collect on the full worked case computes 2026-03-20: 14 days from 6 March, the event day not counted (section 187 (1)), ending with day 14 (section 188 (1)), a Friday, so no weekend roll.
  • collect from 7 March computes 2026-03-23: day 14 lands on Saturday 21 March, so section 193 rolls the end to Monday.
  • truth on the candidate 2026-03-20 is TRUE_ONLY; truth on a wrong candidate is NEITHER — not established, not refuted.
  • truth on a partial case (duration removed) is NEITHER, and the draft follow-up asks for durationDays only; an empty draft asks for both fields.

The same expectations run in both SDKs: evaluate.test.ts and missing.test.ts on the TypeScript side, test_deadline.py on the Python side. If the engine and the hand count disagree, the test fails — the hand count decides, not the snapshot.

These pin the shapes the layers trade, whatever the law says:

  • validateFormInput accepts the example input, keeps candidateEnd optional, and rejects non-calendar dates, empty strings, malformed dates, out-of-range durations, and non-object bodies — each with its own error.
  • DraftInput keeps absent fields absent (never defaulted), validates present fields strictly, and still requires the legal axis.
  • the adapter emits exactly the two example facts; the case pins the axis, the Europe/Berlin time zone, and the policy urn:de:corpus:clir:bgb-fristen#BGB_FRISTEN_TAG; the draft adapter emits only known facts.
  • readCollect/readTruth reduce every engine answer to one of value|claim|not-computed|unknown.
  • the server speaks the fixed-field API only: unknown fields and bad bodies are refused with 400, unknown captures with 404.
  • checkContract opens the pinned spec offline, finds every APP_PREDICATES entry in the model, and recomputes the example case; an unknown spec fails closed with a non-zero exit.

These pin the shape of the model’s report — valid checks, but not independent expectations. A canon refactor that keeps every subject answer may still change them, and then the update is routine, not a legal event:

  • whyNot over the partial case returns 2 blockers, one per candidate rule (missing.test.ts: “whyNot exposes two alternative routes”).
  • the example answer’s grounds name the TagesfristEnde rule application (evaluate.test.ts: “the example answer carries its grounds”).
  • the ASKABLE keys are the engine’s full URNs, spelled exactly as the blockers spell them.

When one of these fails after a canon move, re-read the new report, confirm the subject answers still hold, and update the expectation with the reason — never the other way round.

The blockers’ meaning (which fields to ask) is covered separately by six synthetic checks that feed hand-built blocker graphs — intermediate next to known input, another object, alternative routes, foreign namespace, unrecognized trigger, constrained value — so the adapter logic does not depend on the canon happening to produce every shape.

These pin bytes and hashes within one SDK at fixed pins:

  • a capture round-trips: replay matches the recorded result hash.
  • the tamper matrix (capture-replay.test.ts, ten checks) covers document-byte edits, input-only edits, case edits, query edits, version edits, edited result hashes, document-plus-checksum replacement, foreign-model captures, and host-style captures without bytes — each with its own expected outcome, spelled out in the Save chapter. Three rows pass undetected by design: the input-only edit, the versions edit, and the consistent document-plus-checksum replacement.
  • TypeScript and Python bytes differ by design: each SDK replays its own snapshots, never the other’s.

Refresh a snapshot only from a run you have read, and only for the SDK that produced it.

The view layer is tested on synthetic but valid answers, so no single canon must naturally produce every state:

Output
an unrecognized evaluationStatus reads as unknown
a multi-valued collect is never truncated
every claim status renders its own headline
an empty collect renders as no date, never as a blank
not-computed and unknown render without derived claims

These five live in evaluate.test.ts beside the live scenarios. A new answer shape gets a synthetic check first, then — only if a canon produces it — a live one.

The model handle is a process-wide singleton (openModel caches one LawPackage), and the suite hammers that sharing pattern directly: “two sequential cases stay isolated” re-runs the example around a different case, and “ten parallel cases stay isolated” fires ten concurrent evaluates through Promise.all and matches each answer against its own hand-computed end (including section 193 rolls). The check states its own boundary:

TypeScript
// Ten concurrent same-process evaluates stay isolated: each answer
// matches its own hand-computed end (event day not counted, weekends
// roll forward). This covers one Node process only — threaded or
// multi-process hosts are not tested here.

Threaded hosts, worker pools, and multi-process deployments are not covered — and even inside one Node process, the claim stays narrow: these two scenarios pass with the shared singleton. That is evidence about the scenarios run, not a proof that every interleaving under every load is safe.

Three commands, all from the unpacked deadline-app directory. All three must pass before a change is done:

Output
npm test
npm run typecheck
python3 python/test_deadline.py

npm test runs the 63 TypeScript checks; npm run typecheck runs tsc --noEmit; the Python file runs its 10 checks as a plain script (it also runs under pytest when available) in any environment with arxo and the canon installed — the Python SDK page pins the versions. A typical failure order is: contracts first (a wrong field or kind), subjects next (a wrong value or status), structural and bytes last (the report moved).

A fourth command covers what handleRequest tests cannot — the wire itself. It has two modes: local, where a loud skip is acceptable, and release, where a skip fails:

Output
npm run test:http # local: SKIP loudly when no socket listens
npm run test:http:strict # release: STRICT_HTTP=1, skip fails

It starts the real createServer() on an ephemeral loopback port (a unix socket where TCP listen is forbidden) and asserts client-received statuses: 200 on the worked case, 200 on the capture download for its id with 404/400 on unknown/malformed ids, 400 on malformed JSON, 400 on an extra field, 413 on an oversize body. It lives outside npm test because sandboxes may forbid every listen; there the local mode prints SKIP loudly and passes vacuously. A skipped run proves nothing about the wire — so the release record is only claimed from the strict mode, run where sockets work. Before trusting the Deployment network claims, run the strict mode there.

  1. Write the case as a form: the four fields, with the legal time fixed at 2026-09-17.
  2. State the group first: subject (with the one-sentence hand reason — the calendar count, the missing fact, the wrong candidate), contract (with the promised shape), structural (with the report feature), or bytes (with the producing SDK).
  3. Add the check to the file its group lives in; prefer a synthetic check when the canon need not produce the shape.
  4. If the change alters bytes, record a fresh snapshot from a run you have inspected, for the SDK that produced it.
  5. Run all three suites again: Node, typecheck, and Python.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.