docs← Back to article

Markdown for LLMs

Testing

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Testing

This page describes the suite that ships with the deadline example:
what it checks, in which group each check belongs, how to run all
three suites, and how to add a regression when behavior changes.

```text
docs/build/examples/deadline-app
```

## The four groups

Every check below belongs to exactly one group. The group decides
what a failure means — and what may change it.

| Group | Question it answers | Changed by |
|---|---|---|
| **Subject expectations** | Is the legal answer right? | Only by a corrected hand computation, with the reason stated |
| **Contract checks** | Do the layers trade the shapes they promise? | Only by a deliberate interface change, with callers updated |
| **Structural regressions** | Did the model's report change shape? | By any canon refactor, even one that keeps every subject answer |
| **Reproducibility** | Do the same inputs give the same bytes? | Only by a new build, re-recorded from an inspected run |

The suite is 63 TypeScript checks plus 10 Python checks. Per file:

```text
test/input-schema.test.ts ..... 13   validation + adapter mapping
test/evaluate.test.ts ......... 13   scenarios + synthetic rendering
test/missing.test.ts .......... 14   draft flows + synthetic blockers
test/capture-replay.test.ts ... 10   round-trip + tamper matrix
test/server.test.ts ............ 9   endpoints over handleRequest
test/batch.test.ts ............. 2   fixture run + engine failure
test/contract.test.ts .......... 2   pinned spec + unknown spec
python/test_deadline.py ........ 10   the Python mirror on the engine
```

### Subject expectations

These say what the answer must be for reasons outside the engine.
Each value below was counted by hand from the pleaded provisions
before the engine ever ran:

- collect on the full worked case computes `2026-03-20`: 14 days
  from 6 March, the event day not counted (section 187 (1)),
  ending with day 14 (section 188 (1)), a Friday, so no weekend
  roll.
- collect from 7 March computes `2026-03-23`: day 14 lands on
  Saturday 21 March, so section 193 rolls the end to Monday.
- truth on the candidate `2026-03-20` is `TRUE_ONLY`; truth on a
  wrong candidate is `NEITHER` — not established, not refuted.
- truth on a partial case (duration removed) is `NEITHER`, and the
  draft follow-up asks for `durationDays` only; an empty draft asks
  for both fields.

The same expectations run in both SDKs: `evaluate.test.ts` and
`missing.test.ts` on the TypeScript side, `test_deadline.py` on the
Python side. If the engine and the hand count disagree, the test
fails — the hand count decides, not the snapshot.

### Contract checks

These pin the shapes the layers trade, whatever the law says:

- `validateFormInput` accepts the example input, keeps
  `candidateEnd` optional, and rejects non-calendar dates, empty
  strings, malformed dates, out-of-range durations, and non-object
  bodies — each with its own error.
- `DraftInput` keeps absent fields absent (never defaulted),
  validates present fields strictly, and still requires the legal
  axis.
- the adapter emits exactly the two example facts; the case pins
  the axis, the `Europe/Berlin` time zone, and the policy
  `urn:de:corpus:clir:bgb-fristen#BGB_FRISTEN_TAG`; the draft
  adapter emits only known facts.
- `readCollect`/`readTruth` reduce every engine answer to one of
  `value|claim|not-computed|unknown`.
- the server speaks the fixed-field API only: unknown fields and
  bad bodies are refused with 400, unknown captures with 404.
- `checkContract` opens the pinned spec offline, finds every
  `APP_PREDICATES` entry in the model, and recomputes the example
  case; an unknown spec fails closed with a non-zero exit.

### Structural regressions

These pin the shape of the model's report — valid checks, but not
independent expectations. A canon refactor that keeps every subject answer may still
change them, and then the update is routine, not a legal event:

- `whyNot` over the partial case returns 2 blockers, one per
  candidate rule (`missing.test.ts`: "whyNot exposes two
  alternative routes").
- the example answer's grounds name the `TagesfristEnde` rule
  application (`evaluate.test.ts`: "the example answer carries its
  grounds").
- the `ASKABLE` keys are the engine's full URNs, spelled exactly
  as the blockers spell them.

When one of these fails after a canon move, re-read the new report,
confirm the subject answers still hold, and update the expectation
with the reason — never the other way round.

The blockers' *meaning* (which fields to ask) is covered
separately by six synthetic checks that feed hand-built blocker
graphs — intermediate next to known input, another object,
alternative routes, foreign namespace, unrecognized trigger,
constrained value — so the adapter logic does not depend on the
canon happening to produce every shape.

### Reproducibility

These pin bytes and hashes within one SDK at fixed pins:

- a capture round-trips: replay matches the recorded result hash.
- the tamper matrix (`capture-replay.test.ts`, ten checks) covers
  document-byte edits, input-only edits, case edits, query edits,
  version edits, edited result hashes, document-plus-checksum
  replacement, foreign-model captures, and host-style captures
  without bytes — each with its own expected outcome, spelled out
  in the Save chapter. Three rows pass undetected by design: the
  input-only edit, the versions edit, and the consistent
  document-plus-checksum replacement.
- TypeScript and Python bytes differ by design: each SDK replays
  its own snapshots, never the other's.

Refresh a snapshot only from a run you have read, and only for the
SDK that produced it.

## Rendering without the canon

The view layer is tested on synthetic but valid answers, so no
single canon must naturally produce every state:

```text
an unrecognized evaluationStatus reads as unknown
a multi-valued collect is never truncated
every claim status renders its own headline
an empty collect renders as no date, never as a blank
not-computed and unknown render without derived claims
```

These five live in `evaluate.test.ts` beside the live scenarios.
A new answer shape gets a synthetic check first, then — only if a
canon produces it — a live one.

## Concurrency: what is covered

The model handle is a process-wide singleton (`openModel` caches
one `LawPackage`), and the suite hammers that sharing pattern
directly: "two sequential cases stay isolated" re-runs the example
around a different case, and "ten parallel cases stay isolated"
fires ten concurrent evaluates through `Promise.all` and matches
each answer against its own hand-computed end (including section
193 rolls). The check states its own boundary:

```ts
// Ten concurrent same-process evaluates stay isolated: each answer
// matches its own hand-computed end (event day not counted, weekends
// roll forward). This covers one Node process only — threaded or
// multi-process hosts are not tested here.
```

Threaded hosts, worker pools, and multi-process deployments are
not covered — and even inside one Node process, the claim stays
narrow: these two scenarios pass with the shared singleton. That
is evidence about the scenarios run, not a proof that every
interleaving under every load is safe.

## How to run it

Three commands, all from the unpacked `deadline-app` directory.
All three must pass before a change is done:

```text
npm test
npm run typecheck
python3 python/test_deadline.py
```

`npm test` runs the 63 TypeScript checks; `npm run typecheck`
runs `tsc --noEmit`; the Python file runs its 10 checks as a
plain script (it also runs under pytest when available) in any
environment with `arxo` and the canon installed — the Python SDK
page pins the versions. A typical failure order is: contracts
first (a wrong field or kind), subjects next (a wrong value or
status), structural and bytes last (the report moved).

A fourth command covers what `handleRequest` tests cannot — the
wire itself. It has two modes: local, where a loud skip is acceptable,
and release, where a skip fails:

```text
npm run test:http           # local: SKIP loudly when no socket listens
npm run test:http:strict    # release: STRICT_HTTP=1, skip fails
```

It starts the real `createServer()` on an ephemeral loopback port
(a unix socket where TCP listen is forbidden) and asserts
client-received statuses: 200 on the worked case, 200 on the
capture download for its id with 404/400 on unknown/malformed
ids, 400 on malformed JSON, 400 on an extra field, 413 on an
oversize body. It lives outside `npm test` because sandboxes may
forbid every listen; there the local mode prints `SKIP` loudly
and passes vacuously. A skipped run proves nothing about the
wire — so the release record is only claimed from the strict
mode, run where sockets work. Before trusting the Deployment
network claims, run the strict mode there.

## Adding a regression

1. Write the case as a form: the four fields, with the legal time
   fixed at `2026-09-17`.
2. State the group first: subject (with the one-sentence hand
   reason — the calendar count, the missing fact, the wrong
   candidate), contract (with the promised shape), structural
   (with the report feature), or bytes (with the producing SDK).
3. Add the check to the file its group lives in; prefer a
   synthetic check when the canon need not produce the shape.
4. If the change alters bytes, record a fresh snapshot from a run
   you have inspected, for the SDK that produced it.
5. Run all three suites again: Node, typecheck, and Python.

## Next

- [Security and privacy: data, storage, and logging](/build/application/security-and-privacy/)