Skip to content
docs
Arxo ↗

Writing scenarios

For LLMs11 sections

A scenario pins one question to one expected answer before the run, so a later change either keeps the answer or fails loudly. This stage sits after package declaration, once the static check reports OK, and before review: the model states what follows, scenarios state what counts as correct, and review reads both against the pinned source.

  • The declared package with its static check reporting OK.
  • The pinned fragments and the scope card naming the covered norms.
  • The decision records with their tripwire tests: the decision record template and the filled EAI tariff decision and EAI threshold decision.
  • The test form — given, evaluate, expect — as taught on the writing-tests tutorial page.
  • Plan each norm across seven slots before writing: the main case, missing input, negative test, explicit negation, conflict, exception, and boundary. Not every norm fills every slot; an empty slot needs one line saying why. A negative test feeds a near miss and expects the condition unmet (here: neither established nor denied); an explicit-negation test, which expects a denial, is a different slot — never file one under the other’s name.
  • Write the given part with an explicit context on all four axes plus asserts carrying an id and an origin, so the proof can name the fact behind a wrong answer.
  • Fix every expectation before the run, from the pinned source or from a decision record — never from an observed run.
  • Keep suites small and named by purpose. The accepted candidate keeps five core checks, a world probe plus a property, two finding probes, twenty-two tariff checks, an upper-bound pair, and one fractional probe across six suites.
SlotWhat it probesWhere the expectation comes from
mainthe norm fires and derives the right valuethe pinned article
missingabsent input leaves the question undecidedsilence of a strict rule
negativea near-miss input leaves the condition unmetthe decision on bounds
explicit negationa denied conclusion is derivedthe denial rule, if any
conflicttwo applicable readings meetthe decision on priority, if any
exceptionan escape removes supportthe defeasible construct, if any
boundaryeach bound straddled from both sidesthe decision on inclusive bounds
  • Which slots each norm fills, and the reason for each empty slot.
  • How silence reads: in the running example every rule is strict, so a near miss stays undecided rather than denied, and the expectation says so explicitly.
  • Which decision each boundary pair trips: the thirty-versus-twenty-nine pair trips the lower inclusive bound, the one-hundred-versus-one-hundred-one pair trips the upper inclusive bound — and each pair trips only its own end.

The artifact is the set of suites plus the plan mapping each check to its slot. The check below is quoted from the running-example core suite; it pins the near-miss side of the payout boundary.

Arxo Law
test "EAI-LOSS-TWENTY-NINE-IS-NOT-INSURER-PAYOUT" {
given { context { decision_time @2026-09-13T12:00:00+05:00; knowledge_time @2026-09-13T12:00:00+05:00; legal_time @2026-09-13; timezone "Asia/Almaty"; }
assert "loss": capacity_loss_percent(entity_ref("urn:kz:eai:worker:29"), 29) { origin case_input; }
}
evaluate truth(insurance_payout_due(entity_ref("urn:kz:eai:worker:29")));
expect truth_status == NEITHER;
}

The thirty-seven candidate checks map to the plan as follows.

CheckSlotExpectation
employer must insuremainduty established
class twenty-two premium (core)mainpremium of 29600 KZT derived
late-payment penalty (core)mainpenalty of 3000 KZT derived, exact
loss thirty qualifiesboundarypayout established, trips the lower inclusive bound
loss twenty-nine undecidednegative, boundarycondition unmet, rule stays silent
capacity-loss world probemainpayout established for the probe worker
established-loss propertypropertyevery collected loss-thirty worker is due payout
R1 employer-reimbursement pairmain, negativeloss twenty established, loss four silent
tariff suite, classes 01–22main × 22each premium derived from its source percent
loss one hundred qualifiesboundarypayout established, trips the upper inclusive bound
loss one hundred one silentnegative, boundarycondition unmet, rule stays silent
fractional penaltymainpenalty of exactly 1.5 KZT, trips smuggled rounding
sum-distinguishes premiummainpayroll 1M with sum 2M answers 59200, trips the payroll-base reading
minimum-floor premiumboundarybase 120 against floor 1000 answers the floor
minimum-boundary premiumboundarybase equal to the floor stands

The running example is the employee accident insurance package: name kz.corpus.employee_accident_insurance, version 0.1.0, language 0.2, zero dependencies, explicit local imports. The accepted candidate covers the employer duty to insure, the twenty-two-class tariff with premium base as insured sum times rate plus the minimum floor, insurer payout for capacity loss from thirty through one hundred percent, employer reimbursement for loss five through twenty-nine, and penalty as unpaid times 0.015 times days. Sources are pinned to edition EAI_EDITION with materialization PINNED_UNOFFICIAL_COPY — an Adilet API copy retrieved 2026-09-13, sha256 pinned, local copy kept in the package. The static check reports OK and all thirty-seven checks pass. The per-row tariff checks trip EAI-D1 row by row, the two boundary pairs trip EAI-D2 end by end, the fractional probe trips EAI-D3, the three premium-base probes trip EAI-D4. The source snapshot’s seven checks stay as the before-state: they pass on mutants the candidate kills.

Editing an expectation to match a surprising run. A failing scenario teaches either a model error or a wrong expectation; changing the expected answer without a reason erases the lesson. Two smaller traps come from the tutorial: a context left implicit lets the tool substitute defaults and check something other than meant, and a misspelled expected value fails before any run, naming the file but not the offending word.

The stage is done when the scenario run passes fully and both boundary pairs straddle. Observed on the accepted candidate with the pinned tool (law 0.1.0, semantics law.core/0.2, published build), boundary excerpt:

Terminal
$ law test docs/handbook/files/fixtures/eai-candidate
Output
ok [kz.corpus.employee_accident_insurance#authored] tests/core.lawtest / EAI-LOSS-THIRTY-QUALIFIES
ok [kz.corpus.employee_accident_insurance#authored] tests/core.lawtest / EAI-LOSS-TWENTY-NINE-IS-NOT-INSURER-PAYOUT
ok [kz.corpus.employee_accident_insurance#authored] tests/bounds.lawtest / EAI-LOSS-100-QUALIFIES
ok [kz.corpus.employee_accident_insurance#authored] tests/bounds.lawtest / EAI-LOSS-101-SILENT
total: 37 checked, 37 passed, 0 failed, 0 not run; code 0

Criterion: all thirty-seven lines read ok, and the boundary checks disagree in the expected direction at each end. A bound change that collapses the loss pair fails the checks pinning that pair — the expected failures, no more, no fewer; the premium floor owns its own pair (below-floor, equal) under the same rule.

A suite pins only the points it probes. The candidate leaves honest holes: loss above one hundred one stays silent with no check pinning that silence, negative loss, fractional payroll, and below-payroll insured sums are untested, the missing-input and conflict slots are empty, the explicit-negation slot is empty because no rule denies, and the exception slot is empty by construction since every rule is strict. The established-loss property collects workers at exactly thirty percent, so it re-probes one point rather than sweeping the interval. The property form shown here is narrow: one binder, a collect-shaped domain, and a single-atom expectation. Wider shapes are outside what this page verifies — a two-binder probe in a scratch copy was accepted rather than refused, so probe any wider shape there before relying on it. A passing suite never proves the model matches the act — that reading stays with the author and the reviewer.

Continue with Properties and mutations, which asserts answer-wide properties and probes scenario strength with mutations.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.