Skip to content
docs
Arxo ↗

Coverage and quality

For LLMs11 sections

Passing scenarios confirm the probed points; this stage asks how much of the act those points cover and how well the package is probed. Five measures answer separately — pinned text, fragment depth, executability, scenarios with properties, and review — each with its own denominator. The stage sits after probing and with review: the report below travels to the release decision, and release reads states, not a single percentage.

  • The pinned texts with hashes and the scope card that fixes which articles count.
  • The lowered form of the package for static reads.
  • The suites with their plans and the property runs from the previous stages.
  • The review record template, the filled EAI review opening finding EAI-R1, the candidate review closing it, and the R3 review closing finding EAI-R3.
  • The passport index for the release record and its report history.
  • Measure pinned text first: bytes pinned against bytes of the whole act. The denominator is the act, not the scope — scope draws the line, the measure shows where the line sits.
  • Measure fragment depth next: per-article depth — executable, anchored, or source-only — with byte coverage of the pinned text. The measure call reports both span and depth for a draft or a released package; span without depth still lets full span pair with dead rules.
  • Measure executability apart: rules that fire against rules written. Lowering survival comes from the check run, dead rules by name from the eval run — both moments are shown for the running example on the previous page.
  • Measure scenarios with properties apart: expectations checked against expectations planned, with the domain size on every property. The suite runs below are this measure for the running example: thirty-seven of thirty-seven on the candidate under law 0.1.0 (both builds), nine of thirty-seven under the 0.1.1 debug build.
  • Measure review apart: findings open against findings raised. EAI-R1 was that row while open; the candidate review closes it, and finding EAI-R2 (the 0.1.1 Money divergence) opens as a compatibility note with a recorded support decision.
  • Never merge the five rows or their states: unmeasured is not zero, unsupported is not passing, missing input is not failure, and a true zero (an empty domain, confirmed) is not an unrun check.

Three instruments, each with verified reach:

  • The measure call covers span and depth only: article span by naming convention and markers of executable content, depth per article, byte coverage of the pinned text. It says nothing about firing or scenarios.
  • The static audit covers the written form: construct inventory against the grammar table, premises that point nowhere, dead symbols, and volume inside articles. Every section of its report carries a state — done, failed, or not checked with a reason — and a missing measure stays a missing measure, never a zero. Note the install honestly: the audit skill names a law audit entry that the installed help does not list, so on this tool the static numbers come from the skill scripts, not from a shipped subcommand.
  • The passport reports cover the release: the build compiles the package, runs its scenarios, replays the saved evaluations, rebuilds the release from itself, and writes the release passport; registry-side reports add quality, compatibility, and performance history, failures included. A fragment measure is not whole-text span, and article-span and rule-by-tests rows that read not run or unsupported must never be read as zero or full percent.
  • The scope card fixes the denominator before measuring: which units count and which stay out. The candidate scope card and the candidate source inventory show the shape at branch granularity.
  • Which zeros are true: an empty collected domain, confirmed by the run report, is a true zero; anything unmeasured is not a number at all. The candidate property collects exactly one worker by construction; the runner reports its status only, so the confirmed-by-run domain size reads missing, not one.
  • Who closes review findings and on what evidence: EAI-R1 closed on the three-state re-run plus the v2 inventory row; EAI-R3 closed on the premium-base re-run plus the v3 inventory rows; EAI-R2 stays open as an engineering case with a support decision, not a verdict blocker.
  • Whether release goes with known failures: the report below shows twenty-eight live Money failures under 0.1.1 against thirty-seven passes under 0.1.0 — release reads that split as a support decision (0.1.0 only), not as a typo to fix in prose.

The artifact is the coverage report: one row per measure with denominator, state, and the run behind the number. States stay literal — done, failed, not run, unsupported, missing, true zero. The candidate report, every number from a shown run:

MeasureDenominatorStateGround
text spanpinned bytes vs act bytesnot_runneeds the author-local law_measure seat; no live seat observed
fragment depthper-unit depth plus byte spannot_runsame as text span
executability (lowering)rules lowered vs rules written (29)donelaw engine check exit 0; lower emits 54636 bytes (local 0.1.0 build)
executability (firing)rules fired vs rules writtenmissingno dead-rule listing observed on this candidate; suite green is not a firing proof
scenarios (0.1.0)37 planned = 37 rundone, 37 passedsuite run on both 0.1.0 builds
scenarios (0.1.1 debug)37 planned = 37 runfailed, 9 passed / 28 failedsuite run; every failure a Money expectation reading NEITHER
property domainworkers collected (1 by construction)missingrunner reports status only, not the domain size
reviewfindings raised (R1, R2, R3) vs open (R2)done with one open noteR1 and R3 closed by re-run; R2 open as engineering case

The running example is the employee accident insurance package: name kz.corpus.employee_accident_insurance, version 0.1.0, language 0.2, zero dependencies, explicit local imports. The accepted candidate covers five branches across five articles: the employer duty to insure (Article 8), the late-payment penalty (Article 9), tariff classes with premium base and floor (Article 17, with the insured-sum input assumed under Article 16), insurer payout (Article 19), and employer reimbursement (Article 19). Sources are pinned to edition EAI_EDITION with materialization PINNED_UNOFFICIAL_COPY — an Adilet API copy retrieved 2026-09-13, sha256 pinned, local copy kept in the package. Applied literally: text span and depth read not_run for want of a measure seat; executability rests on a clean check and lower of twenty-nine rules; scenarios pin the per-row tariffs, both boundary pairs, the fractional penalty, and the premium base with its floor; review carries EAI-R1 and EAI-R3 closed and EAI-R2 open as a compatibility note.

The single percentage: “eighty-seven percent covered” merges five denominators into one number that answers nothing. Its twin is the silent denominator — a property over an empty domain, a span over the scope instead of the act. The dullest failure is editing the report to match the frozen claim: the live run below fails twice, and those two rows stay failed no matter what the previous release claimed.

The compatibility case, same candidate inputs under two environments. First observed on the source tree — seven checks, two Money failures under the 0.1.1 debug build:

Output
FAIL [kz.corpus.employee_accident_insurance#authored] tests/core.lawtest / EAI-MINING-CLASS-22-PREMIUM
truth_status == TRUE_ONLY: in the document NEITHER
FAIL [kz.corpus.employee_accident_insurance#authored] tests/core.lawtest / EAI-SPECIAL-LATE-PAYMENT-PENALTY
truth_status == TRUE_ONLY: in the document NEITHER
total: 7 checked, 5 passed, 2 failed, 0 not run; code 1

Confirmed on the accepted candidate with tool version law 0.1.1, semantics law.core/0.2, binary hash sha256:7fddf8d081e7fd527c36cc0393aea6cbf9e00ea814e960a76ac701f74794c01f (debug build):

Terminal
$ law test docs/handbook/files/fixtures/eai-candidate
Output
FAIL [kz.corpus.employee_accident_insurance#authored] tests/tariff.lawtest / EAI-TARIFF-CLASS-01-PREMIUM
truth_status == TRUE_ONLY: in the document NEITHER
...
FAIL [kz.corpus.employee_accident_insurance#authored] tests/penalty-fractional.lawtest / EAI-PENALTY-FRACTIONAL-EXACT
truth_status == TRUE_ONLY: in the document NEITHER
total: 37 checked, 9 passed, 28 failed, 0 not run; code 1

The failing set is exactly the twenty-eight Money-expectation checks — all twenty-two tariff rows, the core class-22 premium, both penalties, all three premium-base probes — each reading NEITHER where TRUE_ONLY is expected, while all nine non-Money checks pass. Same inputs (candidate snapshot id above), two environments, differing answers: the divergence is localized to the Money path and recorded as finding EAI-R2 with status open-engineering-case and the support decision law 0.1.0 only.

Criterion: the report is done when every row names its denominator and its state, no two states share one number, the twenty-eight failures above stand recorded as failed, and the support decision names the environment the candidate keeps. The thirty-seven-pass record describes the same files under law 0.1.0 on both builds, this run describes them under the 0.1.1 debug build, and the report keeps every run labeled with its binary hash.

The measure call needs pinned bytes and says nothing about firing. The static audit reads the written form only: each finding closes by a named scenario or a recorded verdict, never by prose. The passport confirms only recorded expectations and checks. The installed help lists no measure or audit entry, and the suite runner itself runs unlisted — on this install, presence is established by running, and absence is a limit, not a promise. The two scenario records differ by tool release alone — same files, same suite — and only the labeled run tells which is which. The Money-path localization is an observed regularity, not a diagnosed cause: no engine internals were inspected for this handbook. And no measure proves the model matches the act: span, firing, passing, and review each bound a different doubt.

Continue with Semantic review, which reads the model against the pinned text line by line and records findings.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.