Skip to content
docs
Arxo ↗

Semantic review

For LLMs11 sections

Semantic review asks whether the model says what the pinned source says — no more, no less. It sits after scenarios pass: the static check has reported OK and the suites are green, so the remaining risk is meaning, not form. The reviewer reads each covered provision against its rules, tests each reading decision with a counterexample, and records every mismatch as a finding that a re-run must close.

  • Read provision by provision: for each covered paragraph, name the rules that express it and the rules that touch it without expressing it.
  • Hunt the two classic misses: a paragraph with no rule, and a rule with no paragraph. Both are findings even when every scenario passes.
  • Turn each miss into a counterexample: a concrete input where source and model disagree, small enough to become one scenario.
  • Write the fix as a scenario first, then as a rule change, then re-run the full suite: the new scenario must pass and every control scenario must stand. A control is an old scenario whose expectation the source still supports — it never moves. An old scenario whose expectation the source contradicts is not a control but a wrong expectation: fix it with an independent justification from the pinned text, and record the change as its own finding, never as silent editing.
  • Keep the fix out of the reviewed package until the opinion is written: draft it in a copy, so the opinion describes the version it read.
  • Finding or note: a finding claims a source-model mismatch and needs a counterexample; a note records style or doubt and needs none.
  • Scope of the opinion: which articles were read, which were skimmed, and which were left out. An opinion without a scope statement overclaims.
  • Severity per finding: does the mismatch change an answer on a realistic input, or only on an edge the scope card already excludes?

The artifact is the review opinion with its scope: one record per finding with the quoted source line, the counterexample, the fix, and the re-run result, plus a verdict per covered norm. The running example carries three opinions. The filled EAI review record opens finding EAI-R1 against the source snapshot; the candidate review record closes it against the accepted candidate snapshot; the R3 review record closes finding EAI-R3 — premium base read payroll instead of the insured sum, minimum floor missing — against the revised snapshot after an external re-review caught it under a green suite. All three reviews are labeled AGENT reviews: each was produced by an automated reviewer reading the pinned source text, and none claims human approval.

The running example is the employee accident insurance package: name kz.corpus.employee_accident_insurance, version 0.1.0, language 0.2, zero dependencies, explicit local imports. The accepted candidate covers the employer duty to insure, the twenty-two-class tariff with premium base as insured sum times rate plus the minimum floor, insurer payout for capacity loss from thirty through one hundred percent, employer reimbursement for loss five through twenty-nine, and penalty as unpaid times 0.015 times days. Sources are pinned to edition EAI_EDITION with materialization PINNED_UNOFFICIAL_COPY — an Adilet API copy retrieved 2026-09-13, sha256 pinned, local copy kept in the package. The static check reports OK and all thirty-seven candidate checks pass. The per-row tariff checks trip EAI-D1, the two boundary pairs trip EAI-D2, the fractional probe trips EAI-D3, the three premium-base probes trip EAI-D4.

Finding EAI-R1: the source model covers only the insurer branch of article nineteen. The act first assigns lower loss to the employer (pinned bytes, article 19 point 1, line 363; the insurer branch follows on line 364):

Возмещение вреда, связанного с утратой заработка (дохода) работником в связи с установлением ему степени утраты профессиональной трудоспособности от пяти до двадцати девяти процентов включительно, осуществляется страхователем согласно трудовому законодательству Республики Казахстан.

[translation] Compensation for harm from established capacity loss from five through twenty-nine percent inclusive is paid by the policyholder under labour legislation.

Classification: a gap inside a promised article. The v1 scope card lists article 19 as covered at whole-article granularity while this adjacent branch has no rule, no test, and no recorded exclusion — so a reader is entitled to expect it, and its absence is a finding, not a reading dispute. Had the scope card named the insurer branch as the entire contract with the employer branch explicitly excluded, the same absence would have been a scope decision to cite, not a mismatch to fix. The counterexample is a worker with twenty percent loss: the source expects employer-side reimbursement, while the model derives nothing at all — the insurer payout stays undecided and no employer-side conclusion exists. The fix is one strict rule plus two probe checks, drafted in the R1 review copy with its probe checks and carried into the accepted candidate: loss twenty established, loss four silent.

Reviewing the scenarios instead of the source. A green suite proves the model matches its own expectations; only the pinned text can show a missing branch. The EAI suite was fully green while the five-through-twenty-nine branch was absent — the miss surfaced only when the reviewer read article nineteen line by line. A second trap is fixing the package mid-review: the opinion then describes a version nobody can re-read. Draft the fix in a copy, finish the opinion, then hand both over.

The stage is done when every finding carries a counterexample with a re-run. Observed runs with tool version law 0.1.0, semantics law.core/0.2. Before the fix, the probe suite cannot even run — the predicate it asks about does not exist (seven ok lines precede the skip):

Terminal
$ law test docs/handbook/files/fixtures/review-r1
Output
SKIP [kz.corpus.employee_accident_insurance#authored] tests/r1.lawtest
<repo>/docs/handbook/files/fixtures/review-r1/tests/r1.lawtest: тест не лоуверится (LDC-E2102: предикат "low_loss_employer_reimbursement_due" не объявлен ни делом, ни предъявленной программой: голое имя достроилось бы в urn:kz:corpus:clir:employee-accident-insurance#low_loss_employer_reimbursement_due, которого нет в мире (§188))
итого: 8 проверено, 7 прошли, 0 не прошли, 1 не исполнены; код 2

[translation] The suite is skipped: the predicate is declared neither by the case nor by the program. The summary reads: eight checked, seven passed, zero failed, one unexecuted, exit code two. (<repo> stands for the local checkout path, elided here.)

The middle state separates the missing question from the missing answer. With the predicate declared but the rule still absent, the same probe runs and answers neither — the question exists, the derivation does not:

Output
FAIL [kz.corpus.employee_accident_insurance#authored] tests/r1.lawtest / EAI-R1-LOW-LOSS-TWENTY-EMPLOYER-REIMBURSES
truth_status == TRUE_ONLY: в документе NEITHER
итого: 37 проверено, 36 прошли, 1 не прошли, 0 не исполнены; код 1

[translation] The failing line reads: truth status expected established-only, the document holds neither. The summary reads: thirty-seven checked, thirty-six passed, one failed, zero unexecuted, exit code one.

After adding the one-rule fix, the static check and the full candidate suite pass (pinned tool, published build):

Terminal
$ law engine check docs/handbook/files/fixtures/eai-candidate
check OK: docs/handbook/files/fixtures/eai-candidate
$ law test docs/handbook/files/fixtures/eai-candidate
Output
ok [kz.corpus.employee_accident_insurance#authored] tests/r1.lawtest / EAI-R1-LOW-LOSS-TWENTY-EMPLOYER-REIMBURSES
ok [kz.corpus.employee_accident_insurance#authored] tests/r1.lawtest / EAI-R1-LOW-LOSS-FOUR-SILENT
total: 37 checked, 37 passed, 0 failed, 0 not run; code 0

Criterion: the before run shows the gap as an unexecuted suite with LDC-E2102, the middle run shows the declared-but-underived question answering neither, and the after run shows thirty-seven of thirty-seven passing with every control check unmoved. The review-r1 copy keeps the nine-of-nine demonstration; the candidate review record binds the R1 closure to the pre-R3 snapshot, and the R3 review record binds the premium-base closure to the revised snapshot. The real package is untouched throughout.

A semantic review covers only the provisions it names, and it reads only the pinned edition — a later amendment reopens every finding and verdict. The EAI AGENT reviews read five articles and left the rest of the act out of scope; their opinions say nothing about unreviewed articles. Review findings also inherit the source copy’s standing: PINNED_UNOFFICIAL_COPY means the bytes are pinned but unofficial, so a verdict can cite the copy but cannot speak for the official publication.

Continue with Technical review, which checks the package as an engineering artifact: structure, pins, dependencies, and reproducibility.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.