Skip to content
docs
Arxo ↗

Draft workbench

For LLMs11 sections

A draft is a sandbox: it never becomes law, never enters the finished body of acts, and stays invisible to plain questions. This stage takes a draft from pinned bytes to a checked candidate through a fixed round of calls: pin the bytes, check the shape, run it, read it back in words, compare engines, measure span, and diff it against the released edition. The stage sits inside the author loop, after first modeling and before scenarios are declared done: review later reads the verdicts this round produced.

  • The draft text and the source document it claims to follow, held on the author side.
  • Pinned companion bytes for the sources the draft cites, with hashes.
  • The tool reference plus the live tool list of your own seat, which wins over any page when they disagree.
  • A slice that reaches the author-local calls: some public slices leave the whole workbench out.
  • Pin first. The pin call takes the document from you — text or page bytes with an address and a retrieval time — and returns a hash with a ready publication block. The server fetches nothing and invents no time: you hold the document, and its bytes enter every later call as companion sources.
  • Check the shape. The draft check call parses the draft and lowers it; a node that drops out on lowering is a verdict, not a log line — “a rule that can never fire is indistinguishable from silence”. The local check run below shows the same moment for the running example.
  • Run it. The eval call executes the draft in a sandbox and names dead rules one by one; a clean check never promises derivation. The local suite run on the next page is the equivalent moment outside the sandbox.
  • Read it back. The verbalize call pairs pinned text with generated sentences for meaning review, and names the constructs it cannot verbalize rather than skipping them. It never judges correctness — the pairing is the product.
  • Compare engines. The draft comparison call asks the same question through both engines and compares the canonical output bytes; a split is a blocker that gets classified, never quietly patched. Keep this apart from the next two runs: the structural compare below executes nothing — identical lowers report no change — while the suite run executes expectations, and the impact call relates two editions of one act.
  • Measure the span. The measure call reports article span and per-article depth with byte coverage of the pinned text; it takes a draft or a released package. Span alone never proves derivation — full span with dead rules stays possible.
  • Diff the edition. The impact call diffs the draft against the released edition of the same package — same package is a hard rule, never two acts — and reports added, dropped, and changed norms with coverage targets. It does not run scenarios; firing stays with eval.

Exact call names with where each one runs:

Output
(non-runnable sketch; contracts from the tool reference, presence from the live tool list of your seat)
law_pin author-local slice; client brings the document
law_draft_check author-local slice; parse plus lowering survival
law_draft_eval author-local slice; sandboxed execution
law_draft_verbalize author-local slice; text pairs for meaning review
law_draft_differential author-local slice; both engines, byte compare
law_measure draft or package mode; span plus depth
law_draft_impact same package only; edition diff, no scenarios
law engine check local run; observed below, exit zero
law engine lower/diff local runs; observed below, no execution
law test local run; executes the suite, shown on next page
  • Full round or check-only: run the full round when the draft newly claims an article; check-only suffices for wording fixes inside already-covered text.
  • Same-package discipline for the edition diff: a refused cross-act pair first needs the legal-source identity established — renaming the draft to the same package name alone never turns two acts into editions of one.
  • What a split between engines means: stop, classify the cause, and record the class — never smooth it over.
  • Which verdicts travel with the draft: the check verdict, the eval results, the comparison verdict, and the measure numbers go to review; prose summaries of them do not substitute.

The artifact is the checked session: the pin block, the check verdict, the eval results, the verbalization pairs, the comparison verdict, the measure numbers, and the edition diff. The pin block below is quoted from the running-example sources file:

Arxo Law
(quotation, not a run: read from the package sources file)
publication EAI_RU_TEXT of EAI_EDITION { media_type "text/plain; charset=utf-8"; uri "https://zan.gov.kz/api/documents/22880/rus?withHtml=false&page=1"; retrieved_at @2026-09-13T13:00:00+03:00; content_hash "sha256:4cd311a9a3316f3255992116108755f2073aaaaf12b3d41e3c7a1c598c2d3cdb"; local_path "sources/employee-accident-insurance/ru.txt"; }

Record the modeling choices behind the session in the decision record template; the running example carries the tariff decision as EAI-D1 and the threshold decision as EAI-D2, plus EAI-D3 for the penalty with no rounding policy.

The running example is the employee accident insurance package: name kz.corpus.employee_accident_insurance, version 0.1.0, language 0.2, zero dependencies, explicit local imports. The accepted candidate covers five branches across five articles: the employer duty to insure (Article 8), the late-payment penalty (Article 9), tariff classes with premium base and floor (Article 17, with the insured-sum input assumed under Article 16), insurer payout (Article 19), and employer reimbursement (Article 19). Sources are pinned to edition EAI_EDITION with materialization PINNED_UNOFFICIAL_COPY — an Adilet API copy retrieved 2026-09-13, sha256 pinned, local copy kept in the package. The premium base multiplies the insured sum by the class rate: class 22 at 0.0296 gives 29600 KZT on a sum of one million (59200 on a doubled sum), and every other row owns the same shape of check. Payout needs a loss from 30 through 100 percent: loss 30 is established against 29 undecided, loss 100 established against 101 undecided. The penalty multiplies the unpaid sum by 0.015 per day of delay: the suite pins 3000 KZT on round figures and exactly 1.5 KZT on the fractional probe. Modeling choices stand as EAI-D1 (tariff as 22 strict rules), EAI-D2 (inclusive bounds at both ends), EAI-D3 (penalty with no rounding policy), and EAI-D4 (base on the insured sum with the minimum floor).

A clean check that never ran: authors read “parses and lowers” as “derives”, then review finds dead rules the eval call would have named on day one. Its mirror is an edition diff across two acts, which the impact call refuses — establish that both drafts pin the same legal source first, since a shared package name alone proves no such identity. The dullest failure is quoting the tool reference for presence: a call the reference documents can still be absent from your slice, and only the live tool list tells.

Observed runs on the accepted candidate with the pinned tool (law 0.1.0, semantics law.core/0.2, published build):

Terminal
$ law engine check docs/handbook/files/fixtures/eai-candidate
Output
check OK: docs/handbook/files/fixtures/eai-candidate
Terminal
$ law engine lower docs/handbook/files/fixtures/eai-candidate > /tmp/eai-cand-lower.json
$ law engine diff /tmp/eai-cand-lower.json /tmp/eai-cand-lower.json
Output
semantic change: no

Session record for this page — every call of the round with its state:

CallStateGround
law engine checkran, exit 0transcript above
law engine lowerran, exit 0, 54636 bytestranscript above
law engine diff (self)ran, semantic change: notranscript above; structural only, executes nothing
law test (suite)ran, 37 of 37next page, live result
law gen pinning --checkran on the local 0.1.0 build, exit 0 silent; not_run on the published buildpublished 0.1.0 answers unknown command gen (exit 2); pinning runs are labeled with the local binary hash wherever quoted
law engine mutateran on the local 0.1.0 build (67/67/53/14 listing); not_run on the published buildmutate is absent from the published 0.1.0 command surface; full protocol on the properties page
law_pin, law_draft_check, law_draft_eval, law_draft_verbalize, law_draft_differential, law_measure, law_draft_impactnot_runno live MCP seat was observed for this handbook; contracts come from the tool reference, presence is a per-seat condition

Criterion: the check run exits zero, the structural compare of identical lowers reports no change without executing anything, and the record above names a state and a ground for every call — no silent gaps, no claimed runs. The suite run on the next page shows the executing counterpart with its live result.

Drafts never become law, never enter the finished body, and stay invisible to plain questions. The pin call fetches nothing: no document from you, no pin. The installed top-level help lists no pin, measure, audit, or verbalize entry — those calls live on the server side, and absence from your slice is a stated limit, not a hidden feature. The low-level usage names one more comparison-shaped entry whose contract this page does not cover: same name family, no claim made here. The edition diff refuses cross-act pairs. Verbalization never judges correctness. Two builds share one version string but not one command surface: the published 0.1.0 lacks gen and mutate, the local 0.1.0 has them — every run on these pages names its build. And the same candidate under two tool releases disagrees: thirty-seven passing checks under law 0.1.0, twenty-eight Money failures under the 0.1.1 debug build — the next page records that run instead of repeating the pinned claim.

Continue with Coverage and quality, which turns these runs into a report with explicit states.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.