Skip to content
docs
Arxo ↗

L4: the charities case bank

For LLMs9 sections

Status: the protocol is prepared and frozen; the sixteen cases and the L4 model files are pinned with hashes. Expected outcomes below are predictions from the statute text and from expectations recorded as comments in the L4 tests — not results. The comparative run has not been performed: neither the L4 engine nor the Arxo side was executed on this bank.

This page is the detailed case bank behind the L4 comparison. Its twelve base cases also feed the five-system scenario on One charities text, five systems.

The Charities (Jersey) Law 2014 has been formalized twice, independently: once as a published L4 model by an outside author, once as an Arxo package. The experiment asks a narrow question: where the text uses an evaluative term — public benefit, an analogous purpose, a connection with Jersey, an empty list of purposes — does the L4 model compute a value while Arxo returns a request for the deciding body’s judgment, and do the signs and grounds diverge only where that difference predicts?

The comparison is between two models of the Law, not between two languages. Shortcuts in the published model — a non-empty string standing for an approved statement, a substring standing for a legal connection — are choices of that model; the L4 language has other ways to encode them.

  • Edition A — the statute alone. The consolidation “from 16 October 2025 to Current”, published by the Jersey Legal Information Board at jerseylaw.je/laws/current/l_41_2014, pinned as one plain-text extraction by hash and size. All twelve base cases use edition A.
  • Edition B — the statute plus a cited secondary order. The L4 model draws on several Jersey orders made under the Law, kept in its own reference folder. Edition B adds exactly the order a case cites: R&O 144/2019 on the timing of annual returns (jerseylaw.je/laws/current/ro_144_2019), R&O 27/2025 on reportable matters, or R&O 60/2018 on the restricted section. The order PDFs are recorded by hash only, because the publisher’s reuse terms were not checked.
  • L4 — the engine at a pinned commit evaluates the model’s test queries (the model’s tests hold 94 evaluation lines, 93 of them with a recorded expected value). Output is the evaluated value with its trace. Recorded expectations are reading material: an L4 test run exits successfully even when an assertion fails, so only the executed values count.
  • Arxo — the existing Jersey package, version 0.1.0, answers the same questions through its own scenarios: a status (true, false, requires judgment, or no answer under an unfixed reading), a request for judgment naming the deciding body where the text names one, and a proof graph.

The L4 model and its tests are pinned with the experiment under the model’s MIT licence. Two of its test queries, with their recorded expectations, as they appear there:

Output
#EVAL `all purposes are charitable` completeCharity
-- Expected: TRUE
#EVAL `has Jersey connection` incompleteCharity
-- Expected: FALSE
CaseSituationRecorded L4 expectationWhat the text saysPredicted Arxo answerComparability
C01Complete applicant: meets the charity test?truethe test of Article 5(1)requires judgment on public benefit; true once the body has answeredcomparable
C02Empty purposes, invalid constitutionfalsefails under Article 5(2)false, on the Article 5(2) groundcomparable; the ground is checked against 5(2), not the empty list
C03Public benefit on the complete applicantcomputed trueto be determined by the body named in Article 7(1)requires judgment, body namedcomparable as a pair: computation against request
C04Valid public benefit statementnon-empty string, trueneeds the Commissioner’s approval (Article 8(3)(f))boundarynot comparable: string check against approval
C05All purposes charitable, on an empty listvacuously trueopen reading of Article 5(1)(a)three answers under three declared readingscomparable once the reading is fixed
C06Advancement of educationtruelisted purpose (Article 6(1)(b))truecomparable
C07An analogous purposecomputed true“may reasonably be regarded as analogous” (Article 6(1)(p))requires judgmentcomparable as a pair
C08Constitution allows government control, plus an exempting orderboolean fieldneeds “acting in that capacity”, and Article 5(3) lets an order disapply itanswer with both conditionscomparable; the ground is checked against the article
C09Connection with Jerseytrue or false by address substringa legal connection, or the Commissioner’s opinion (Articles 11(4)(c), 2(3))connection by Article 2(3) or the Commissioner’s opinionnot comparable: address string against legal connection
C10Appeal deadline, timely or not (28 and 56 days)computed dates and flagsthe Law leaves the deadline to an order (Article 36(2)(a))boundary: delegated to another instrumentnot comparable: numbers with no support in the text
C11Governor with a spent conviction under this Lawfalse, under an assumption about spent convictionsreportable “whether or not spent” (Article 19(1)(f))reportablecomparable; the text decides
C12Risk level, “compliant”computed strings and flagsno such concepts in the Lawrefusal: no such questionnot comparable: question outside the Law

Three outcome predictions follow. C01, C02, and C06 are where the two sides are expected to agree. C03 and C07 are where a computed value meets a request for judgment — a recorded pair, not a defect. C04, C08, C09, and C11 are where the model’s reading and the text can part ways, so the article decides.

Each edit case moves from edition A to edition B and asks how each side learns which answers changed.

CaseWhat changesUnder edition AUnder edition BHow each side learns
R1Annual-return timing: the Timing Order sets “the period of 2 months following the end of the year”a number with no support: not comparablecomparable; the model’s 60-day window approximates two calendar monthsL4: the rule is already written against the order. Arxo: its package states the annual-return duty and declares the deadline delegated to the order; covering it means adding the order and a two-month rule
R2Reportable matters widen: the 2025 order adds convictions involving vulnerable persons and any unspent convictiononly Article 19(1)(a)–(g) appliesthe field widens; a spent conviction involving a vulnerable person becomes reportable by the orderL4: reading the reportable-matter rule against the order. Arxo: seven kinds from the Law, the ordered kind declared as delegated
R3Solicitation defined: the 2018 order gives a four-limbed definitionthe content is left to the orderthe model’s substring test now differs from the order’s definitionL4: reading the solicitation rule. Arxo: declared boundary until the order is added
R4Financial year end: anniversary of registration or an agreed dateas R1the same charity is due or overdue depending on the date chosenL4: a different fixture parameter. Arxo: the same boundary as R1

None of the edit cases touches the evaluative-term pairs of C03 and C07: the orders do not speak to public benefit or analogy.

  1. An address string is not a legal connection (Article 11(4)(c)).
  2. A number with no support in edition A: 28 and 56 days, two months, 7 and 30 days, risk levels, “compliant”.
  3. A string feature against an offence with a bearer, a measure, and an intent (prohibited words, solicitation).
  4. A question outside the Law (risk, compliance).
  5. A case comparable only under edition B: under A it is marked not comparable, naming the missing order — never scored as a win for either side.

Signs and statuses must be strictly equal. There is no numeric tolerance: the Law itself computes no amounts, and numbers appear only in the not-comparable controls. A missing fact on the Arxo side and an empty list or string in the L4 model are recorded explicitly as different things. A computed value against a request for judgment is a pair to settle against the article. Explanations — the Arxo proof graph and the L4 trace — are compared qualitatively, not byte for byte.

The same case bank is also prepared for the machine-author side of the program: a model writes an L4 model and an Arxo package from the Law’s text alone, with an equal documentation budget, never seeing these cases. What is measured is the cost of accepting and maintaining the result — silent errors after convergence, what each system’s own check catches, the cost of the R1–R4 edits — not first-try correctness, which is kept for calibration only. Because the L4 Jersey model is public, any resemblance in names or structure is examined as possible recall, and the Arxo author’s copy excludes the existing Jersey package. This axis is prepared, not run.

  • The full set of 94 test lines; the bank is a twelve-case subset.
  • Secondary orders as norms in the base cases (they enter only through R1–R4).
  • Event-driven simulation of duties over time, beyond single points.
  • Performance measurement.
  • Changes to the Arxo package or its migration.

Materials: experiments/comparisons/l4/ in the project repository. See also L4, One charities text, five systems, and the methodology.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.