Skip to content
docs
Arxo ↗

Explaining yes, no, and unknown

For LLMs6 sections

Status: the protocol is prepared and frozen, the seventeen cases are fixed with hashes, and the judging criteria were written before any run. Expected outcomes below are predictions from the statute text, not results. The comparative run has not been performed.

Every rules system can say “yes”. The harder parts are saying why not, and saying what is still missing. This scenario asks four systems — Blawx with its answer-set engine, Logical English, L4, and Arxo — to explain the same outcomes three ways, and compares the explanations by checks that can be repeated from saved output rather than by taste.

The cases run on the Charities (Jersey) Law 2014, in the consolidation “from 16 October 2025 to Current”, published at jerseylaw.je/laws/current/l_41_2014 and pinned by hash. The questions come from the question catalogue of the Arxo package for this Law: does an applicant meet the charity test, is a purpose charitable, do all purposes qualify, is public benefit provided and determined, is a purpose purely ancillary, is a name undesirable. Twelve of the cases repeat the shared bank of One charities text, five systems.

ExplanationArxoBlawx with its answer-set engineLogical EnglishL4
Why yesproof graph: the chain of rule applications, each anchored to a fragment of the source textthe model plus a justification tree, including which section defeated whichthe answer plus a proof tree in the words of the rules, with a source badgethe evaluated value plus a trace and its graph
Why nostatus plus the list of blockers: which condition failed or is not establisheda proved negation with its justification (a constructive proof of the negative)“it is not the case that …” as failure to proveclassical negation, exceptions as “and not”; unknown on partial evaluation
What is missinga request for judgment naming the deciding body, a declared package boundary, or an open readinga hypothetical query: sets of assumptionsunknown, assumable, or judgment-needed markers; expected unknowns in the scenarioempty values or lists; how unknown propagates was not found in the docs

The differences in form are predictable from each system’s design and are not defects. “Blockers” against “a proved negation” is a pair to record, not a contest. When a system computes an evaluative term that the text leaves to a named body, that is recorded as a pair too, and settled only against the article.

  1. Faithful to the outcome. The explanation states the same sign or status the engine gave on the same inputs, and its grounds reproduce the outcome when rerun on the same pin.
  2. Anchored to the text. Every ground names an article or paragraph that resolves in the pinned text. A ground without an address cannot be compared; an address that does not resolve counts as a mismatch.
  3. Minimal. Grounds are removed one at a time. The explanation is minimal if removing any ground breaks the conclusion. Extra grounds are counted, not condemned.
  4. Says what would change the outcome. The explanation comes with counterfactual edits: for “why no”, which facts would lift each blocker; for “what is missing”, which answer from the deciding body or which fact would close the question. Each claimed edit is executed and must give the claimed sign.

Each check is a script over saved artifacts. Clarity and ease of reading are deliberately not scored here; any such remark is filed as an opinion, outside the table.

CaseSituationExpected outcome (from the text)Comparability
EXP-C01Complete applicant, listed purposemeets the charity test: why yes (Article 5(1))comparable
EXP-C02Empty purposes, invalid constitutiondoes not meet the test: why no (Article 5(2))comparable
EXP-C03Public benefit, no determination yetrequires judgment by the body named in Article 7(1): what is missingcomparable as a pair (computed value against a request)
EXP-C04Non-empty benefit statementthe statement exists; its sufficiency is a boundary (Article 8(3)(f))comparable, with a boundary
EXP-C05“All purposes qualify” on an empty listopen: vacuously true, false on a purposive reading, or no answer until a reading is fixedcomparable once the reading is fixed
EXP-C06Listed purpose: advancement of educationtrue: why yes (Article 6(1)(b))comparable
EXP-C07Purpose argued to be analogousrequires judgment (Article 6(1)(p)): what is missingcomparable as a pair
EXP-C08Government control in the constitutiondoes not meet the test: why no (Article 5(2))comparable
EXP-C09Connection with Jerseytrue on the fixture; what is missing when only an address string is given (Articles 11(4)(c), 2(3))sign comparable; a substring ground is not
EXP-C10Appeal deadlines of 28 and 56 daysno numbers in the Law; only recorded L4 dates existnot comparable
EXP-C11Governor with a spent conviction under this Lawreportable, “whether or not spent” (Article 19(1)(f)): why yescomparable; the text decides against the recorded L4 value
EXP-C12Risk level, “compliant”no such question in the Lawnot comparable
EXP-C13Is a name undesirablerequires judgment: “in the opinion of the Commissioner” (Article 12(1))comparable as a pair
EXP-C14Is a purpose purely ancillaryrequires judgment (Article 5(1)(a)(ii))comparable as a pair
EXP-R01EXP-C03 plus an affirmative determinationtrue: why yes (Articles 7(1), 5(1)(b))comparable
EXP-R02EXP-C07 plus determinations denying the analogyfalse: why no (Article 6(1)(p))comparable
EXP-R03EXP-C08 plus an order disapplying the exclusion, as a fact of the casetrue: why yes (Article 5(3))comparable

The three edit cases keep the same text and add facts: each one closes a “what is missing” or flips a “why no”, so the counterfactual criterion has something real to check.

  • Arxo and L4 can be judged on the whole bank: Arxo through its Jersey package, L4 through the published Jersey model at its pin.
  • Blawx and Logical English have no Jersey model. On this bank they are compared by explanation forms only, not executed; executing them is a separate future task.

To calibrate the forms, each direction’s own prepared experiment serves as a reference set, not as part of the scored bank: twelve cases on a teaching act about flying birds for Blawx, twelve on section 1 of the British Nationality Act 1981 for Logical English, and twelve charity-test cases for L4.

The Blawx calibration set uses a five-section teaching act shipped with the Blawx editor (MIT licence): penguins are birds; birds fly; penguins do not fly, subject to the next section; penguins on planes can fly; and cartoon penguins with jetpacks can fly, except for one named penguin. Its cases cover a plain positive answer, a defeat chain, a double ground, an empty case, classical negation, a hypothetical query, and an explanation that must name both the winning and the defeated section.

One case shows why “recorded answer” and “text” are kept apart. For the jetpack penguin named in the exception, the act’s wording says it cannot fly; a recorded answer from a sibling project says it can, because that encoding never wires in the exception. The case gets its own outcome — text versus recorded answer, with both comparisons written down — rather than a verdict on either side.

A small Arxo model of the bird act exists for this calibration set. It passes the language’s static check with no diagnostics; its scenarios are written but not executed. Hypothetical queries have no direct Arxo counterpart: Arxo shows what would continue a derivation, not sets of assumptions, and that case is recorded as a boundary.

  • Comparing answer signs for their own sake — that belongs to each direction’s own experiment.
  • Performance measurement.
  • Readability of the rules as a literary property.
  • Secondary orders as norms; cases on numbers outside the Law are marked not comparable, not deleted.
  • Changes to the Arxo package or to any direction’s materials.

Materials: experiments/comparisons/shared/explanations/ in the project repository. See also the methodology.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.