Markdown for LLMs
Explaining yes, no, and unknown
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Explaining yes, no, and unknown **Status:** the protocol is prepared and frozen, the seventeen cases are fixed with hashes, and the judging criteria were written before any run. Expected outcomes below are predictions from the statute text, not results. The comparative run has **not** been performed. Every rules system can say "yes". The harder parts are saying why not, and saying what is still missing. This scenario asks four systems — [Blawx](/comparisons/blawx/) with its answer-set engine, [Logical English](/comparisons/logical-english/), [L4](/comparisons/l4/), and Arxo — to explain the same outcomes three ways, and compares the explanations by checks that can be repeated from saved output rather than by taste. ## The text and the questions The cases run on the **Charities (Jersey) Law 2014**, in the consolidation "from 16 October 2025 to Current", published at [jerseylaw.je/laws/current/l_41_2014](https://www.jerseylaw.je/laws/current/l_41_2014) and pinned by hash. The questions come from the question catalogue of the Arxo package for this Law: does an applicant meet the charity test, is a purpose charitable, do all purposes qualify, is public benefit provided and determined, is a purpose purely ancillary, is a name undesirable. Twelve of the cases repeat the shared bank of [One charities text, five systems](/comparisons/jersey-charities/). ## The three explanations | Explanation | Arxo | Blawx with its answer-set engine | Logical English | L4 | |---|---|---|---|---| | Why yes | proof graph: the chain of rule applications, each anchored to a fragment of the source text | the model plus a justification tree, including which section defeated which | the answer plus a proof tree in the words of the rules, with a source badge | the evaluated value plus a trace and its graph | | Why no | status plus the list of blockers: which condition failed or is not established | a proved negation with its justification (a constructive proof of the negative) | "it is not the case that ..." as failure to prove | classical negation, exceptions as "and not"; unknown on partial evaluation | | What is missing | a request for judgment naming the deciding body, a declared package boundary, or an open reading | a hypothetical query: sets of assumptions | unknown, assumable, or judgment-needed markers; expected unknowns in the scenario | empty values or lists; how unknown propagates was not found in the docs | The differences in form are predictable from each system's design and are not defects. "Blockers" against "a proved negation" is a pair to record, not a contest. When a system computes an evaluative term that the text leaves to a named body, that is recorded as a pair too, and settled only against the article. ## The criteria, fixed before the run 1. **Faithful to the outcome.** The explanation states the same sign or status the engine gave on the same inputs, and its grounds reproduce the outcome when rerun on the same pin. 2. **Anchored to the text.** Every ground names an article or paragraph that resolves in the pinned text. A ground without an address cannot be compared; an address that does not resolve counts as a mismatch. 3. **Minimal.** Grounds are removed one at a time. The explanation is minimal if removing any ground breaks the conclusion. Extra grounds are counted, not condemned. 4. **Says what would change the outcome.** The explanation comes with counterfactual edits: for "why no", which facts would lift each blocker; for "what is missing", which answer from the deciding body or which fact would close the question. Each claimed edit is executed and must give the claimed sign. Each check is a script over saved artifacts. Clarity and ease of reading are deliberately **not** scored here; any such remark is filed as an opinion, outside the table. ## The case bank | Case | Situation | Expected outcome (from the text) | Comparability | |---|---|---|---| | EXP-C01 | Complete applicant, listed purpose | meets the charity test: why yes (Article 5(1)) | comparable | | EXP-C02 | Empty purposes, invalid constitution | does not meet the test: why no (Article 5(2)) | comparable | | EXP-C03 | Public benefit, no determination yet | requires judgment by the body named in Article 7(1): what is missing | comparable as a pair (computed value against a request) | | EXP-C04 | Non-empty benefit statement | the statement exists; its sufficiency is a boundary (Article 8(3)(f)) | comparable, with a boundary | | EXP-C05 | "All purposes qualify" on an empty list | open: vacuously true, false on a purposive reading, or no answer until a reading is fixed | comparable once the reading is fixed | | EXP-C06 | Listed purpose: advancement of education | true: why yes (Article 6(1)(b)) | comparable | | EXP-C07 | Purpose argued to be analogous | requires judgment (Article 6(1)(p)): what is missing | comparable as a pair | | EXP-C08 | Government control in the constitution | does not meet the test: why no (Article 5(2)) | comparable | | EXP-C09 | Connection with Jersey | true on the fixture; what is missing when only an address string is given (Articles 11(4)(c), 2(3)) | sign comparable; a substring ground is not | | EXP-C10 | Appeal deadlines of 28 and 56 days | no numbers in the Law; only recorded L4 dates exist | not comparable | | EXP-C11 | Governor with a spent conviction under this Law | reportable, "whether or not spent" (Article 19(1)(f)): why yes | comparable; the text decides against the recorded L4 value | | EXP-C12 | Risk level, "compliant" | no such question in the Law | not comparable | | EXP-C13 | Is a name undesirable | requires judgment: "in the opinion of the Commissioner" (Article 12(1)) | comparable as a pair | | EXP-C14 | Is a purpose purely ancillary | requires judgment (Article 5(1)(a)(ii)) | comparable as a pair | | EXP-R01 | EXP-C03 plus an affirmative determination | true: why yes (Articles 7(1), 5(1)(b)) | comparable | | EXP-R02 | EXP-C07 plus determinations denying the analogy | false: why no (Article 6(1)(p)) | comparable | | EXP-R03 | EXP-C08 plus an order disapplying the exclusion, as a fact of the case | true: why yes (Article 5(3)) | comparable | The three edit cases keep the same text and add facts: each one closes a "what is missing" or flips a "why no", so the counterfactual criterion has something real to check. ## Who can be judged on what - **Arxo** and **L4** can be judged on the whole bank: Arxo through its Jersey package, L4 through the published Jersey model at its pin. - **Blawx** and **Logical English** have no Jersey model. On this bank they are compared by explanation forms only, not executed; executing them is a separate future task. To calibrate the forms, each direction's own prepared experiment serves as a reference set, not as part of the scored bank: twelve cases on a teaching act about flying birds for Blawx, twelve on section 1 of the British Nationality Act 1981 for Logical English, and twelve charity-test cases for L4. ### The bird act, briefly The Blawx calibration set uses a five-section teaching act shipped with the Blawx editor (MIT licence): penguins are birds; birds fly; penguins do not fly, subject to the next section; penguins on planes can fly; and cartoon penguins with jetpacks can fly, except for one named penguin. Its cases cover a plain positive answer, a defeat chain, a double ground, an empty case, classical negation, a hypothetical query, and an explanation that must name both the winning and the defeated section. One case shows why "recorded answer" and "text" are kept apart. For the jetpack penguin named in the exception, the act's wording says it cannot fly; a recorded answer from a sibling project says it can, because that encoding never wires in the exception. The case gets its own outcome — text versus recorded answer, with both comparisons written down — rather than a verdict on either side. A small Arxo model of the bird act exists for this calibration set. It passes the language's static check with no diagnostics; its scenarios are written but not executed. Hypothetical queries have no direct Arxo counterpart: Arxo shows what would continue a derivation, not sets of assumptions, and that case is recorded as a boundary. ## Out of scope - Comparing answer signs for their own sake — that belongs to each direction's own experiment. - Performance measurement. - Readability of the rules as a literary property. - Secondary orders as norms; cases on numbers outside the Law are marked not comparable, not deleted. - Changes to the Arxo package or to any direction's materials. Materials: experiments/comparisons/shared/explanations/ in the project repository. See also the [methodology](/comparisons/methodology/).