docs← Back to article

Markdown for LLMs

One charities text, five systems

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# One charities text, five systems

**Status:** the protocol is prepared and frozen, and the twelve shared
cases are fixed with hashes. Expected outcomes below are predictions from
the statute text and from expectations recorded in another project's
tests — not results. The comparative run has **not** been performed for
any of the five systems.

Five systems meet one pinned statute: [L4](/comparisons/l4/),
[docassemble](/comparisons/docassemble/), [Blawx](/comparisons/blawx/),
[Logical English](/comparisons/logical-english/), and Arxo. They do very
different jobs — computing rules, interviewing a person, explaining
answers, writing rules in controlled English — so the point of the
scenario is not one score. It is a classification: for each case, what
each system can take from it, and where it is not comparable, on what
ground.

## The text

The **Charities (Jersey) Law 2014**, in the consolidation in force "from
16 October 2025 to Current", published by the Jersey Legal Information
Board at
[jerseylaw.je/laws/current/l_41_2014](https://www.jerseylaw.je/laws/current/l_41_2014).
One plain-text extraction of that page is pinned by hash and size, and it
is the only arbiter for every system. Secondary orders made under the Law
are not part of this shared scenario; the L4 study treats them as a
separate edition, described on [L4: the charities case
bank](/comparisons/l4-charities-case-bank/).

The cases touch the charity test (Article 5), the list of charitable
purposes (Article 6), public benefit (Article 7), the benefit statement
on registration (Article 8), the Jersey connection (Article 11), matters
a governor must report (Article 19), and appeal deadlines (Article 36).

## Who brings what

- **L4** — a published L4 model of this Law by an independent author
  (MIT licence), with tests whose expected values are written as
  comments. Those recorded expectations are one source of the predictions
  below; they are reading material until the pinned engine and model are
  run.
- **docassemble** — a new guided interview for the charity test, to be
  written for the run. It asks a person the questions the text leaves to
  a person, branches on the answers, and assembles a document.
- **Blawx** and **Logical English** — no Jersey model exists for either.
  Their prepared experiments use other texts (a teaching act about birds
  for Blawx, section 1 of the British Nationality Act 1981 for Logical
  English). Here they take part through **patterns** they share with the
  Jersey cases — a defeating exception, a proved denial, a scene fact
  versus a request for judgment — and every case that needs a Jersey
  model is not comparable for them.
- **Arxo** — the existing Arxo package for this Law, version 0.1.0, with
  its own scenarios for the charity test, the register, and governors.
  Where the text names a deciding body for an evaluative term, Arxo
  returns a request for that body's judgment instead of a computed value.

## The case bank

"Expected" comes from the statute text, or from the recorded L4 test
expectation when the text and that expectation agree; where they differ,
the text decides and the difference is kept on record.

| Case | Situation | Expected (from the text) | L4 | docassemble | Blawx / Logical English | Arxo |
|---|---|---|---|---|---|---|
| J01 | Complete applicant, full list of purposes | meets the charity test (Article 5(1)) | computed | normal interview path to a document | pattern only: normal path | answer, with the public-benefit judgment supplied or requested |
| J02 | Empty purpose list, invalid constitution | does not meet the test (Article 5(2)) | computed; the ground must be checked against 5(2) | refusal screen, no further questions | pattern only: proved denial | false, on the Article 5(2) ground |
| J03 | Public benefit on a complete applicant | for the Commissioner, tribunal, or court to determine (Articles 5(1)(b), 7(1)) | computed true | question put to a person | Logical English pattern: scene fact versus request | request for judgment, deciding body named |
| J04 | Non-empty public benefit statement | the text requires the Commissioner's approval (Article 8(3)(f)) | non-empty string counts as valid | path continues | not comparable | boundary: approval is not a string check |
| J05 | "All purposes are charitable" on an empty list | open: depends on how the quantifier is read | vacuously true | default branch, reading recorded | not comparable | three answers, one per declared reading |
| J06 | "Advancement of education" | a charitable purpose (Article 6(1)(b)) | true | normal path | not comparable | true |
| J07 | A purpose analogous to the listed ones | an evaluation: "may reasonably be regarded as analogous" (Article 6(1)(p)) | computed true | must show a judgment question | Logical English pattern: scene fact versus request | request for judgment |
| J08 | Government control in the constitution, plus an exempting order | excluded unless an order disapplies it (Article 5(2)–(3)) | boolean field | refusing or continuing branch, citing the exemption | pattern: defeating exception or default with rebuttal | answer with the "acting in that capacity" condition and the exemption |
| J09 | Connection with Jersey | a legal connection or the Commissioner's opinion (Articles 11(4)(c), 2(3)) | substring of the address: not comparable on the ground | opinion question; a string-based version is not comparable | not comparable | connection by Article 2(3) or the Commissioner's opinion |
| J10 | Appeal deadlines of 28 and 56 days | the Law leaves the deadline to an order; no numbers in the text (Article 36(2)(a)) | not comparable | not comparable | not comparable | boundary: delegated to another instrument |
| J11 | Governor with a spent conviction under this Law | reportable: "whether or not spent" (Article 19(1)(f)) | recorded false, under an assumption about spent convictions; checked against the text | reporting or disqualification branch | not comparable | reportable |
| J12 | Risk level, "compliant" | concepts the Law does not contain | not comparable | not comparable | not comparable | refusal: no such question in the Law |

A skipped answer is not a separate case: it runs through every case as a
shared property. Arxo reports the missing facts; the interview asks
again.

## Not comparable, and on what ground

A case is marked, never deleted. Six grounds are fixed before any run:

1. **An address string is not a legal connection.** Testing whether an
   address contains "Jersey" answers a different question from Article
   11(4)(c).
2. **A number with no support in the pinned text.** The 28 and 56 days,
   a two-month window, risk levels, and "compliant" come from secondary
   orders or from the model author, not from the Law.
3. **A string feature versus an element with a bearer, a measure, and an
   intent.** Prohibited words and solicitation are offences in the text,
   not substrings.
4. **A question outside the Law.**
5. **Different readings of an empty list, with no reading fixed.** Once a
   reading is fixed, the case becomes comparable.
6. **A pattern-only system.** Blawx and Logical English have no Jersey
   model; executing them on this bank is a separate future task.

## When answers count as the same

- **Match** — the same sign or status, strictly equal. There is no
  numeric tolerance, because the Law contains no amounts to compute.
- **Mismatch** — the same question got different answers. A computed
  value against a request for judgment is recorded as a pair and is not
  a defect on either side until it is checked against the article.
- **Not comparable** — one of the six grounds above, named per case.

Grounds are compared separately from signs: a matching "false" on J02
still records whether it rests on Article 5(2) or on the empty list.
Explanations — an Arxo proof graph, an L4 evaluation trace, the
interview's screen log, a Blawx justification or a Logical English
explanation on their own texts — are compared as qualitative pairs; the
dedicated comparison is on [Explaining yes, no, and
unknown](/comparisons/explanations/).

## What each system is predicted to take

- **L4** takes J01–J09 and J11. Not comparable: the ground of J09, J10,
  and J12. J04, J05, and J08 need the ground checked against the article.
- **docassemble** takes J01–J09 and J11 as interview branches. Not
  comparable: J10, J12, J09 when implemented as a string, and J02 or J05
  without a fixed reading.
- **Blawx** takes nothing executable from this bank. Comparable patterns:
  defeating exception (J08), proved denial (J02), the scope of an
  exception.
- **Logical English** takes nothing executable from this bank.
  Comparable patterns: scene fact versus request for judgment (J03, J07,
  J09), default with rebuttal (J08).
- **Arxo** takes J01–J09 and J11 through its Jersey package. Boundaries:
  J04 (Commissioner's approval), J10 (deadline delegated to an order),
  J12 (no such question). J03, J07, and J09 return an observable request
  for judgment; J05 returns one answer per declared reading.

## Out of scope

- Secondary orders as norms, and the edit cases they create (handled on
  the L4 case-bank page as a second edition).
- Comparing two revisions of the Law's text (a separate docassemble
  experiment).
- Performance measurement.
- Changes to the Arxo package or its migration.
- Writing new Jersey models for Blawx or Logical English.

Agreement on these twelve cases, when the run happens, will not show that
any two of these systems formalize the Law the same way.

Materials: experiments/comparisons/shared/je-charities/ in the project
repository. See also the [methodology](/comparisons/methodology/).