docs← Back to article

Markdown for LLMs

Methodology

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Methodology

A comparison here is an experiment, not an essay. Before anything runs, the
question, the hypothesis, and the boundaries are written down. Only then are
the two models built. An Arxo advantage may come out of that process — it is
never assumed going in.

## What gets fixed first

Every experiment starts with a **contract**: the inputs, the entities, the
units, the periods, the editions, the computation stages, and the rounding.
For logical questions the contract separately pins down missing facts,
explicit denial, conflict, judgment requests, and refusal. Anything one side
cannot express is recorded as a limit of comparability, not smuggled in.

For legal calculation the source of meaning is the pinned text of the act.
Other implementations' tests supply expectations to check against. Their
numbers and rules are never treated as the norm, and never copied into the
subject package without the text behind them.

## Reproducibility

The protocol records versions and commits of both systems, dependencies, the
environment, input and model hashes, the generator seed, and the install and
repeat commands. Secrets and personal data never enter the materials. Each
run gets its own results directory with the commands, the actual versions,
the raw output, and a normalized table per case. A historical run is never
overwritten: it keeps its version attached.

## Outcomes per case

Each case keeps both sides' answers and one outcome:

- **match** — the answers agree under the contract;
- **mismatch** — the same question got different answers, cause open;
- **not comparable** — the sides answer different questions;
- **execution error** — a side failed to run, with the cause recorded;
- **not checked** — nothing ran; outside the score, never counted as success.

Numbers agree by exact equality or by a justified tolerance, and the raw
difference stays on record even when a tolerance hides it. A mismatch is
explained through the contract and the sources first: our model may be
wrong, the neighbor's model may be wrong, the editions or stages may
differ, or the semantics may genuinely diverge. An unsettled cause stays
open. A different answer alone never proves a defect.

Timing, when measured, names the mode, the warmup, the repetitions, and the
cost split — build, load, compute, proof construction, serialization —
because different stages answer different questions. A historical timing
never confirms a current engine.

## What agreement does not prove

The close of every experiment names the checked property, the sample size,
the non-comparable cases, and the limits. Agreement on a bank of cases is
not a proof that two languages are equivalent, that two models formalize the
same norm, or that a structural transfer preserved inferential meaning.
Each of those claims needs its own evidence: answer equality needs an
executed reference, a cited article needs a checked formalization, a valid
document needs a meaning-preserving mapping.

## Shared scenarios

Three scenarios run across several systems at once, so related pages can
share one story instead of inventing three:

- **One contract lifecycle** — Symboleo, Accord, Stipula, and Arxo walk the
  same paid-delivery trace: agreement, delivery and payment deadlines, late
  payment and its monetary consequence, suspension and resumption,
  termination, including force majeure, timeouts, and refusals.
- **One charities text** — L4, docassemble, Blawx, Logical English, and Arxo
  work from the same Charities (Jersey) Law text on a shared bank of cases,
  each with its own non-comparable remainder and grounds.
- **One explanation comparison** — Blawx with its answer-set engine,
  Logical English, L4, and Arxo explain the same outcome three ways: why
  yes, why no, and what is missing — judged by fixed criteria, not taste.

All three are designed, frozen, and reproducible. Their comparative runs
have not been performed yet. Each has its own page with the case bank:
[contract lifecycle](/comparisons/contract-lifecycle/), [charities
text](/comparisons/jersey-charities/), [explanations](/comparisons/explanations/).
The runs that were executed elsewhere in the program are collected under
[Measurements](/comparisons/measurements/).

## The machine-author axis

Arxo is built for a world where models write formalizations and people
accept them, so the program also compares how representations treat a
machine author: the cost of reviewing a model against the act text, what a
stock check catches before execution, how an edition change propagates, and
how stable independent writings of the same norm are. The pilot task is one
charities act with several target representations; the author never sees the
case bank, only the act text and a fixed documentation budget. First-try
correctness is recorded for calibration only — it measures familiarity, not
language quality.

## Openness rules

A property the docs do not mention stays open: "not found in the
documentation" never becomes "impossible". A missing fact never becomes an
explicit denial. A language property, an engine-and-version property, a
particular model's choice, extra tooling, and a modelling assumption are
kept apart on every page. When evidence is thin, the page narrows the
claim instead of strengthening the adjective.