Markdown for LLMs
Methodology
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Methodology A comparison here is an experiment, not an essay. Before anything runs, the question, the hypothesis, and the boundaries are written down. Only then are the two models built. An Arxo advantage may come out of that process — it is never assumed going in. ## What gets fixed first Every experiment starts with a **contract**: the inputs, the entities, the units, the periods, the editions, the computation stages, and the rounding. For logical questions the contract separately pins down missing facts, explicit denial, conflict, judgment requests, and refusal. Anything one side cannot express is recorded as a limit of comparability, not smuggled in. For legal calculation the source of meaning is the pinned text of the act. Other implementations' tests supply expectations to check against. Their numbers and rules are never treated as the norm, and never copied into the subject package without the text behind them. ## Reproducibility The protocol records versions and commits of both systems, dependencies, the environment, input and model hashes, the generator seed, and the install and repeat commands. Secrets and personal data never enter the materials. Each run gets its own results directory with the commands, the actual versions, the raw output, and a normalized table per case. A historical run is never overwritten: it keeps its version attached. ## Outcomes per case Each case keeps both sides' answers and one outcome: - **match** — the answers agree under the contract; - **mismatch** — the same question got different answers, cause open; - **not comparable** — the sides answer different questions; - **execution error** — a side failed to run, with the cause recorded; - **not checked** — nothing ran; outside the score, never counted as success. Numbers agree by exact equality or by a justified tolerance, and the raw difference stays on record even when a tolerance hides it. A mismatch is explained through the contract and the sources first: our model may be wrong, the neighbor's model may be wrong, the editions or stages may differ, or the semantics may genuinely diverge. An unsettled cause stays open. A different answer alone never proves a defect. Timing, when measured, names the mode, the warmup, the repetitions, and the cost split — build, load, compute, proof construction, serialization — because different stages answer different questions. A historical timing never confirms a current engine. ## What agreement does not prove The close of every experiment names the checked property, the sample size, the non-comparable cases, and the limits. Agreement on a bank of cases is not a proof that two languages are equivalent, that two models formalize the same norm, or that a structural transfer preserved inferential meaning. Each of those claims needs its own evidence: answer equality needs an executed reference, a cited article needs a checked formalization, a valid document needs a meaning-preserving mapping. ## Shared scenarios Three scenarios run across several systems at once, so related pages can share one story instead of inventing three: - **One contract lifecycle** — Symboleo, Accord, Stipula, and Arxo walk the same paid-delivery trace: agreement, delivery and payment deadlines, late payment and its monetary consequence, suspension and resumption, termination, including force majeure, timeouts, and refusals. - **One charities text** — L4, docassemble, Blawx, Logical English, and Arxo work from the same Charities (Jersey) Law text on a shared bank of cases, each with its own non-comparable remainder and grounds. - **One explanation comparison** — Blawx with its answer-set engine, Logical English, L4, and Arxo explain the same outcome three ways: why yes, why no, and what is missing — judged by fixed criteria, not taste. All three are designed, frozen, and reproducible. Their comparative runs have not been performed yet. Each has its own page with the case bank: [contract lifecycle](/comparisons/contract-lifecycle/), [charities text](/comparisons/jersey-charities/), [explanations](/comparisons/explanations/). The runs that were executed elsewhere in the program are collected under [Measurements](/comparisons/measurements/). ## The machine-author axis Arxo is built for a world where models write formalizations and people accept them, so the program also compares how representations treat a machine author: the cost of reviewing a model against the act text, what a stock check catches before execution, how an edition change propagates, and how stable independent writings of the same norm are. The pilot task is one charities act with several target representations; the author never sees the case bank, only the act text and a fixed documentation budget. First-try correctness is recorded for calibration only — it measures familiarity, not language quality. ## Openness rules A property the docs do not mention stays open: "not found in the documentation" never becomes "impossible". A missing fact never becomes an explicit denial. A language property, an engine-and-version property, a particular model's choice, extra tooling, and a modelling assumption are kept apart on every page. When evidence is thin, the page narrows the claim instead of strengthening the adjective.