# Measurements Seven comparative runs have been executed so far. All other comparisons in this section are protocols that have been prepared but not run. Each run below has its own page with its date, versions, setup, numbers and limits. This page lists what each run shows and what it does not. | Run | Date | What ran | Size | |---|---|---|---| | [Catala: the parity run](/comparisons/catala-parity-run/) | 9 September 2026, extended 22 September 2026 | Catala 1.2.0 interpreter and Arxo on one act | 65 cases, 240 random inputs, mutation sets | | [LLM / RAG: the US tax code run](/comparisons/llm-rag-irc-run/) | 2 October 2026 | Two models in four modes: no help, retrieval, tools, retrieval plus tools | 4 main questions plus two side checks; one request per model, question and mode | | [Science tasks: models with and without tools](/comparisons/science-with-and-without-tools/) | 1 October 2026 | Three models on three tasks, without tools and with Arxo tools | One run per task and mode | | [WebAssembly: the i32 spec-tests run](/comparisons/wasm-spec-tests-run/) | 2 October 2026 | Arxo against the official integer test script | 19 line-cases from 14 directives | | [Computer algebra: the textbook-steps run](/comparisons/computer-algebra-run/) | 2 October 2026 | SymPy, Wolfram and Arxo on a bank and textbook cases | 129 bank rows plus 15 OpenStax cases | | [Lean: the divisibility-instances run](/comparisons/lean-run/) | 2 October 2026 | Lean core and Arxo on divisibility instances | 14 cases | | [FIDE: the Dutch-pairing check run](/comparisons/fide-pairing-run/) | 2 October 2026 | bbpPairings and Arxo on presented pairings | 13 hand-built cases plus 20 generated tournaments | ## Catala: the parity run The vehicle-damage rules of the National Bank of Kazakhstan were modelled twice: in Catala and in Arxo. The two models were built independently from the same pinned text. Both returned the expected values on all 65 hand-written cases and agreed on all 240 random inputs. Mutation sets measured how many deliberately broken rules the case bank detects. **What it shows:** on this act, under an agreed contract for money, dates, missing data and negation, the two models compute the same values. The cases detect specific classes of model errors. **What it does not show:** that the two languages are equivalent; anything about Catala's compiled backends; anything about other acts. The timing numbers compare different products of one call: a value against a full evaluation document with a proof. ## LLM / RAG: the US tax code run Two inexpensive models answered synthetic questions on the US Internal Revenue Code (earned income credit, a corporate reorganization, a wash sale, and an earned income credit case with income missing). They answered with no help, with lexical retrieval over a narrow pool of pinned sections, with Arxo calculation tools, and with both. **What it shows:** in this run, the tools mode returned all requested amounts in 6 of 6 question-and-model pairs. Retrieval alone did so in 1 of 6 and no help in 0 of 6. Adding retrieval on top of tools gave 4 of 6. In one case the model skipped the tool and invented an amount. **What it does not show:** a statistical ranking of models or modes. Each cell is one request. The retrieval pool was narrow and chosen in advance. The amount score does not cover the legal correctness of the full answer. ## Science tasks: models with and without tools Three models solved an elliptic-curve chain, a homology certificate and a measurement uncertainty budget, once per task without tools and once with Arxo science tools. **What it shows:** without tools, one model (Luna) solved all three tasks; on the third it noted a table-rounding difference. A second model solved two after correcting its own arithmetic. A third model made method and arithmetic errors. The claim that a plain model fails without tools was **not confirmed** on these tasks. For Luna the tools did not change the final answer. They add an executable check of each claim and the provenance of each number. **What it does not show:** behaviour on larger inputs, longer traces or random parameter sets. It is not a ranking: there is one run per task and mode, and the with-tools controls were recorded earlier under a different context. ## WebAssembly: the i32 spec-tests run Nineteen line-cases from fourteen directives of the official integer test script were checked against the `w3c.wasm_core` package. Every recorded value and trap was returned as written. **What it shows:** on the sampled directives — wrap-around arithmetic, division and remainder traps, comparisons — the package executes the scripted expectations with a proof graph per case. **What it does not show:** validation of modules, unsupported instructions, or the rest of the script. Fourteen directives are a sample, not the specification. ## Computer algebra: the textbook-steps run SymPy 1.14.0 and Wolfram 15.0.1 agree on all 129 bank rows with zero false accepts by Arxo. On fifteen OpenStax cases the references diverge once: on the geometric series without its convergence condition. **What it shows:** Arxo checks presented textbook steps without accepting a wrong one, and refuses outside its catalog with a reason instead of a value. The one divergence sits between a reference and the textbook, not in Arxo. **What it does not show:** the breadth of either algebra system, or an independent certificate from any side. All three answers are verdict records, cross-checked against each other. ## Lean: the divisibility-instances run The Lean side ran on the core library; mathlib statements were not executed. Ten of fourteen instances match; three permission cases are not comparable by design and one case stays unchecked. **What it shows:** applying an imported theorem statement to data, certified by Arxo with provenance to the mathlib pin, meets kernel proofs of the same instances where the two overlap — and both sides refuse honestly where their premises run out. **What it does not show:** a mathlib execution, anything about the unchecked projection case, or either language as a whole. ## FIDE: the Dutch-pairing check run Thirty-three presented pairings were checked against bbpPairings 6.0.0: zero false accepts and zero false refusals on the absolute criteria, with six differences of checking subject. **What it shows:** Arxo accepts exactly the pairings that are clean under the absolute criteria and names the violated criterion with a proof whenever it rejects, while the check mode accepts only its own optimum. **What it does not show:** pairing construction, colour allocation, or anything beyond the Dutch system with standard points. ## Directions without runs The other directions in this section have a prepared protocol (frozen cases, expected outcomes from the source text, reproduction commands) but no executed comparative run. Expected outcomes in those protocols are predictions, not results. The setup is described in the [methodology](/comparisons/methodology/).