Skip to content
docs
Arxo ↗

Measurements

For LLMs8 sections

Seven comparative runs have been executed so far. All other comparisons in this section are protocols that have been prepared but not run. Each run below has its own page with its date, versions, setup, numbers and limits. This page lists what each run shows and what it does not.

RunDateWhat ranSize
Catala: the parity run9 September 2026, extended 22 September 2026Catala 1.2.0 interpreter and Arxo on one act65 cases, 240 random inputs, mutation sets
LLM / RAG: the US tax code run2 October 2026Two models in four modes: no help, retrieval, tools, retrieval plus tools4 main questions plus two side checks; one request per model, question and mode
Science tasks: models with and without tools1 October 2026Three models on three tasks, without tools and with Arxo toolsOne run per task and mode
WebAssembly: the i32 spec-tests run2 October 2026Arxo against the official integer test script19 line-cases from 14 directives
Computer algebra: the textbook-steps run2 October 2026SymPy, Wolfram and Arxo on a bank and textbook cases129 bank rows plus 15 OpenStax cases
Lean: the divisibility-instances run2 October 2026Lean core and Arxo on divisibility instances14 cases
FIDE: the Dutch-pairing check run2 October 2026bbpPairings and Arxo on presented pairings13 hand-built cases plus 20 generated tournaments

The vehicle-damage rules of the National Bank of Kazakhstan were modelled twice: in Catala and in Arxo. The two models were built independently from the same pinned text. Both returned the expected values on all 65 hand-written cases and agreed on all 240 random inputs. Mutation sets measured how many deliberately broken rules the case bank detects.

What it shows: on this act, under an agreed contract for money, dates, missing data and negation, the two models compute the same values. The cases detect specific classes of model errors.

What it does not show: that the two languages are equivalent; anything about Catala’s compiled backends; anything about other acts. The timing numbers compare different products of one call: a value against a full evaluation document with a proof.

Two inexpensive models answered synthetic questions on the US Internal Revenue Code (earned income credit, a corporate reorganization, a wash sale, and an earned income credit case with income missing). They answered with no help, with lexical retrieval over a narrow pool of pinned sections, with Arxo calculation tools, and with both.

What it shows: in this run, the tools mode returned all requested amounts in 6 of 6 question-and-model pairs. Retrieval alone did so in 1 of 6 and no help in 0 of 6. Adding retrieval on top of tools gave 4 of 6. In one case the model skipped the tool and invented an amount.

What it does not show: a statistical ranking of models or modes. Each cell is one request. The retrieval pool was narrow and chosen in advance. The amount score does not cover the legal correctness of the full answer.

Science tasks: models with and without tools

Section titled “Science tasks: models with and without tools”

Three models solved an elliptic-curve chain, a homology certificate and a measurement uncertainty budget, once per task without tools and once with Arxo science tools.

What it shows: without tools, one model (Luna) solved all three tasks; on the third it noted a table-rounding difference. A second model solved two after correcting its own arithmetic. A third model made method and arithmetic errors. The claim that a plain model fails without tools was not confirmed on these tasks. For Luna the tools did not change the final answer. They add an executable check of each claim and the provenance of each number.

What it does not show: behaviour on larger inputs, longer traces or random parameter sets. It is not a ranking: there is one run per task and mode, and the with-tools controls were recorded earlier under a different context.

Nineteen line-cases from fourteen directives of the official integer test script were checked against the w3c.wasm_core package. Every recorded value and trap was returned as written.

What it shows: on the sampled directives — wrap-around arithmetic, division and remainder traps, comparisons — the package executes the scripted expectations with a proof graph per case.

What it does not show: validation of modules, unsupported instructions, or the rest of the script. Fourteen directives are a sample, not the specification.

SymPy 1.14.0 and Wolfram 15.0.1 agree on all 129 bank rows with zero false accepts by Arxo. On fifteen OpenStax cases the references diverge once: on the geometric series without its convergence condition.

What it shows: Arxo checks presented textbook steps without accepting a wrong one, and refuses outside its catalog with a reason instead of a value. The one divergence sits between a reference and the textbook, not in Arxo.

What it does not show: the breadth of either algebra system, or an independent certificate from any side. All three answers are verdict records, cross-checked against each other.

The Lean side ran on the core library; mathlib statements were not executed. Ten of fourteen instances match; three permission cases are not comparable by design and one case stays unchecked.

What it shows: applying an imported theorem statement to data, certified by Arxo with provenance to the mathlib pin, meets kernel proofs of the same instances where the two overlap — and both sides refuse honestly where their premises run out.

What it does not show: a mathlib execution, anything about the unchecked projection case, or either language as a whole.

Thirty-three presented pairings were checked against bbpPairings 6.0.0: zero false accepts and zero false refusals on the absolute criteria, with six differences of checking subject.

What it shows: Arxo accepts exactly the pairings that are clean under the absolute criteria and names the violated criterion with a proof whenever it rejects, while the check mode accepts only its own optimum.

What it does not show: pairing construction, colour allocation, or anything beyond the Dutch system with standard points.

The other directions in this section have a prepared protocol (frozen cases, expected outcomes from the source text, reproduction commands) but no executed comparative run. Expected outcomes in those protocols are predictions, not results. The setup is described in the methodology.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.