Markdown for LLMs
Catala: the parity run
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Catala: the parity run **Result:** the computational part of one act was modelled independently in Catala 1.2.0 and in Arxo. On 65 hand-written cases the Catala model returned the expected value in 65 of 65, and so did Arxo. On 240 random inputs the two agreed in 240 of 240. All 20 hand-written mutations of the Arxo model were detected. The run dates from 9 September 2026 and was extended for the public artifact on 22 September 2026. Its numbers belong to Catala 1.2.0 with the interpreter and the standard runtime, on this one act. For the comparison of the two languages as a whole, see [Catala](/comparisons/catala/). For a concept-by-concept mapping, see [Coming from Catala](/language/coming-from-catala/). ## The act and the two models The act is Resolution No. 14 of the Board of the National Bank of the Republic of Kazakhstan of 28 January 2016, "Rules for Determining the Amount of Damage Caused to a Vehicle". Both sides use the same pinned copy of the text. The Catala model was derived from the text without reusing the Arxo model. It covers six areas: total loss, the value of a part, the choice of appraisers, the admissibility of an inspection, deadlines, and the final amount. The 2026 working-day calendar enters Arxo as a pinned dataset with a hash. In Catala it is a structure passed as input, and working days are written by hand as a fold over candidate dates. Before any run, a written contract fixed the semantics of money, dates, missing data and negation for both sides. External facts that the act delegates to other tools, such as the amount of wear, enter both models as case facts. ## What "agreement" means here The two sides return different products from one call. Catala returns a value. Arxo returns an evaluation document: the manifest, the results, the proof graph, the deontic positions, conflicts, issues and a result hash. Agreement is therefore defined on a projection. A case agrees when the value projected from Arxo's answer equals the value mapped from Catala's output for the same question. The mapping from Arxo predicates to Catala fields is published with the artifact. A round trip through that adapter gave 65 of 65 and 240 of 240. Three different comparisons appear in the run: - **Against hand-set expectations.** The expected answers for the 65 cases were established manually from the act text and the calendar, with a rationale per case. Each model is checked against them. - **Between the two models on random inputs.** Random cases for total loss, part value, choice of appraisers and deadlines on the 2026 calendar were generated with seeds 1 and 2, 120 each. - **Between Arxo's two implementations.** The Rust engine and the Python reference implementation compare the full evaluation document byte for byte, not only the value. This is internal to Arxo and is not a comparison with Catala. Comparison is limited to whole-tenge amounts. Catala rounds a money-by-decimal product to the nearest hundredth in its runtime. Arxo computes exactly. The act requires the final amount in whole tenge. This difference was written into the contract before the run and is pinned by a dedicated case. ## Results | Check | Catala 1.2.0 | Arxo | |---|---|---| | 65 cases against hand-set expectations | 65 / 65 | 65 / 65 | | 240 random inputs (seeds 1 and 2), Catala against the Python implementation | 240 / 240 | — | | 240 random inputs, Rust engine against the Python implementation | — | 240 / 240 | | Adapter round trip | 65 / 65 and 240 / 240 | | | Alternative Catala encoding (one exception with a disjunction of two grounds) | 65 / 65; 120 / 120 on each seed | — | | Hand-written mutations of the Arxo model | — | 20 / 20 detected, none rejected by the compiler | No discrepancies between the two models were found on either the hand-set or the random cases. In 610 Catala runs (305 per encoding) there was no conflict error. The independent proof checker written in Lean also received the evaluation documents of the 65 cases: 219 documents (118 without the calendar, 101 with it). All 219 were accepted under its partial-verification policy. Money arithmetic in these documents is admitted by the checker, not proved. ## Mutation results Mutations measure the strength of the case bank, not of either language. A mutation changes one rule of the Arxo model. It counts as detected when at least one case gives a different answer. **Hand-written mutations (20).** These target specific losses: the 80 % and 70 % thresholds, the 15,000 and 20,000 km boundaries, deduction of salvage, 5 to 6 working days, working to calendar days, a strict to non-strict comparison at a limit, a single appraiser, the rounding check, and others. All 20 were detected. The 65 parity cases alone detected 16 of the 20. **Systematic mutations (420).** Operators changed comparisons, dropped or negated conditions, changed constants and units, and removed exceptions across all 13 modules of the package. | Outcome | Count | |---|---:| | Detected | 133 | | Survived | 97 | | Rejected by the compiler's static checks | 190 | Of the 97 survivors, 84 are behaviour that the scenario set does not exercise. Four are equivalent mutations: a strict rule makes the mutated one irrelevant. The rest are structural: linking conditions and elements that the scenarios assert together. **Comparable core (153).** A separate matrix mutated only the 26 rules that feed the 14 predicates compared with Catala. Of 153 mutants, 97 compiled and 56 were rejected. | Case set | Mutants detected (of 97) | |---|---:| | 65 parity cases | 76 | | 240 random inputs | 59 | | 104 additional authored scenarios | 70 | | All 169 scenarios | 78 | | Scenarios plus random inputs | 84 | Four of the 97 are equivalent. Random inputs detected 6 of the 11 non-equivalent mutants that all 169 scenarios missed. Nine survived every set. All nine are structural: three linking conditions, and six conditions of one clause that the scenarios assert as a block. ## Timing Measured on 9 September 2026 on Apple Silicon under macOS, warm runs. These numbers do not compare languages. The three implementations produce different products: Catala a value, Arxo a canonical evaluation document with a proof graph and hashes. | Measurement | Catala 1.2.0, interpreter | Arxo, Rust engine | Arxo, Python implementation | |---|---|---|---| | One calculation inside a process | about 16 µs | about 0.4 to 0.5 ms | about 2.5 ms | | Process start and model load | about 40 ms | about 10 ms | about 0.3 s | How the numbers were taken: - Catala: 10,000 calls of one scope folded over a list. The time of the same file with a single call was subtracted. - Arxo: the scenario runner over the case file. The time of a one-scenario file was subtracted. - Catala's build directory must sit inside the project. With an external build directory the standard library is rechecked on every start, at 2.0 to 3.7 s instead of 0.04 s. - The command-line wrapper adds about 0.7 s per call. Measurements call the engine binary directly. - Catala's compiled backends were not measured. A Python module was generated, but there was no runtime to execute it on the test machine. ## Limits - One act and its computational part. The run does not cover everything Catala can do, nor all of Arxo. - Forms outside the comparable model were not part of the comparison: ten duties, one liberty, the defeasible layer, established negations in the proof, the document lists of the annexes. They are checked by Arxo's own scenarios. - Contradictory input cannot be posed to a Catala variable, which holds one value. In one scenario the Catala model needs an exclusivity guard where Arxo keeps two supports. - The expected answers for the 65 cases were set by the author from the text. There was no external legal review of the cases and no review of the Catala model by an experienced Catala user. - Whole-tenge amounts only. ## Reproducing the run The public artifact is [github.com/arxohq/arxo-catala-parity](https://github.com/arxohq/arxo-catala-parity), Apache-2.0. It contains: - the Arxo package and the pinned calendar bytes; - the Catala model and an English description of the agreed semantics; - the 65 cases and the stand-alone comparison and mutation scripts, which use only the Python standard library; - the recorded reports and an English report with the qualification of every observation. The Catala side reproduces from the repository with Catala 1.2.0 installed through opam. The Arxo results are recorded in the repository, with the SHA-256 of the engine binary that produced them. Revisions of the accompanying paper are tagged in the repository.