Markdown for LLMs
LLM / RAG
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# LLM / RAG **In short:** This direction does not contrast Arxo with an abstract machine mind — Arxo itself assumes models write formalizations. It compares three fixed architectures on one held-out case bank with one model and one knowledge slice: the model with the act text, the model with retrieval over the act, and the model with an executable model behind tools. This page is for teams deciding where a language model belongs in a legal-answering pipeline. ## What the three configurations are for - **Model with text** answers freely phrased questions with zero formalization to start — fast to build, grounded in nothing but context. - **Model with retrieval** grounds wording in source chunks and reduces invented citations, at the cost of indexing and chunking choices. - **Model with an executable model behind tools** reaches computed answers without retelling: the model calls the engine through a tool transport instead of paraphrasing the act. Retrieval-grounding and tool-use literature, evaluation methods for generated answers, and the tool-transport specification back the setup; answer-quality scores from the literature never serve as acceptance here — the pinned act text plus an independent annotator does. ## Where it meets Arxo All three answer the same held-out tariff cases, judged against the pinned act text on signs, source links, refusal behavior on thin data, and cost. The Arxo side is an adapter covering readiness, refusals, and edition windows; the tariff arithmetic stays in the pinned subject package rather than being duplicated. The run ban is load-bearing: the same model and the same knowledge slice everywhere, so the engine is never credited for what a model or source change produced. ## Key differences - **Proof versus quote.** Text and retrieval configurations cite string quotes or chunk identifiers — unverifiable as proof. The tool configuration answers with the article plus the derivation graph. The comparison is a judged quality pair, never byte equality. - **Thin data.** The first two configurations are expected to refuse or ask for clarification — a silent answer is scored as a defect. The engine side refuses with a named outcome. Retrieval failures and tool errors share an outcome class with different cause codes. - **Revisions are re-runs, not re-reads.** A reform phase re-runs the same stand: a rewritten article, a changed chunking setup, new monetary parameters, a relocated rule with a new package version. Model and index versions are fixed at freeze time, with fingerprints reported. - **Stochasticity and cost are measured, not assumed.** Fixed temperature and seed, five repeats, all repeats reported; cost splits into preparation and indexing, inference tokens per case, latency percentiles, and revision maintenance. ## A concrete scenario The prepared experiment sets twelve tariff cases — tables, deadline coefficients, bonus-malus, an ambiguous city wording kept open, missing region and class, pre-window dates, off-tariff questions, paraphrase and dirty-input probes — plus four revision cases. Expectations come from the act text and the recorded bank; open readings stay open. The published 93-case bank is not held out, so leakage mitigation (post-cutoff paraphrases, a closed bank, slice dates) is part of the protocol, not an afterthought. ## Choosing and combining Choose model-with-text for exploration and zero-setup triage; model with retrieval when answers must show their source passages and the question load is free-form; model with an executable model when the answer must compute, refuse cleanly, and replay. Arxo is not the opponent of any of these — it is the third configuration's engine, and the author of the models the other two paraphrase. The honest question is never "model or rules" but which architecture owns which question. ## Evidence and open questions - Sources checked: September 2026 (retrieval and evaluation literature, tool-use docs, the transport specification slice; case hashes recomputed). - Studied profile: no model, embedding, chunking, or engine pin chosen — all fixed at freeze and run time, with fingerprints. - Basis: confirmed by documentation plus a prepared protocol; comparative run not performed. - Open: the frozen revision bytes for one reform case, several negative scenarios that are honest package debt, and the run itself. ## Sources and reproducible materials - Companion page: [LLM / RAG: the US tax code run](/comparisons/llm-rag-irc-run/) — an executed run on US tax questions, separate from the prepared protocol. - Retrieval grounding: [arxiv.org/abs/2005.11401](https://arxiv.org/abs/2005.11401) - Answer evaluation methods: [arxiv.org/abs/2309.15217](https://arxiv.org/abs/2309.15217) - Leakage-aware construction: [arxiv.org/abs/2412.13670](https://arxiv.org/abs/2412.13670)