docs← Back to article

Markdown for LLMs

LLM / RAG

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# LLM / RAG

**In short:** This direction does not contrast Arxo with an abstract
machine mind — Arxo itself assumes models write formalizations. It
compares three fixed architectures on one held-out case bank with one
model and one knowledge slice: the model with the act text, the model
with retrieval over the act, and the model with an executable model behind
tools. This page is for teams deciding where a language model belongs in a
legal-answering pipeline.

## What the three configurations are for

- **Model with text** answers freely phrased questions with zero
  formalization to start — fast to build, grounded in nothing but context.
- **Model with retrieval** grounds wording in source chunks and reduces
  invented citations, at the cost of indexing and chunking choices.
- **Model with an executable model behind tools** reaches computed answers
  without retelling: the model calls the engine through a tool transport
  instead of paraphrasing the act.

Retrieval-grounding and tool-use literature, evaluation methods for
generated answers, and the tool-transport specification back the setup;
answer-quality scores from the literature never serve as acceptance here —
the pinned act text plus an independent annotator does.

## Where it meets Arxo

All three answer the same held-out tariff cases, judged against the pinned
act text on signs, source links, refusal behavior on thin data, and cost.
The Arxo side is an adapter covering readiness, refusals, and edition
windows; the tariff arithmetic stays in the pinned subject package rather
than being duplicated. The run ban is load-bearing: the same model and the
same knowledge slice everywhere, so the engine is never credited for what
a model or source change produced.

## Key differences

- **Proof versus quote.** Text and retrieval configurations cite string
  quotes or chunk identifiers — unverifiable as proof. The tool
  configuration answers with the article plus the derivation graph. The
  comparison is a judged quality pair, never byte equality.
- **Thin data.** The first two configurations are expected to refuse or
  ask for clarification — a silent answer is scored as a defect. The
  engine side refuses with a named outcome. Retrieval failures and tool
  errors share an outcome class with different cause codes.
- **Revisions are re-runs, not re-reads.** A reform phase re-runs the same
  stand: a rewritten article, a changed chunking setup, new monetary
  parameters, a relocated rule with a new package version. Model and index
  versions are fixed at freeze time, with fingerprints reported.
- **Stochasticity and cost are measured, not assumed.** Fixed temperature
  and seed, five repeats, all repeats reported; cost splits into
  preparation and indexing, inference tokens per case, latency
  percentiles, and revision maintenance.

## A concrete scenario

The prepared experiment sets twelve tariff cases — tables, deadline
coefficients, bonus-malus, an ambiguous city wording kept open, missing
region and class, pre-window dates, off-tariff questions, paraphrase and
dirty-input probes — plus four revision cases. Expectations come from the
act text and the recorded bank; open readings stay open. The published
93-case bank is not held out, so leakage mitigation (post-cutoff
paraphrases, a closed bank, slice dates) is part of the protocol, not an
afterthought.

## Choosing and combining

Choose model-with-text for exploration and zero-setup triage; model with
retrieval when answers must show their source passages and the question
load is free-form; model with an executable model when the answer must
compute, refuse cleanly, and replay. Arxo is not the opponent of any of
these — it is the third configuration's engine, and the author of the
models the other two paraphrase. The honest question is never "model or
rules" but which architecture owns which question.

## Evidence and open questions

- Sources checked: September 2026 (retrieval and evaluation literature,
  tool-use docs, the transport specification slice; case hashes
  recomputed).
- Studied profile: no model, embedding, chunking, or engine pin chosen —
  all fixed at freeze and run time, with fingerprints.
- Basis: confirmed by documentation plus a prepared protocol; comparative
  run not performed.
- Open: the frozen revision bytes for one reform case, several negative
  scenarios that are honest package debt, and the run itself.

## Sources and reproducible materials

- Companion page: [LLM / RAG: the US tax code run](/comparisons/llm-rag-irc-run/) — an executed run on US tax questions, separate from the prepared protocol.
- Retrieval grounding: [arxiv.org/abs/2005.11401](https://arxiv.org/abs/2005.11401)
- Answer evaluation methods: [arxiv.org/abs/2309.15217](https://arxiv.org/abs/2309.15217)
- Leakage-aware construction: [arxiv.org/abs/2412.13670](https://arxiv.org/abs/2412.13670)