LLM / RAG: the US tax code run
Result: on 2 October 2026, two inexpensive models answered synthetic questions on the US Internal Revenue Code in four modes. All requested amounts matched the reference in:
- 0 of 6 question-and-model pairs with no help;
- 1 of 6 with retrieval;
- 6 of 6 with calculation tools;
- 4 of 6 with retrieval plus tools.
On a question with income deliberately missing, every mode refused to invent the figure in 2 of 2, except retrieval plus tools, at 1 of 2. This is one request per model, question and mode. It is not a ranking.
This run is separate from the prepared protocol on the LLM / RAG page. It uses different models, a different case bank and one request per cell instead of repeated runs.
Models and modes
Section titled “Models and modes”The models were liquid/lfm-2.5-2.6b:free and mistralai/mistral-nemo.
Each request ran at temperature 0, with low reasoning requested and a limit
of 6,000 output tokens. Providers may treat the reasoning setting
differently: the Liquid model sometimes spent the whole limit on reasoning
and gave no answer.
| Mode | What the model received |
|---|---|
| No help | The question and a shared system prompt |
| Retrieval | The same, plus the top 8 passages from lexical retrieval |
| Tools | The existing application adapter: source selection, a server-side check of input facts, and calculation by the Arxo tax model through tool calls |
| Retrieval plus tools | The tools adapter plus the same top 8 passages as in retrieval mode |
The difference between no help and tools also includes the adapter’s prompt and its deterministic fact preparation. The comparison is between finished application modes. It does not isolate the effect of tool calls alone.
The retrieval pool
Section titled “The retrieval pool”Retrieval was local lexical ranking, close to BM25, returning the top 8 passages. The pool was limited in advance to sections 32, 356, 358, 382, 1091, 6081 and 6151 of the Internal Revenue Code, Rev. Proc. 2024-40 section 3.06, and Rev. Rul. 2026-19 Table 3. This is a favourable, narrow pool. It does not test retrieval across US legislation.
Each subsection was its own passage. The annual table and the monthly Table 3 were kept whole. Section numbers in a question gave a ranking bonus. The index held no rules, no engine results, no expected amounts and no earlier model answers.
The statute text is the pinned edition USCODE-2024-title26, with a source currency date of 6 January 2025. The run does not establish that this text is current for 2025 or 2026. The dates, the annual table and the rate are conditions of this experiment.
Questions and reference answers
Section titled “Questions and reference answers”| Question | Reference |
|---|---|
| Earned income credit, where adjusted gross income ($42,500) is above earned income ($19,000) | $4,618.48 |
| Corporate reorganization with gain less than cash received | Section 356: $180,000; section 358: $360,000; section 382: $108,625 |
| Wash sale with the replacement bought exactly 30 days before the sale | Loss $2,500; replacement basis $10,820.75 |
| Earned income credit with adjusted gross income missing | Ask for the income; do not compute an amount |
The earned income credit reference is the engine’s formula with rounding to cents. It is not the amount from the Form 1040 lookup table.
Scoring
Section titled “Scoring”An amount scores when every requested amount matches the checked engine result. The corporate question needs all three amounts. The missing-income question scores when the model does not invent the income. The amount score does not cover the legal correctness of the full answer or its explanation. Provider errors and missing final answers count in the denominator, because the run measures whether a result reached the user.
Results
Section titled “Results”| Mode | Liquid, amounts | Nemo, amounts | Total | Income not invented |
|---|---|---|---|---|
| No help | 0/3 | 0/3 | 0/6 | 2/2 |
| Retrieval | 0/3 | 1/3 | 1/6 | 2/2 |
| Tools | 3/3 | 3/3 | 6/6 | 2/2 |
| Retrieval plus tools | 2/3 | 2/3 | 4/6 | 1/2 |
What retrieval changed
Section titled “What retrieval changed”- Earned income credit. Nemo with no help gave $2,245.71. With retrieval it found the correct maximum credit ($7,152) and threshold ($30,470), but chose a 7.65 % phase-out rate instead of 21.06 % and gave $6,232.45. The source was found and applied wrongly. Liquid with retrieval used all 6,000 tokens on reasoning and gave no final answer.
- Wash sale. With no help, both models left the replacement basis at $8,320.75. Nemo with retrieval applied section 1091(d) and reached $10,820.75. Liquid found the right formula but made a subtraction error and gave $11,820.75. It also gave the wrong start of the 30-day window.
- Corporate question. The top 8 passages missed sections 382(b) and 382(f) and the monthly rate. This retrieval gap was kept as it was. Liquid noted that Table 3 was missing but still gave wrong section 358 and section 382 results. Nemo gave an invented section 382 amount of $1,100,000.
Retrieval plus tools
Section titled “Retrieval plus tools”- Nemo, on the missing-income question, did not call the tool. It chose the maximum credit for three children instead of two and gave $7,965.41, although the income was explicitly missing. In tools mode the same question ended with a request for the income. Having a tool available does not make the model use it.
- On the corporate question, the Liquid request ended with a provider error. Nemo stated that it would call a function, but made no call and gave no amounts.
Explanations around correct amounts
Section titled “Explanations around correct amounts”In tools mode Liquid gave the correct corporate amounts, but described the role of the 54.5 percentage points inaccurately. In retrieval plus tools, Liquid explained the wash-sale arithmetic correctly. It first described a later purchase and then an earlier one. On the missing-income question it confused a passage from the user’s question with the statute.
Side checks
Section titled “Side checks”Source coverage control. Both models were given sections 356(a), 358(a), 382(a), (b), (f) and (g) and Table 3 directly, chosen by hand rather than retrieved. Liquid then gave $270,000 for section 356 and $2,660,000 for section 358. It read the rate and the section 382 amount of $108,625 correctly. Nemo identified the boot gain correctly, but gave a basis of $270,000 and $108,125 for section 382. Better retrieval alone did not close the gap. This control is not part of the four-mode table.
A rule question without arithmetic. Does an ordinary extension of time to file (section 6081) extend the time to pay (section 6151)? Both models, with no help and with retrieval, answered “no”. The reasoning differed:
- Liquid with no help invented a quotation of section 6151.
- With retrieval, Liquid quoted real passages but added unsupported material.
- Nemo with retrieval quoted section 6151(a), but wrongly offered estimated tax payments as an alternative to an extension of payment.
The full four-mode test was not run for this question.
Availability and cost
Section titled “Availability and cost”On the first launch, all 16 tools and retrieval-plus-tools requests failed locally with the tool server unavailable (HTTP 502), before reaching the model. After the tool server was started, only those 16 requests were repeated. The original failure is kept on record and is not counted as a reasoning error. No successful model answer was rerun.
In total, 54 observations were saved:
- 32 from the first launch;
- 16 after the tool server was restored;
- 4 for the payment question;
- 2 for the source control.
The provider-reported cost was $0.00260702. The Liquid model reported a cost of zero. This is the API cost of this run only. Local computation and hosting are not included.
Limits
Section titled “Limits”- Small sample, one run, no repeats, no independent larger bank.
- Two specific models through one provider; other versions may behave differently.
- The fact parser and the calculation scope were narrow by design.
- Source-linked input means a fact matches the user’s text. It is not an independent check that the fact is true.
- A computed status confirms that the calculation ran. It does not confirm the model’s free text around it.
- The retrieval pool was narrow and favourable, and the statute text has a fixed currency date.
- The production model and interface were not changed for the run.
In this run, retrieval helped find and show the rule, and calculation tools helped apply formulas and arithmetic. Adding eight long passages to the tools mode did not improve the result automatically.
Documentation for Arxo. Writings — blog.arxo.io.
Anonymous visit counts on stats.arxo.io, no cookies.