Skip to content
docs
Arxo ↗

Show grounds: proof, sources, and hashes

For LLMs10 sections
  • Goal: show the user why the answer is what it is — the rule that fired, the source anchors behind it, and the hashes that pin the computation — each presented for exactly what it warrants, no more, and list the checks actually performed, not the checks the format would allow.

  • You need: Calculate and check: the proof and hashes below come from the COMPUTED answers built there.

  • Run:

    Terminal
    node --import tsx --test test/evaluate.test.ts
  • Files:

    Output
    src/to-view.ts grounds section of the view model
    src/runtime.ts (answers already carry proof/sources/hashes)

The answer document carries its own grounds; the application does not compute them. On the example truth call (event 2026-03-06, 14 days, candidate 2026-03-20) the grounds section of the view model reads:

Output
rule: TagesfristEnde (urn:de:corpus:clir:bgb-fristen#TagesfristEnde)
sources: No sources anchored in this canon build
program: sha256:468e3fe1… semantic: sha256:90e53b45… result: sha256:4c543151…
via: local document: 4691 bytes

Four things sit next to the answer: the rule that fired, the source list, the three hashes, and the checks that ran on this answer. The next four sections take them in that order. The full answer document with untruncated hashes is in Verify the numbers at the bottom.

TagesfristEnde is the rule that counts a period in days; its full identifier is urn:de:corpus:clir:bgb-fristen#TagesfristEnde. The view shows that rule identifier as the address of the reasoning — it is stable across runs and versions, so a reviewer can look it up. It is not a link into the source text: this build attaches no source anchors, and the rule id does not become one.

The sources list is empty, and the application says so plainly: “this canon build attaches no source anchors to the answer”. An empty list is information, not a failure — but hiding it would imply anchors that do not exist.

Each hash pins a different thing, and the view labels them separately:

HashPinsChanges when
programthe canon build that answeredthe canon version changes
semanticthe evaluated case with its contextfacts, legal time, timezone, or policy change — not the question kind
resultthe outcome of this querythe outcome changes

Read it as: semantic answers “same case?”, result answers “same outcome from this SDK?”, program answers “same build?”. None of them answers “same document bytes?” — that is the checksum’s job in chapter 7. The document bytes do not hash to result: result is the engine’s outcome hash, while the checksum over the stored bytes is a separate value the application computes itself (chapter 7 names it documentSha256).

The middle row is the one newcomers misread. Collect and truth on the same case share semantic while their result hashes differ: the question kind is not covered. Change the facts, the legal time, the timezone, or the policy, and semantic moves.

Seven trust properties meet in one answer. The view shows each as the status of a check actually performed — never as potential checkability, and never collapsed into a single badge:

Output
Proof graph: present (the TagesfristEnde application the view reads)
Proof checker: not run — the app displays the graph, it does not re-verify it
Document checksum: passed over the stored bytes (input and metadata not covered)
Replay: passed under the recorded environment (see chapter 7) — after Replay ran on this capture; before that, not run
Producer authentication: not configured — 'via: local' names the execution
path inside the document; it does not authenticate who wrote the file
Source anchors: absent in this build
Applicable law: legalTime 2026-09-17, Europe/Berlin, BGB_FRISTEN_TAG —
outside those assumptions the answer says nothing

The replay line is the state after Replay ran on that capture; before any replay it reads not run, and the application never renders a static “passed” — the replay outcome appears only in the Replay report for the capture it ran on.

A single “verified” badge would fuse all seven into one glow. The view shows seven short lines instead, each checkable on its own. Every line belongs to one capture: new input recomputes, stores a new capture, and starts un-replayed — a passed replay never carries over to a new calculation.

One row per change, each probed on the worked case:

Changeprogramsemanticresult
collect → truth, same casesamesamediffers
a fact changedsamemovesrecomputed — compare, never assume
legal time or timezone changedsamemovesrecomputed — compare, never assume
policy TAG → KALENDERsamemovesrecomputed — the verdict itself may flip
another model versionmovesdo not comparedo not compare — replay refuses the mix by spec
another SDK, same casesamesamediffers — each SDK replays its own captures

The policy row is why the app pins one policy per case: the new semantic value is shared by that policy’s collect and truth alike.

The last row is the cross-language comparison. The same case through the Python SDK yields the same program and semantic hashes but a different result hash and different document bytes: same statuses, same values, same case hash — different outcome bytes. Replay therefore stays within one SDK (chapter 7).

  • Proof nodes are present on COMPUTED answers; on other statuses the graph may be partial or absent, and the view must say that.
  • Hash equality across SDKs does not hold: same statuses, same values — different bytes and different result hashes.
  • The canon manifest (questions(), passport()) describes what the model can answer; it is inventory, not grounds, and belongs in the form builder, not in the answer view.
  • via: local describes the execution path recorded in the answer. It says nothing about who produced the file holding the answer or whether that file changed afterwards.

The example truth answer, with the fields the view reads:

JSON
{
"evaluationStatus": "COMPUTED",
"truthStatus": "TRUE_ONLY",
"rulesApplied": ["TagesfristEnde"],
"sources": [],
"hashes": {
"program": "sha256:468e3fe17c6a28371c6ce8258ae7928c496593673d4b73a52c2a9c40cb50ee33",
"semantic": "sha256:90e53b453107dfee9422a5fe24774823430b5c82c3577b9321a0838a6e54c695",
"result": "sha256:4c5431514533d2afe380ec92a6427831120fdbeca7315f731b5e374c490f759a"
},
"via": "local",
"documentBytes": 4691
}

The probe values behind the change table, on the worked case:

Output
collect vs truth, same case ..... semantic 90e53b45… on both; result ae9a7b19… (collect) vs 4c543151… (truth)
facts: event 2026-05-01, 30 days semantic af79c78f…
legal time 2026-01-01 only ...... semantic 26d9fd0a…
timezone UTC only ............... semantic 63a056fb…
policy BGB_FRISTEN_KALENDER ..... semantic ac8223da… (collect and truth alike)
Python SDK, collect ............. same program and semantic; result 1f0abeda…; document 4884 vs 4820 bytes

Run the grounds test and confirm rule, empty sources, and hashes:

Terminal
node --import tsx --test test/evaluate.test.ts
Output
ok 3 - the example answer carries its grounds
# tests 13
# pass 13
# fail 0

Predict before you open each answer.

You ask collect and then truth on the same case. Which of the three hashes match?

program and semantic match: same build, same case. result differs (ae9a7b19… for collect, 4c543151… for truth), because the question kind is part of the outcome, not of the case.

You change only the legal time to 2026-01-01. Which hash moves?

semantic moves (to 26d9fd0a…): the legal time is part of the case context. program stays, because the canon build is the same. The result is recomputed, so compare it rather than assume it.

The answer has `sources: []`. Is it ungrounded?

No. This canon build attaches no source anchors, and the application says so plainly. The grounds it can show are the applied rule identifier, which is the address of the reasoning, and the three hashes.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.