Markdown for LLMs
Collect facts and evidence
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Collect facts and evidence
Rules work on typed facts. Users bring stories and paperwork. This page covers the agent's middle job: turning "she gets a 25-minute break and has one baby" or an uploaded bank confirmation into the exact facts a question needs. Each fact stays traceable to the words it came from. The agent accepts "I don't know" and never asks the user for what the engine should derive.
## One document, start to finish
An employee asks whether purchase P42 is reimbursable. The approval order A17 and the receipt R25 are already in the case. The employee now uploads a bank confirmation, B91, to show that the purchase was paid. The teaching package `demo.expense-evidence` (an invented reimbursement rule, not legislation) runs an active evidence policy, so every step is visible in its answers.
1. **The user attaches a document.** B91 is genuine and verified.
2. **The agent extracts a claim from it:** `paid(P42)`, meaning the purchase was paid.
3. **The claim is linked to the document.** The case now holds a support link: B91 supports `paid(P42)`.
4. **The policy checks the link.** B91 confirms a different purchase, P99, so the policy rejects it as support for `paid(P42)` (`NO_ACCEPTANCE_RULE_APPLIED`).
5. **The answer changes accordingly.** `reimbursable(P42, EMP7)` comes back `NEITHER`.
`NEITHER` means "not established", not "unpaid". **Rejecting one support does not prove the opposite claim**: the case simply has no admitted ground for saying P42 was paid.
When the employee uploads the right confirmation, B92, the policy admits it and the same question returns `TRUE_ONLY`. The rejected B91 link stays in the record; the finding rests on B92 alone.
Status: ran locally on the reference engine with the release pinned to the package's own `0.2.4` (scenes 1 and 2 of the [admission table](#admission-when-a-support-counts); see the validation record).
## From story to typed facts
The labour-law example starts from a story: an employee with one child under eighteen months gets a 25-minute feeding break. The rules behind `feeding_break_too_short` (package `kz-labour-code`, Labour Code of the Republic of Kazakhstan) need two typed facts:
```json
// the elicited fact set — ran via MCP law_ask, kind truth, legal time 2026-09-01
[
{ "predicate": "children_under_eighteen_months",
"args": ["urn:kz:tk:employee:1", 1] },
{ "predicate": "child_feeding_break_minutes",
"args": ["urn:kz:tk:employee:1", "urn:kz:tk:employer:1", 25] }
]
```
Status: ran via MCP `law_ask` on the repo server (jurisdiction Republic of Kazakhstan as reported by the server; no profile name is exposed). The captured answer establishes the too-short finding on this set.
Each conversion is a small decision the agent makes visible:
| Story fragment | Typed fact | Decision |
|---|---|---|
| "one baby" | count `1` on `children_under_eighteen_months` | The baby must be under eighteen months; the agent confirms this instead of assuming it |
| "25-minute break" | minutes `25` on `child_feeding_break_minutes` | Integer minutes, per the signature |
| "she", "her employer" | `urn:kz:tk:employee:1`, `urn:kz:tk:employer:1` | Stable identities reused across facts |
### Types come from the canon, and errors come before evaluation
The signature decides the type, not the agent: the count and the minutes are integers, so "about half an hour" needs a clarification first. `law_rules` with `contract: true` returns a machine input contract over the same dependency world as `law_ask`; read required inputs from it rather than from rule text by hand.
On the local SDK route (`@arxo/law`), the package declarations type each bare value: a date as `'2026-03-06'`, a decimal as a string, money as `'730000 KZT'`, an enum as its variant name, an entity as a short id or a full URN. An unknown relation, a wrong number of arguments, a value of the wrong type or an unknown enum variant throws `FactError` **before the engine runs**. The error carries the failing path and the nearest declared names, so a wrong fact never becomes a confident answer.
Status: documented in the `@arxo/law` README and source; observed for a misspelled relation in [Build your first agent](/agent-engineering/first-agent/) (SDK run, 2026-10-04, canon `de.bgb.fristen@0.1.0`).
A `FactError` is a failure of your input, not a legal answer. Report it as "this fact does not fit the canon", fix the name or value from the `nearest` list or ask the user, and never retry by guessing a different relation.
### Object identity: the same thing, the same name
Both facts name the same employee with the same identifier, which lets the rules join them. With a fresh identifier per fact, the rules would see two strangers and stay silent. Mint one identifier per real-world object and reuse it everywhere; never share one between two objects (keep the employer distinct even in a one-person story); record the mapping so a reviewer can check it.
Short labels are fine inside one case if used consistently: the SDK expands a short id such as `'frist'` into a namespaced entity of the package. Across cases and acts, full URNs keep objects apart.
## Source facts versus derived conditions
This is the most important line in elicitation: the agent collects source facts and never asks the user for derived conditions. In the example, the two input relations are source facts that the case supplies. `feeding_break_too_short` is a derived condition that the rules produce. The user is never asked "is the break too short?", because that is the question the engine answers.
If a predicate appears in rule heads, the agent does not ask the user for it. If it appears only in rule bodies with no producing rule, the agent asks for it or leaves it explicitly unknown. When in doubt, re-read the rules or the input contract instead of asking the user for a computed result.
Status: ran via MCP `law_rules` on the repo server. Both rules read the two input relations in their bodies and produce `feeding_break_too_short` in their heads. The local producers query under [How to verify](#how-to-verify) shows the same split.
## Three fact states
A fact is in one of three states, not two:
| State | Meaning | How to send it |
|---|---|---|
| yes | The fact holds | List it (the default state of a listed fact) |
| no | An explicit negative assertion | SDK: `state: 'no'` |
| unknown | Not given | Omit it, or SDK `state: 'unknown'`, which is dropped before evaluation |
Nothing is defaulted. A fact you do not list is unknown, and the engine reads unknown as undefined, never as false.
**"I don't know" is a valid answer.** The agent omits the fact and asks with a partial set. It never fills the gap with a guess, a default or a "typical" value.
Status: ran via MCP `law_ask` with the minutes fact supplied and the child-count fact omitted (legal time 2026-09-01). The captured answer is `NEITHER`, with a per-premise report marking the child-count premise undefined and stating that undefined is not "false".
The agent reports it as it is: "Not established either way; the missing piece is the number of children under eighteen months." What to do next with each status is in [Read answers](/agent-engineering/handle-answers/).
**Use the negative state with care.** What a rule does with `no` depends on the package. Status: NOT RUN for any negative-fact call; reproduce by reading how the package treats negative input first. Until then, ask for the positive counterpart ("she has one child", not "she has no second child") or leave the fact out and report the gap.
## Clarifications: ask early, ask once
A good question names the gap and why it matters: "How many children under eighteen months does the employee have? The feeding-break rules need that count." Batch independent questions into one round (child count, ages, break length, timezone of recorded times). Each answer becomes a typed fact or an explicit unknown.
## Units and quantities
The unit is part of the fact. The feeding-break minutes are plain integers with the unit fixed by the label, so "half an hour" becomes `30`. Richer quantities carry the unit inside the value: the example cases of package `jcgm.gum` (a measurement methodology, not state law) assert an estimate of `150 uV` with an uncertainty of `12 uV`, and the SDK writes a quantity as `'14 calendar_day'`. Status: read from the committed `jcgm.gum` example case and the `@arxo/law` README.
Convert to the unit the signature expects, name the conversion in the report ("30 minutes, converted from 'half an hour'"), and do not assert a fact whose unit is ambiguous ("the bottle holds 2"). A wrong unit makes a wrong fact.
## From document to evidence
A document on the user's desk is not yet evidence in the case. It passes four stations, and a package policy may then decide one more matter, admission:
| Station | What it is | Example |
|---|---|---|
| Document | A file or record the user provided | A May timesheet PDF |
| Extracted proposal | A fact the agent drafts from the document | "Break length is 25 minutes" |
| Accepted input | A fact actually sent to and accepted by the engine | `child_feeding_break_minutes(emp1, er1, 25)` with origin `case_input` |
| Support link | The link between one fact and one document, not yet a decision that it counts | Provenance naming the timesheet plus a quote |
**Extracted is not accepted.** A proposal in the agent's notes affects nothing until it is sent and accepted. The agent never describes a proposal as "in the case".
Admission is not a station the agent walks through. It is a decision the package's evidence policy makes about each support link, and only an admitted link counts toward its claim. The decision is per link, not per document or per fact, so one document may be admitted for one claim and rejected for another. A rejected link counts for neither side.
Each station fails on its own (unreadable file, misreading, rejected input or `FactError`, rejected link). Say which one failed, not "the document didn't work".
### Provenance: name the document, quote the words
The verified call attached the support link in this shape:
```json
// evidence declaration (law_ask "evidence" argument)
[{ "id": "urn:doc:timesheet-may",
"payload": { "kind": "timesheet", "confidence": "0.9" } }]
```
```json
// provenance on one case fact
{ "predicate": "child_feeding_break_minutes",
"args": ["urn:kz:tk:employee:1", "urn:kz:tk:employer:1", 25],
"provenance": {
"evidence": "urn:doc:timesheet-may",
"span": { "quote": "перерыв 25 минут" } } }
```
Status: ran via MCP `law_ask` on the repo server (legal time 2026-09-01). Each grounding fact showed one attached evidence item, and the too-short finding was unchanged.
The quote is short, verbatim and in the document's own language. Copy the exact words that carry the value, never paraphrase into the quote field, and keep it tight enough for a reviewer to find on the page. The confidence figure travels as data. The engine does not grade it, and the agent must not present it as a verified probability.
Package `kz-labour-code` selects no evidence policy. Without one, document and support material stays transport data: it is stored in the case and enters its hash, but gives the fact no support, and the answer bytes are the same as if the document were absent. Facts without provenance are still accepted for ordinary predicates, but nobody can check them. Attach provenance to every fact drafted from a document, and say plainly which facts came from the user's bare assertion.
### Admission: when a support counts
With an active policy, `case_input` is no longer enough. Package `demo.expense-evidence` protects three predicates (`approved`, `expense_amount`, `paid`) behind policy `ExpenseEvidenceV1`. A protected fact enters support only through an admitted support, whatever its origin, even `adjudicated`. The same case, asked five ways:
| Scene | Support edges for `paid(P42)` | Admission outcome | `reimbursable(P42, EMP7)` |
|---|---|---|---|
| 1. Genuine bank confirmation of ANOTHER purchase (B91) | B91 supports | rejected: `NO_ACCEPTANCE_RULE_APPLIED` | `NEITHER` |
| 2. Correct confirmation (B92) added | B91 supports, B92 supports | B91 rejected, B92 admitted | `TRUE_ONLY` |
| 3. B92 excluded again | B91 supports | rejected | `NEITHER` |
| 4. Second independent confirmation (B93), B92 excluded | B91 supports, B93 supports | B91 rejected, B93 admitted | `TRUE_ONLY` |
| 5. No bank document; bare `paid(P42)` assertion, origin `adjudicated` | none | assertion blocked: `PROTECTED_PREDICATE_REQUIRES_EVIDENCE`, issue `PROTECTED_ASSERTION_BLOCKED` | `NEITHER` |
Status: ran locally on the reference engine, release pinned to the package's own `0.2.4`, 2026-10-04. In all five scenes the approval (A17) and amount (R25) supports were admitted, and only the payment leg varied.
Three readings follow:
- A rejected support supports neither polarity and does not prove the denial. In scene 1, `paid(P42)` reads `NEITHER`, not `FALSE_ONLY`.
- A finding can stand next to rejected material only on other admissible grounds. Scene 2 still carries the rejected B91 edge, and the finding rests on B92.
- A bare assertion of a protected predicate changes nothing at any origin. Scene 5 names the blocking cause and the issue in the answer itself.
Before saying "the document is in the case", check whether the question's inputs are ordinary or protected.
### Ambiguity: one passage, several readings
A timesheet cell reading "25/30" could be two breaks, a corrected entry or a range. The agent does not pick a reading silently:
1. Quote the passage verbatim and name the candidate readings.
2. Ask the user which reading holds, offering the candidates.
3. If the user cannot resolve it, extract no fact from that passage and report the gap.
4. Never assert two contradictory facts "to cover both options". That creates a conflict where none existed.
The same holds for illegible scans and unverifiable machine-extracted text: record the doubt at the proposal station and draft no support link until a reading is settled.
### Conflicting documents
The timesheet says 25 minutes, the manager's memo says 30. Two competing `case_input` assertions conflict as proposed inputs, and the engine does not referee them: no automatic rule prefers the newer, signed or more confident document, and any such preference is a policy the agent states, not a computation it runs. Two *admitted* supports of opposite polarity under an active policy are different: a genuine dispute, the literal's state `BOTH`, resolved by a norm, a priority or an authority, never by admission itself.
Status: NOT RUN. Reproduce by sending the full-input feeding-break truth call with two competing child-count facts (counts 1 and 2, each with its own evidence and quote) and reading how the answer reports the clash.
Either way, surface both documents with their quotes, name the conflict and route the choice to its owner ([Human decisions, interpretations and scenarios](/agent-engineering/decisions-and-scenarios/)). Until then, ask with each variant separately and report both answers side by side, each with its own result hash; one merged ask hides what the reviewer needs.
## Limits of the document link
State these plainly whenever they matter:
1. **Acceptance is not truth.** An accepted `case_input` fact means the engine received it, not that the claim is true in the world. The engine does not audit the timesheet.
2. **Paper does not become proof by itself.** Only a derivation from accepted facts carries weight, and for a protected predicate only through an admitted support (see [Read answers](/agent-engineering/handle-answers/)).
3. **Quotes can be checked; paraphrases cannot.** Page numbers, file names and verbatim spans keep a link checkable. Summaries do not.
## Source text versus case documents
For fragment `TK_ART82` of `kz-labour-code`, `law_sources` returns the article 82 excerpt with a hash over the full source text: that is the canon's ground. Case documents differ in role, not in pinnability. `law_case_document` copies a file into the case as evidence with status `presented`, times, a URI and a content hash, and pins it in the case lock, but creates no support edge (see [Cases and processes](/agent-engineering/case-lifecycle/)). Report a rule's grounding in pinned text and a fact's support in a case document separately; they carry different trust.
## How to verify
```sh
./law version
./law query producers \
--param predicate='urn:kz:corpus:clir:labour-code#feeding_break_too_short' \
--param pkg=kz.corpus.labour_code --format human
```
The producers query shows the two rule heads that make `feeding_break_too_short` derived, anchored on article 82. The body relations it lists are the elicited inputs.
MCP reproductions: repeat each MCP call on this page with legal time 2026-09-01 (full facts, child count omitted, with provenance) and call `law_sources` for fragment `urn:kz:corpus:clir:labour-code#TK_ART82`; expected results are in the validation record.
Local admission reproduction (read-only):
```text
1. Lower packs/demo/expense-evidence/scenes/base.lawcase (the lower-case step the demo script uses).
2. Evaluate truth of reimbursable(P42, EMP7) with the semantic release pinned to the committed
IR's own semanticVersion (packs/demo/expense-evidence/clir/expense-evidence.lawir.json),
varying only the document composition per scene; scene 5 also drops every bank document with
its edges and verifications and asserts paid(P42) with origin adjudicated.
Expected: NEITHER / TRUE_ONLY / NEITHER / TRUE_ONLY / NEITHER, with the admission outcomes above.
```
Status: labeled pseudocode for the steps; results as in the validation record. The committed script `python3 packs/demo/expense-evidence/run_demo.py` currently fails with `SEMANTIC_RELEASE_MISMATCH`: it requests the moving `0.2` line, now bound to `0.2.6`, while the committed IR pins `0.2.4`. Pin the request to the IR's release or re-lower the IR first.
## Validation record
Local engine `law 0.1.1`, semantics `law.core/0.2` (`./law version`). MCP runs on the repo server, jurisdiction Republic of Kazakhstan as reported; no profile name exposed.
| Check | Result |
|---|---|
| MCP `law_ask` truth: full facts / child count omitted / with provenance | Ran (authoring session): finding established / `NEITHER`, premise undefined / one evidence item per grounding fact |
| MCP `law_rules`, `law_sources` (`TK_ART82`) | Ran (authoring session): inputs in bodies, verdict in heads; excerpt with full-text hash |
| `./law query producers` feeding-break question | Ran 2026-10-04: two strict producers anchored on article 82 |
| Expense-evidence scenes 1–5 (release pinned to `0.2.4`) | Ran 2026-10-04: `NEITHER / TRUE_ONLY / NEITHER / TRUE_ONLY / NEITHER`; outcomes as in the table |
| `run_demo.py` as committed | Fails 2026-10-04: `SEMANTIC_RELEASE_MISMATCH` |
| SDK `FactError` on a misspelled relation | Observed 2026-10-04 in the first-agent run |
| Negative fact; conflicting child-count facts | NOT RUN |
Previous: [Pin context, editions and time](/agent-engineering/context-and-time/)
Next: [Read answers](/agent-engineering/handle-answers/)