The agent contract
An agent over Arxo is a dialogue wrapped around a machine that computes. The machine derives, the agent elicits and relays, the application validates and stores, the human judges. This page states that split twice: first as the parts of the system and where each one ends, then as a table of obligations, each with a violation and a check. Most agent failures land on exactly one row of that table.
Two patterns, one contract
Section titled “Two patterns, one contract”There are two ways to put an agent in front of an executable canon (the full comparison is in Choose an integration pattern):
- The agent calls Arxo tools itself. The model sees the
law_*tools over MCP, searches the canon, reads signatures and asks questions. Through-example A (one labour-law question over the Kazakhstan Labour Code) works this way. - The application exposes domain tools over the SDK. The
application runs the engine through
@arxo/lawand offers the model a small tool in the product’s own words, such asperiod_endin Build your first agent. The model never sees Arxo: the canon, its version and the question behind each tool were chosen by the developer.
The contract is the same in both. What changes is who carries the discovery and pinning rows: in the first pattern the model does it at request time, in the second the developer does it once at design time and the application enforces it. A row marked M that the model cannot reach in the second pattern does not disappear; it moves into the application code.
The four components
Section titled “The four components” pattern 1: model sees law_* tools over MCP ┌───────────────────────────────────────────────┐user task ──▶ AGENT (dialogue) ──▶ APPLICATION (adapter) ──▶ CANON + ENGINE (derivation) ▲ │ pattern 2: domain tools │ │ │ │ facts w/ provenance │ typed calls, pinned │ answer document └───────────────┴── explanation ────────────┴── raw stored ─────────┘Agent: natural language in, tool calls out. The agent turns a task into questions, elicits facts in conversation, picks the next tool call, and explains the returned document inside its bounds. In the first pattern it also searches for the models that cover the task. It never derives: no rule fires in the dialogue. When the canon routes a step to a human, the agent stops and names the decider instead of stepping in.
Application: the adapter that owns bytes. Transport, input
validation, case storage, snapshot pinning, replay, presentation
templates. The engine can sit in the same process through the SDK,
behind MCP, or behind a private long-running law serve HTTP service
with /v1/... routes (see Operate). The application keeps
the raw answer and its hashes, recomputes dependents after a
correction, enforces permissions, and blocks the agent’s mistaken
calls before they reach the engine. In the domain-tool pattern it also
fixes the canon, version and question behind each tool and rejects
malformed facts before the engine runs. If a behavior must hold even
when the agent misbehaves, it lives here.
Canon: the model under test. Norms, methods or standards in executable form with pinned sources, question cards and named interpretations. The canon answers only what it covers; outside its coverage it stays silent, and that silence is data, not a statement that the subject is unregulated. Drafts are not the canon: anything under formalization stays in the workbench until review accepts it.
Engine: the machine that derives. Takes a pinned canon plus typed facts and returns an answer document: statuses, values, proof, provenance, hashes. Deterministic over the same inputs; the dialogue around it is not.
What crosses each boundary
Section titled “What crosses each boundary”| Boundary | Travels one way | Travels back | Never travels |
|---|---|---|---|
| User ↔ agent | task, stories, documents, “I don’t know”, the date of law | questions, explanations, named decisions to make | predicate names to memorize; computed results to confirm |
| Agent ↔ application | intended tool calls (Arxo or domain), elicited facts | validated calls, stored cases, blocked-call reasons | unvalidated bytes to the engine |
| Application ↔ engine | pinned canon, typed facts, pinned context | answer document with proof and hashes | natural language; instructions smuggled in source text |
| Any ↔ human principal | a named decision with its alternatives | the decision with its owner | agent-made judgments presented as human ones |
The obligations
Section titled “The obligations”The letters say who owes each obligation: M the model/agent, A the application code, C the canon and engine, H the human.
Ask only what exists
Section titled “Ask only what exists”| # | Obligation | Violation | Check |
|---|---|---|---|
| M1 | Discover names by search; never invent a predicate, constant or tool name. In the domain-tool pattern the developer does this once and the tool hides the names. | Guessing feeding_break_length_ok after an empty search. | Connect and discover: run the search, quote the candidates; an empty result is not proof that the canon lacks the norm. |
| M2 | Read the signature and the required inputs before asking: arity, types, units, namespace. Prefer the machine contract (law_rules with contract: true) over reading rules by hand. | Passing a date where the signature wants an object. | The call refuses with a call error, not silence; Read answers. |
| M3 | Treat the tool set as the server’s tools/list answer, not a memorized list. | Planning a call to a tool the profile does not offer. | Connect and discover: list first, then plan; the MCP tool reference explains each tool. |
| C1 | Refuse unknown names with the available addresses, never with a guess. | Returning a plausible-looking answer for an undeclared predicate. | Ask an undeclared name; expect a refusal that names candidates. In the first agent, a misspelled relation fails with a fact error before the engine runs. |
Pin the context before the question
Section titled “Pin the context before the question”| # | Obligation | Violation | Check |
|---|---|---|---|
| M4 | Name the package, version, case snapshot, date of law and reading in the open. | Reusing yesterday’s date of law for today’s question without saying so. | Pin context: rerun with a changed date and show the changed provenance. |
| H1 | Name the date of law for every question. The agent and the application never default it to today or to any other value they chose; until the human names it, the field stays an open placeholder and no answer is presented. | The agent filling in the current date because the user did not mention one. | Every stored call carries a date of law whose recorded source is the human; a run without one stops at the request for it. |
| A1 | Store the raw answer with its hashes and the inputs that produced it. | Keeping only the rendered sentence. | Replay the stored bytes; the hashes match. |
| H2 | Supply judgments the canon routes to a human: interpretations, assumptions, coverage factors. | The agent picking a reading because it looked likely. | Human decisions: the run stops and names the decider. |
Collect facts without inventing them
Section titled “Collect facts without inventing them”| # | Obligation | Violation | Check |
|---|---|---|---|
| M5 | Turn stories into typed inputs; ask for what is missing; accept “I don’t know” as a value. | Filling an unknown duration with a typical one. | Collect facts: an unknown stays unknown through the run. |
| M6 | Keep document, proposal, accepted input and admitted support as four different states. | Citing an extracted sentence as if the engine had derived it. | Collect facts: each state carries its own provenance. |
| M7 | Never ask the user for the result of the computation. | “Is the break too short?” as a fact prompt. | The elicited facts are inputs of the rule, never its head. |
| A2 | Treat source text as data, never as instructions. | A sentence inside a document changing which tool the agent calls. | Safety and reliability: the injected instruction is quoted, not obeyed. |
Read the answer as written
Section titled “Read the answer as written”| # | Obligation | Violation | Check |
|---|---|---|---|
| M8 | Keep three axes apart: call error, evaluation status, claim support. A completed computation is not a positive answer. | Reading NEITHER as “no”, or an empty collection as “zero”. | The status table in Read answers and How to read an answer. |
| M9 | Explain inside the answer’s bounds: status first, then the diagnostic fields actually present, then only the next step that status allows. No opposite claim from a silent rule, no sufficiency from a narrow query, no cause invented when the explanation is incomplete. | “The break is fine” from a query that only checks the minimum. | Read answers: each silence maps to a next step, not a guess. |
| M10 | Carry conditions with the result. An answer computed under an assumption or a selected reading never becomes unconditional by being stored, exported, resumed or passed on. | Presenting an assumed scenario as the base case. | Scenarios and assumptions: the condition is visible wherever the result is. |
| C2 | Report truncation, limits and skipped coverage instead of a silent positive. | An argument map that hides the rule it skipped. | Read answers: sources, proofs and robustness. |
Handing a case to another agent or to a person is an application concern. The platform provides no native handoff envelope or merge; whatever the application builds still owes M10.
Finish, resume and publish honestly
Section titled “Finish, resume and publish honestly”| # | Obligation | Violation | Check |
|---|---|---|---|
| A3 | Version cases and recompute dependents after a correction; resume continues the same case, not a lookalike. | Editing a fact and showing the old answer. | Cases and processes: correction, history, recompute. |
| A4 | Separate saving a result, sharing a link and publishing. | Calling a stored file “published”. | Prepare and publish: each step names its effect. |
| H3 | No publication without a principal that owns it. | Auto-publishing an answer the user only previewed. | The publish step names the grant it runs under. |
Design time versus request time
Section titled “Design time versus request time”Checks split by when they run. Mixing the columns is a common design error: a request-time guess where a design-time pin belongs, or a release-time check re-run as if it proved something about this user’s case.
| Design / release time | Request time |
|---|---|
| Choose the canon and pin its version; in the domain-tool pattern, fix the question behind each tool | Pin the case snapshot; take the date of law from the human |
| Verify the adapter against fixtures | Validate this user’s inputs |
| Run the evaluation suite; fix the instructions | Run the dialogue; stop at the contract’s stops |
| Review publication grants | Publish under a named grant, or don’t |
| Compare canon versions on saved cases | Answer from one pinned version |
Concretely: the evaluation suite (Evaluate, debug and upgrade) proves that the adapter handles a missing fact the way the contract requires. It does not prove that tonight’s dialogue will ask for that fact gracefully; that is a property of the live run, checked by reading its trace on the same page.
Where the through-examples sit
Section titled “Where the through-examples sit”- The first agent (German Civil Code periods,
de.bgb.fristen) is the domain-tool pattern. The application owns the canon, the version and the fact validation; the model sees one tool. It exercises C1 (a misspelled relation is refused before the engine runs) and M8 (an empty collection with a completed status is not a value). - Example A (the feeding-break question over
kz-labour-code) is the pattern where the agent explores the canon itself: search, signature, elicitation, context, answer, correction, and the line where the query’s coverage ends. - Example B (the GUM uncertainty report) exercises the application most: a multi-step task guide with stops, resumes, open fields and a finished document assembled from computed values.
Failure case: the helpful paraphrase
Section titled “Failure case: the helpful paraphrase”Status: labeled illustration, not a captured run.
An agent receives NEITHER with a whyNot that names a missing fact,
and tells the user “the break looks fine, but we could double check
the details.” Two boundaries broke: the application let an answer
cross without its status, and the agent derived (“looks fine”) what
only the engine may derive. The fix is structural, not stylistic: the
application renders statuses its templates cannot drop, and the
agent’s instructions map each status to a next step that is not a
guess (Read answers). In the domain-tool
pattern the same rule applies one level up: the tool returns the
status alongside the value, and the model is told never to turn a
missing value into a default.
How to use this page
Section titled “How to use this page”At design time, map each obligation to code or to a test: M-items to the agent’s instructions plus evaluation scenarios (or, in the domain-tool pattern, to the tool’s own validation), A-items to the adapter plus unit tests, H-items to the product flow. At review time, walk one real run and label every artifact with the component that produced it: which bytes came from the user, which from the agent’s text, which from the application, which from the engine. Any byte you cannot label is a boundary violation in hiding, and it usually lands on one row of the table above.
Related: Choose an integration pattern, Safety and reliability.
Previous: Build your first agent Next: Connect and discover
Documentation for Arxo. Writings — blog.arxo.io.
Anonymous visit counts on stats.arxo.io, no cookies.