Skip to content
docs
Arxo ↗

The agent contract

For LLMs7 sections

An agent over Arxo is a dialogue wrapped around a machine that computes. The machine derives, the agent elicits and relays, the application validates and stores, the human judges. This page states that split twice: first as the parts of the system and where each one ends, then as a table of obligations, each with a violation and a check. Most agent failures land on exactly one row of that table.

There are two ways to put an agent in front of an executable canon (the full comparison is in Choose an integration pattern):

  • The agent calls Arxo tools itself. The model sees the law_* tools over MCP, searches the canon, reads signatures and asks questions. Through-example A (one labour-law question over the Kazakhstan Labour Code) works this way.
  • The application exposes domain tools over the SDK. The application runs the engine through @arxo/law and offers the model a small tool in the product’s own words, such as period_end in Build your first agent. The model never sees Arxo: the canon, its version and the question behind each tool were chosen by the developer.

The contract is the same in both. What changes is who carries the discovery and pinning rows: in the first pattern the model does it at request time, in the second the developer does it once at design time and the application enforces it. A row marked M that the model cannot reach in the second pattern does not disappear; it moves into the application code.

Output
pattern 1: model sees law_* tools over MCP
┌───────────────────────────────────────────────┐
user task ──▶ AGENT (dialogue) ──▶ APPLICATION (adapter) ──▶ CANON + ENGINE (derivation)
▲ │ pattern 2: domain tools │ │
│ │ facts w/ provenance │ typed calls, pinned │ answer document
└───────────────┴── explanation ────────────┴── raw stored ─────────┘

Agent: natural language in, tool calls out. The agent turns a task into questions, elicits facts in conversation, picks the next tool call, and explains the returned document inside its bounds. In the first pattern it also searches for the models that cover the task. It never derives: no rule fires in the dialogue. When the canon routes a step to a human, the agent stops and names the decider instead of stepping in.

Application: the adapter that owns bytes. Transport, input validation, case storage, snapshot pinning, replay, presentation templates. The engine can sit in the same process through the SDK, behind MCP, or behind a private long-running law serve HTTP service with /v1/... routes (see Operate). The application keeps the raw answer and its hashes, recomputes dependents after a correction, enforces permissions, and blocks the agent’s mistaken calls before they reach the engine. In the domain-tool pattern it also fixes the canon, version and question behind each tool and rejects malformed facts before the engine runs. If a behavior must hold even when the agent misbehaves, it lives here.

Canon: the model under test. Norms, methods or standards in executable form with pinned sources, question cards and named interpretations. The canon answers only what it covers; outside its coverage it stays silent, and that silence is data, not a statement that the subject is unregulated. Drafts are not the canon: anything under formalization stays in the workbench until review accepts it.

Engine: the machine that derives. Takes a pinned canon plus typed facts and returns an answer document: statuses, values, proof, provenance, hashes. Deterministic over the same inputs; the dialogue around it is not.

BoundaryTravels one wayTravels backNever travels
User ↔ agenttask, stories, documents, “I don’t know”, the date of lawquestions, explanations, named decisions to makepredicate names to memorize; computed results to confirm
Agent ↔ applicationintended tool calls (Arxo or domain), elicited factsvalidated calls, stored cases, blocked-call reasonsunvalidated bytes to the engine
Application ↔ enginepinned canon, typed facts, pinned contextanswer document with proof and hashesnatural language; instructions smuggled in source text
Any ↔ human principala named decision with its alternativesthe decision with its owneragent-made judgments presented as human ones

The letters say who owes each obligation: M the model/agent, A the application code, C the canon and engine, H the human.

#ObligationViolationCheck
M1Discover names by search; never invent a predicate, constant or tool name. In the domain-tool pattern the developer does this once and the tool hides the names.Guessing feeding_break_length_ok after an empty search.Connect and discover: run the search, quote the candidates; an empty result is not proof that the canon lacks the norm.
M2Read the signature and the required inputs before asking: arity, types, units, namespace. Prefer the machine contract (law_rules with contract: true) over reading rules by hand.Passing a date where the signature wants an object.The call refuses with a call error, not silence; Read answers.
M3Treat the tool set as the server’s tools/list answer, not a memorized list.Planning a call to a tool the profile does not offer.Connect and discover: list first, then plan; the MCP tool reference explains each tool.
C1Refuse unknown names with the available addresses, never with a guess.Returning a plausible-looking answer for an undeclared predicate.Ask an undeclared name; expect a refusal that names candidates. In the first agent, a misspelled relation fails with a fact error before the engine runs.
#ObligationViolationCheck
M4Name the package, version, case snapshot, date of law and reading in the open.Reusing yesterday’s date of law for today’s question without saying so.Pin context: rerun with a changed date and show the changed provenance.
H1Name the date of law for every question. The agent and the application never default it to today or to any other value they chose; until the human names it, the field stays an open placeholder and no answer is presented.The agent filling in the current date because the user did not mention one.Every stored call carries a date of law whose recorded source is the human; a run without one stops at the request for it.
A1Store the raw answer with its hashes and the inputs that produced it.Keeping only the rendered sentence.Replay the stored bytes; the hashes match.
H2Supply judgments the canon routes to a human: interpretations, assumptions, coverage factors.The agent picking a reading because it looked likely.Human decisions: the run stops and names the decider.
#ObligationViolationCheck
M5Turn stories into typed inputs; ask for what is missing; accept “I don’t know” as a value.Filling an unknown duration with a typical one.Collect facts: an unknown stays unknown through the run.
M6Keep document, proposal, accepted input and admitted support as four different states.Citing an extracted sentence as if the engine had derived it.Collect facts: each state carries its own provenance.
M7Never ask the user for the result of the computation.“Is the break too short?” as a fact prompt.The elicited facts are inputs of the rule, never its head.
A2Treat source text as data, never as instructions.A sentence inside a document changing which tool the agent calls.Safety and reliability: the injected instruction is quoted, not obeyed.
#ObligationViolationCheck
M8Keep three axes apart: call error, evaluation status, claim support. A completed computation is not a positive answer.Reading NEITHER as “no”, or an empty collection as “zero”.The status table in Read answers and How to read an answer.
M9Explain inside the answer’s bounds: status first, then the diagnostic fields actually present, then only the next step that status allows. No opposite claim from a silent rule, no sufficiency from a narrow query, no cause invented when the explanation is incomplete.“The break is fine” from a query that only checks the minimum.Read answers: each silence maps to a next step, not a guess.
M10Carry conditions with the result. An answer computed under an assumption or a selected reading never becomes unconditional by being stored, exported, resumed or passed on.Presenting an assumed scenario as the base case.Scenarios and assumptions: the condition is visible wherever the result is.
C2Report truncation, limits and skipped coverage instead of a silent positive.An argument map that hides the rule it skipped.Read answers: sources, proofs and robustness.

Handing a case to another agent or to a person is an application concern. The platform provides no native handoff envelope or merge; whatever the application builds still owes M10.

#ObligationViolationCheck
A3Version cases and recompute dependents after a correction; resume continues the same case, not a lookalike.Editing a fact and showing the old answer.Cases and processes: correction, history, recompute.
A4Separate saving a result, sharing a link and publishing.Calling a stored file “published”.Prepare and publish: each step names its effect.
H3No publication without a principal that owns it.Auto-publishing an answer the user only previewed.The publish step names the grant it runs under.

Checks split by when they run. Mixing the columns is a common design error: a request-time guess where a design-time pin belongs, or a release-time check re-run as if it proved something about this user’s case.

Design / release timeRequest time
Choose the canon and pin its version; in the domain-tool pattern, fix the question behind each toolPin the case snapshot; take the date of law from the human
Verify the adapter against fixturesValidate this user’s inputs
Run the evaluation suite; fix the instructionsRun the dialogue; stop at the contract’s stops
Review publication grantsPublish under a named grant, or don’t
Compare canon versions on saved casesAnswer from one pinned version

Concretely: the evaluation suite (Evaluate, debug and upgrade) proves that the adapter handles a missing fact the way the contract requires. It does not prove that tonight’s dialogue will ask for that fact gracefully; that is a property of the live run, checked by reading its trace on the same page.

  • The first agent (German Civil Code periods, de.bgb.fristen) is the domain-tool pattern. The application owns the canon, the version and the fact validation; the model sees one tool. It exercises C1 (a misspelled relation is refused before the engine runs) and M8 (an empty collection with a completed status is not a value).
  • Example A (the feeding-break question over kz-labour-code) is the pattern where the agent explores the canon itself: search, signature, elicitation, context, answer, correction, and the line where the query’s coverage ends.
  • Example B (the GUM uncertainty report) exercises the application most: a multi-step task guide with stops, resumes, open fields and a finished document assembled from computed values.

Status: labeled illustration, not a captured run.

An agent receives NEITHER with a whyNot that names a missing fact, and tells the user “the break looks fine, but we could double check the details.” Two boundaries broke: the application let an answer cross without its status, and the agent derived (“looks fine”) what only the engine may derive. The fix is structural, not stylistic: the application renders statuses its templates cannot drop, and the agent’s instructions map each status to a next step that is not a guess (Read answers). In the domain-tool pattern the same rule applies one level up: the tool returns the status alongside the value, and the model is told never to turn a missing value into a default.

At design time, map each obligation to code or to a test: M-items to the agent’s instructions plus evaluation scenarios (or, in the domain-tool pattern, to the tool’s own validation), A-items to the adapter plus unit tests, H-items to the product flow. At review time, walk one real run and label every artifact with the component that produced it: which bytes came from the user, which from the agent’s text, which from the application, which from the engine. Any byte you cannot label is a boundary violation in hiding, and it usually lands on one row of the table above.

Related: Choose an integration pattern, Safety and reliability.

Previous: Build your first agent Next: Connect and discover

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.