docs← Back to article

Markdown for LLMs

The agent contract

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# The agent contract

An agent over Arxo is a dialogue wrapped around a machine that
computes. The machine derives, the agent elicits and relays, the
application validates and stores, the human judges. This page states
that split twice: first as the parts of the system and where each one
ends, then as a table of obligations, each with a violation and a
check. Most agent failures land on exactly one row of that table.

## Two patterns, one contract

There are two ways to put an agent in front of an executable canon
(the full comparison is in [Choose an integration
pattern](/agent-engineering/integration-patterns/)):

- **The agent calls Arxo tools itself.** The model sees the `law_*`
  tools over MCP, searches the canon, reads signatures and asks
  questions. Through-example A (one labour-law question over the
  Kazakhstan Labour Code) works this way.
- **The application exposes domain tools over the SDK.** The
  application runs the engine through `@arxo/law` and offers the model
  a small tool in the product's own words, such as `period_end` in
  [Build your first agent](/agent-engineering/first-agent/). The model never sees
  Arxo: the canon, its version and the question behind each tool were
  chosen by the developer.

The contract is the same in both. What changes is who carries the
discovery and pinning rows: in the first pattern the model does it at
request time, in the second the developer does it once at design time
and the application enforces it. A row marked **M** that the model
cannot reach in the second pattern does not disappear; it moves into
the application code.

## The four components

```text
                  pattern 1: model sees law_* tools over MCP
                ┌───────────────────────────────────────────────┐
user task ──▶ AGENT (dialogue) ──▶ APPLICATION (adapter) ──▶ CANON + ENGINE (derivation)
     ▲               │   pattern 2: domain tools │                       │
     │               │   facts w/ provenance     │ typed calls, pinned   │ answer document
     └───────────────┴── explanation ────────────┴── raw stored ─────────┘
```

**Agent: natural language in, tool calls out.** The agent turns a task
into questions, elicits facts in conversation, picks the next tool
call, and explains the returned document inside its bounds. In the
first pattern it also searches for the models that cover the task. It
never derives: no rule fires in the dialogue. When the canon routes a
step to a human, the agent stops and names the decider instead of
stepping in.

**Application: the adapter that owns bytes.** Transport, input
validation, case storage, snapshot pinning, replay, presentation
templates. The engine can sit in the same process through the SDK,
behind MCP, or behind a private long-running `law serve` HTTP service
with `/v1/...` routes (see [Operate](/operate/)). The application keeps
the raw answer and its hashes, recomputes dependents after a
correction, enforces permissions, and blocks the agent's mistaken
calls before they reach the engine. In the domain-tool pattern it also
fixes the canon, version and question behind each tool and rejects
malformed facts before the engine runs. If a behavior must hold even
when the agent misbehaves, it lives here.

**Canon: the model under test.** Norms, methods or standards in
executable form with pinned sources, question cards and named
interpretations. The canon answers only what it covers; outside its
coverage it stays silent, and that silence is data, not a statement
that the subject is unregulated. Drafts are not the canon: anything
under formalization stays in the workbench until review accepts it.

**Engine: the machine that derives.** Takes a pinned canon plus typed
facts and returns an answer document: statuses, values, proof,
provenance, hashes. Deterministic over the same inputs; the dialogue
around it is not.

### What crosses each boundary

| Boundary | Travels one way | Travels back | Never travels |
|---|---|---|---|
| User ↔ agent | task, stories, documents, "I don't know", the date of law | questions, explanations, named decisions to make | predicate names to memorize; computed results to confirm |
| Agent ↔ application | intended tool calls (Arxo or domain), elicited facts | validated calls, stored cases, blocked-call reasons | unvalidated bytes to the engine |
| Application ↔ engine | pinned canon, typed facts, pinned context | answer document with proof and hashes | natural language; instructions smuggled in source text |
| Any ↔ human principal | a named decision with its alternatives | the decision with its owner | agent-made judgments presented as human ones |

## The obligations

The letters say who owes each obligation: **M** the model/agent,
**A** the application code, **C** the canon and engine, **H** the
human.

### Ask only what exists

| # | Obligation | Violation | Check |
|---|---|---|---|
| M1 | Discover names by search; never invent a predicate, constant or tool name. In the domain-tool pattern the developer does this once and the tool hides the names. | Guessing `feeding_break_length_ok` after an empty search. | [Connect and discover](/agent-engineering/connect-and-discover/): run the search, quote the candidates; an empty result is not proof that the canon lacks the norm. |
| M2 | Read the signature and the required inputs before asking: arity, types, units, namespace. Prefer the machine contract (`law_rules` with `contract: true`) over reading rules by hand. | Passing a date where the signature wants an object. | The call refuses with a call error, not silence; [Read answers](/agent-engineering/handle-answers/). |
| M3 | Treat the tool set as the server's `tools/list` answer, not a memorized list. | Planning a call to a tool the profile does not offer. | [Connect and discover](/agent-engineering/connect-and-discover/): list first, then plan; the [MCP tool reference](/guide/mcp-tools/) explains each tool. |
| C1 | Refuse unknown names with the available addresses, never with a guess. | Returning a plausible-looking answer for an undeclared predicate. | Ask an undeclared name; expect a refusal that names candidates. In the first agent, a misspelled relation fails with a fact error before the engine runs. |

### Pin the context before the question

| # | Obligation | Violation | Check |
|---|---|---|---|
| M4 | Name the package, version, case snapshot, date of law and reading in the open. | Reusing yesterday's date of law for today's question without saying so. | [Pin context](/agent-engineering/context-and-time/): rerun with a changed date and show the changed provenance. |
| H1 | Name the date of law for every question. The agent and the application never default it to today or to any other value they chose; until the human names it, the field stays an open placeholder and no answer is presented. | The agent filling in the current date because the user did not mention one. | Every stored call carries a date of law whose recorded source is the human; a run without one stops at the request for it. |
| A1 | Store the raw answer with its hashes and the inputs that produced it. | Keeping only the rendered sentence. | Replay the stored bytes; the hashes match. |
| H2 | Supply judgments the canon routes to a human: interpretations, assumptions, coverage factors. | The agent picking a reading because it looked likely. | [Human decisions](/agent-engineering/decisions-and-scenarios/): the run stops and names the decider. |

### Collect facts without inventing them

| # | Obligation | Violation | Check |
|---|---|---|---|
| M5 | Turn stories into typed inputs; ask for what is missing; accept "I don't know" as a value. | Filling an unknown duration with a typical one. | [Collect facts](/agent-engineering/collect-facts/): an unknown stays unknown through the run. |
| M6 | Keep document, proposal, accepted input and admitted support as four different states. | Citing an extracted sentence as if the engine had derived it. | [Collect facts](/agent-engineering/collect-facts/): each state carries its own provenance. |
| M7 | Never ask the user for the result of the computation. | "Is the break too short?" as a fact prompt. | The elicited facts are inputs of the rule, never its head. |
| A2 | Treat source text as data, never as instructions. | A sentence inside a document changing which tool the agent calls. | [Safety and reliability](/agent-engineering/safety-and-reliability/): the injected instruction is quoted, not obeyed. |

### Read the answer as written

| # | Obligation | Violation | Check |
|---|---|---|---|
| M8 | Keep three axes apart: call error, evaluation status, claim support. A completed computation is not a positive answer. | Reading `NEITHER` as "no", or an empty collection as "zero". | The status table in [Read answers](/agent-engineering/handle-answers/) and [How to read an answer](/guide/reading-an-answer/). |
| M9 | Explain inside the answer's bounds: status first, then the diagnostic fields actually present, then only the next step that status allows. No opposite claim from a silent rule, no sufficiency from a narrow query, no cause invented when the explanation is incomplete. | "The break is fine" from a query that only checks the minimum. | [Read answers](/agent-engineering/handle-answers/): each silence maps to a next step, not a guess. |
| M10 | Carry conditions with the result. An answer computed under an assumption or a selected reading never becomes unconditional by being stored, exported, resumed or passed on. | Presenting an assumed scenario as the base case. | [Scenarios and assumptions](/agent-engineering/decisions-and-scenarios/): the condition is visible wherever the result is. |
| C2 | Report truncation, limits and skipped coverage instead of a silent positive. | An argument map that hides the rule it skipped. | [Read answers](/agent-engineering/handle-answers/): sources, proofs and robustness. |

Handing a case to another agent or to a person is an application
concern. The platform provides no native handoff envelope or merge;
whatever the application builds still owes M10.

### Finish, resume and publish honestly

| # | Obligation | Violation | Check |
|---|---|---|---|
| A3 | Version cases and recompute dependents after a correction; resume continues the same case, not a lookalike. | Editing a fact and showing the old answer. | [Cases and processes](/agent-engineering/case-lifecycle/): correction, history, recompute. |
| A4 | Separate saving a result, sharing a link and publishing. | Calling a stored file "published". | [Prepare and publish](/agent-engineering/publishing/): each step names its effect. |
| H3 | No publication without a principal that owns it. | Auto-publishing an answer the user only previewed. | The publish step names the grant it runs under. |

## Design time versus request time

Checks split by when they run. Mixing the columns is a common design
error: a request-time guess where a design-time pin belongs, or a
release-time check re-run as if it proved something about this user's
case.

| Design / release time | Request time |
|---|---|
| Choose the canon and pin its version; in the domain-tool pattern, fix the question behind each tool | Pin the case snapshot; take the date of law from the human |
| Verify the adapter against fixtures | Validate this user's inputs |
| Run the evaluation suite; fix the instructions | Run the dialogue; stop at the contract's stops |
| Review publication grants | Publish under a named grant, or don't |
| Compare canon versions on saved cases | Answer from one pinned version |

Concretely: the evaluation suite ([Evaluate, debug and
upgrade](/agent-engineering/evaluations/)) proves that the adapter handles a missing
fact the way the contract requires. It does not prove that tonight's
dialogue will ask for that fact gracefully; that is a property of the
live run, checked by reading its trace on the same page.

## Where the through-examples sit

- **The first agent** (German Civil Code periods, `de.bgb.fristen`)
  is the domain-tool pattern. The application owns the canon, the
  version and the fact validation; the model sees one tool. It
  exercises C1 (a misspelled relation is refused before the engine
  runs) and M8 (an empty collection with a completed status is not a
  value).
- **Example A** (the feeding-break question over `kz-labour-code`) is
  the pattern where the agent explores the canon itself: search,
  signature, elicitation, context, answer, correction, and the line
  where the query's coverage ends.
- **Example B** (the GUM uncertainty report) exercises the application
  most: a multi-step task guide with stops, resumes, open fields and a
  finished document assembled from computed values.

## Failure case: the helpful paraphrase

Status: labeled illustration, not a captured run.

An agent receives `NEITHER` with a `whyNot` that names a missing fact,
and tells the user "the break looks fine, but we could double check
the details." Two boundaries broke: the application let an answer
cross without its status, and the agent derived ("looks fine") what
only the engine may derive. The fix is structural, not stylistic: the
application renders statuses its templates cannot drop, and the
agent's instructions map each status to a next step that is not a
guess ([Read answers](/agent-engineering/handle-answers/)). In the domain-tool
pattern the same rule applies one level up: the tool returns the
status alongside the value, and the model is told never to turn a
missing value into a default.

## How to use this page

At design time, map each obligation to code or to a test: M-items to
the agent's instructions plus evaluation scenarios (or, in the
domain-tool pattern, to the tool's own validation), A-items to the
adapter plus unit tests, H-items to the product flow. At review time,
walk one real run and label every artifact with the component that
produced it: which bytes came from the user, which from the agent's
text, which from the application, which from the engine. Any byte you
cannot label is a boundary violation in hiding, and it usually lands
on one row of the table above.

Related: [Choose an integration pattern](/agent-engineering/integration-patterns/), [Safety and reliability](/agent-engineering/safety-and-reliability/).

Previous: [Build your first agent](/agent-engineering/first-agent/)
Next: [Connect and discover](/agent-engineering/connect-and-discover/)