Markdown for LLMs
Backup and restore
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Backup and restore
## Goal
Capture everything needed to rebuild a working deployment on a clean
machine, and prove the restore works before it is needed.
## Scope
Component `law` CLI deployment: project files, bundles, journals, keys,
and serve configuration; tool `0.1.1`. Scenario: scheduled backups plus
a restore drill on a clean machine. Paths and schedules are synthetic
created-examples.
## Prerequisites
- A backup store the deployment host can write to but cannot silently
rewrite (operator-policy-example: append-only object prefix
`s3://example-backups/law/` — created-example location).
- The production inventory: project dir, bundle files, `--journal` dir,
signing/registry keys, serve invocation, trusted-key configuration.
- A clean drill machine with the pinned `law` version installed
(see [Install and trust artifacts](/operate/install-trust-artifacts/)).
## Steps
1. Back up, as one labeled set (created-example label `bk-2026-10-03`):
- project sources: `law.toml`, `law.lock`, `.law` files, queries,
`pinning.toml` / `*.deontic-window.toml` inputs if used;
- release bundles: every deployed `.arxo` file;
- decision journals: the full `--journal` directories;
- saved evaluations: `law ask --out` directories kept as replay
corpus;
- keys: signing key files and registry credentials — separately
encrypted (see step 3);
- serve configuration: the exact `law serve` invocation, port,
`--token-file` reference, `trustedKeys`, registry `ID=LOCATION`
map.
2. Copy consistency: journals are append-only during serve and each
decision is served only after its record is fsync'd (real serve
behavior). Quiesce writers with this ordered sequence — fencing
the port alone is not enough, because already-accepted calls keep
running and writing after accept stops:
1. stop intake (disable the ingress route);
2. stop the process and confirm the PID is gone (`systemctl stop`
then check). There is no graceful drain: in-flight calls are
killed, and only decisions completed before the kill have
fsynced records — which is exactly the consistent set to copy;
3. confirm no writer holds the journal open (`lsof +D <journal>`
empty, or the process table shows no server);
4. copy or snapshot the volume;
5. resume serving.
If the backup includes mutable case packages, their external
writers (editors, authoring tools) must be quiesced the same way.
Never back up a journal dir with a plain recursive copy while the
server appends to it.
3. Data + secret protection: bundles and journals are integrity-checked
by hash/signature at restore (see Result check); key files keep mode
`0600` (real `law key new` behavior) and travel encrypted under a
key the backup store never sees (operator-policy-example: age/PGP
envelope — created-example tooling).
4. Record a manifest for the set: filename, bytes, `sha256` per file,
`law version --json`, bundle `--verify` lines (created-example
manifest format; the hashes themselves are the check) — plus a
replay map: for each distinct (recorded pin, executor identity)
pair found in the journal, which saved world directory AND
which runtime executable replay it. The pin alone is NOT the
group key: a runtime-only upgrade keeps the pin while the
executable changes, so one pin routinely spans two
executors (decision A on E1, decision B on E2) that must
not share a replay row without a tested compatibility
statement. Replay's guarantee is conditional by
construction: the same world (pin), the same input, and a
compatible executor (real, `replay.rs` header) — the tool
pin-checks the world but does not verify executor
compatibility itself. Each record carries its executor
identity in `engine: {host, version, binaryDigest}` plus an
execution receipt (`implementation:
law-serve-core/<version>`, real, `record.rs`); copy that
identity per (pin, executor) group into the replay map. The
rule is explicit: a group replays only under its named
executor (same version + binary digest), or under a version
covered by a tested compatibility statement recorded in the
manifest. Without either, keep the old binary with the set
— "the current `law` replays everything" is neither
assumed nor denied here. A shared journal normally spans
several (pin, executor) pairs, so the manifest must say
which worlds and which executors cover which records. It
must also declare the contract: full-history replay (every
pair has its world AND its executor saved) or
current-generation-only (older pairs restore as evidence
but have no saved world to replay against).
5. Restore drill on the clean machine: install `law` from pinned supply,
restore the set, then verify in this order (all real commands):
`law inspect <bundle> --verify --trust <key>`,
`law test` over the restored project,
`law eval <saved-calculation-dir>` for saved calculations,
then the grouped journal replay from step 7 (never one
unscoped `replay-record` over a multi-generation journal).
6. Start `law serve` on the drill machine with the restored world +
journal copy and confirm `GET /healthz`, `GET /readyz`, and
`GET /v1/world` answer as in production.
7. Saved-decision replay, grouped by (recorded pin, executor
identity). `law replay-record` re-executes journal entries on
the loaded world and compares `resultHash`/`outcomeHash` per
entry (real) — but it pin-checks every record against that
world first (`replay.rs`: record `/world/pin` versus the
loaded pin), so a record from another generation reports
`PIN_MISMATCH` instead of replaying. Keep the journal whole
and split only the selection:
1. list decision ids per (pin, executor) group from the
restored segments (real record fields `decisionId`,
`world.pin`, `engine`) with the torn-tail-safe kit
helper — never a plain `jq` over the segments, which
prints partial output and exits `5` on the crash state
the reader itself tolerates:
`kit/bin/arxo-list-decisions <journal-copy>` (stdout TSV:
`decisionId`, `world`, `resolutionHash`, `engineVersion`,
`binaryDigest`, in segment order).
Every newline-terminated line must be a valid record or the
helper refuses naming file and offset; trailing bytes without
a newline are accepted ONLY in the last segment and reported
as `TORN_TAIL` (file, offset, length). The helper is
read-only; a server restart instead cuts the tail and logs
it, so run the listing on the backup copy before any restart
consumes the evidence;
2. for each (pin, executor) group, pick the saved world AND the
named executor binary the manifest replay map records for
that pair and run `<that-law> replay-record --world
<that-world> --journal <dir> --decision <id>…` for the
group's ids. A group replayed under a different,
undeclared executor proves nothing about the backup —
executor identity is part of the acceptance, not a
footnote. After a runtime-only upgrade this means two
replay runs for one pin: pre-upgrade ids under the old
binary, post-upgrade ids under the new one;
3. confirm the groups cover every id the manifest claims for the
contract (full history, or the current generation plus older
evidence kept without replay worlds — see step 4). A clean
grouped replay is the restore's acceptance proof.
## Expected result
- Every manifest hash matches; every check command named in the
manifest exits `0`; served endpoints answer with the production pin.
- The drill machine can answer the replay corpus identically to the
backup source.
## Result check
- Byte check: recompute `sha256` over restored files vs the manifest.
- Semantic check: the grouped replay from step 7 reports matching
hashes for every entry of every covered (pin, executor) group,
and the union of the groups equals the manifest's claimed
coverage; spot-check `law case diff` between a production
evaluation and the drill's re-execution — expect an empty diff.
- Any mismatch fails the drill; do not "accept with notes".
## Failures and diagnostics
- Hash mismatch on restore: storage or transfer corruption (or a
tampered set); re-pull from the append-only store, investigate, do
not serve from the mismatched copy.
- `law replay-record` mismatch on an entry: first distinguish
`PIN_MISMATCH` (the entry belongs to another (pin, executor)
group — replay it against the world AND the executor binary
the manifest names for its recorded pair) from a hash mismatch
on the right world, which means world drift (wrong bundle or
lock), the wrong executor for a runtime-only split group, or
journal truncation.
- Permission errors on keys (`--key`, `--token-file`): restore
preserved the wrong mode/owner; keys must be readable only by the
service account (operator-policy-example: dedicated `lawsvc` user —
created-example name).
- Partial journal copy (server was appending): date the set as
inconsistent and retake it quiesced; a torn journal is not repaired
by retrying the copy.
- `TORN_TAIL` from the listing helper after a crash (not a quiesce
failure): expected and acceptable — the tail bytes were never a
confirmed decision. Keep the backup bytes as taken (evidence),
record file/offset/length in the drill report, and let the drill
restart cut the tail with its log line. Any `REFUSAL` instead
fails the drill: investigate the named file and offset, do not
hand-edit the segment into parseability.
## Support boundaries
- Supported: hash/signature verification, scenario tests, saved-decision
replay as restore acceptance.
- Reference setting: store location, labels, schedule
(created-example: daily sets, 30-day retention), and the service
account are operator-policy-examples.
- "Verified" in this article means only "hashes match and replays are
clean" — it does not attest the normative correctness of the canon.
## Next step
Schedule the drill cadence in the operations calendar; for undoing a
bad switch rather than a lost host, see [Rollback](/operate/rollback/).