# Backup and restore ## Goal Capture everything needed to rebuild a working deployment on a clean machine, and prove the restore works before it is needed. ## Scope Component `law` CLI deployment: project files, bundles, journals, keys, and serve configuration; tool `0.1.1`. Scenario: scheduled backups plus a restore drill on a clean machine. Paths and schedules are synthetic created-examples. ## Prerequisites - A backup store the deployment host can write to but cannot silently rewrite (operator-policy-example: append-only object prefix `s3://example-backups/law/` — created-example location). - The production inventory: project dir, bundle files, `--journal` dir, signing/registry keys, serve invocation, trusted-key configuration. - A clean drill machine with the pinned `law` version installed (see [Install and trust artifacts](/operate/install-trust-artifacts/)). ## Steps 1. Back up, as one labeled set (created-example label `bk-2026-10-03`): - project sources: `law.toml`, `law.lock`, `.law` files, queries, `pinning.toml` / `*.deontic-window.toml` inputs if used; - release bundles: every deployed `.arxo` file; - decision journals: the full `--journal` directories; - saved evaluations: `law ask --out` directories kept as replay corpus; - keys: signing key files and registry credentials — separately encrypted (see step 3); - serve configuration: the exact `law serve` invocation, port, `--token-file` reference, `trustedKeys`, registry `ID=LOCATION` map. 2. Copy consistency: journals are append-only during serve and each decision is served only after its record is fsync'd (real serve behavior). Quiesce writers with this ordered sequence — fencing the port alone is not enough, because already-accepted calls keep running and writing after accept stops: 1. stop intake (disable the ingress route); 2. stop the process and confirm the PID is gone (`systemctl stop` then check). There is no graceful drain: in-flight calls are killed, and only decisions completed before the kill have fsynced records — which is exactly the consistent set to copy; 3. confirm no writer holds the journal open (`lsof +D ` empty, or the process table shows no server); 4. copy or snapshot the volume; 5. resume serving. If the backup includes mutable case packages, their external writers (editors, authoring tools) must be quiesced the same way. Never back up a journal dir with a plain recursive copy while the server appends to it. 3. Data + secret protection: bundles and journals are integrity-checked by hash/signature at restore (see Result check); key files keep mode `0600` (real `law key new` behavior) and travel encrypted under a key the backup store never sees (operator-policy-example: age/PGP envelope — created-example tooling). 4. Record a manifest for the set: filename, bytes, `sha256` per file, `law version --json`, bundle `--verify` lines (created-example manifest format; the hashes themselves are the check) — plus a replay map: for each distinct (recorded pin, executor identity) pair found in the journal, which saved world directory AND which runtime executable replay it. The pin alone is NOT the group key: a runtime-only upgrade keeps the pin while the executable changes, so one pin routinely spans two executors (decision A on E1, decision B on E2) that must not share a replay row without a tested compatibility statement. Replay's guarantee is conditional by construction: the same world (pin), the same input, and a compatible executor (real, `replay.rs` header) — the tool pin-checks the world but does not verify executor compatibility itself. Each record carries its executor identity in `engine: {host, version, binaryDigest}` plus an execution receipt (`implementation: law-serve-core/`, real, `record.rs`); copy that identity per (pin, executor) group into the replay map. The rule is explicit: a group replays only under its named executor (same version + binary digest), or under a version covered by a tested compatibility statement recorded in the manifest. Without either, keep the old binary with the set — "the current `law` replays everything" is neither assumed nor denied here. A shared journal normally spans several (pin, executor) pairs, so the manifest must say which worlds and which executors cover which records. It must also declare the contract: full-history replay (every pair has its world AND its executor saved) or current-generation-only (older pairs restore as evidence but have no saved world to replay against). 5. Restore drill on the clean machine: install `law` from pinned supply, restore the set, then verify in this order (all real commands): `law inspect --verify --trust `, `law test` over the restored project, `law eval ` for saved calculations, then the grouped journal replay from step 7 (never one unscoped `replay-record` over a multi-generation journal). 6. Start `law serve` on the drill machine with the restored world + journal copy and confirm `GET /healthz`, `GET /readyz`, and `GET /v1/world` answer as in production. 7. Saved-decision replay, grouped by (recorded pin, executor identity). `law replay-record` re-executes journal entries on the loaded world and compares `resultHash`/`outcomeHash` per entry (real) — but it pin-checks every record against that world first (`replay.rs`: record `/world/pin` versus the loaded pin), so a record from another generation reports `PIN_MISMATCH` instead of replaying. Keep the journal whole and split only the selection: 1. list decision ids per (pin, executor) group from the restored segments (real record fields `decisionId`, `world.pin`, `engine`) with the torn-tail-safe kit helper — never a plain `jq` over the segments, which prints partial output and exits `5` on the crash state the reader itself tolerates: `kit/bin/arxo-list-decisions ` (stdout TSV: `decisionId`, `world`, `resolutionHash`, `engineVersion`, `binaryDigest`, in segment order). Every newline-terminated line must be a valid record or the helper refuses naming file and offset; trailing bytes without a newline are accepted ONLY in the last segment and reported as `TORN_TAIL` (file, offset, length). The helper is read-only; a server restart instead cuts the tail and logs it, so run the listing on the backup copy before any restart consumes the evidence; 2. for each (pin, executor) group, pick the saved world AND the named executor binary the manifest replay map records for that pair and run ` replay-record --world --journal --decision …` for the group's ids. A group replayed under a different, undeclared executor proves nothing about the backup — executor identity is part of the acceptance, not a footnote. After a runtime-only upgrade this means two replay runs for one pin: pre-upgrade ids under the old binary, post-upgrade ids under the new one; 3. confirm the groups cover every id the manifest claims for the contract (full history, or the current generation plus older evidence kept without replay worlds — see step 4). A clean grouped replay is the restore's acceptance proof. ## Expected result - Every manifest hash matches; every check command named in the manifest exits `0`; served endpoints answer with the production pin. - The drill machine can answer the replay corpus identically to the backup source. ## Result check - Byte check: recompute `sha256` over restored files vs the manifest. - Semantic check: the grouped replay from step 7 reports matching hashes for every entry of every covered (pin, executor) group, and the union of the groups equals the manifest's claimed coverage; spot-check `law case diff` between a production evaluation and the drill's re-execution — expect an empty diff. - Any mismatch fails the drill; do not "accept with notes". ## Failures and diagnostics - Hash mismatch on restore: storage or transfer corruption (or a tampered set); re-pull from the append-only store, investigate, do not serve from the mismatched copy. - `law replay-record` mismatch on an entry: first distinguish `PIN_MISMATCH` (the entry belongs to another (pin, executor) group — replay it against the world AND the executor binary the manifest names for its recorded pair) from a hash mismatch on the right world, which means world drift (wrong bundle or lock), the wrong executor for a runtime-only split group, or journal truncation. - Permission errors on keys (`--key`, `--token-file`): restore preserved the wrong mode/owner; keys must be readable only by the service account (operator-policy-example: dedicated `lawsvc` user — created-example name). - Partial journal copy (server was appending): date the set as inconsistent and retake it quiesced; a torn journal is not repaired by retrying the copy. - `TORN_TAIL` from the listing helper after a crash (not a quiesce failure): expected and acceptable — the tail bytes were never a confirmed decision. Keep the backup bytes as taken (evidence), record file/offset/length in the drill report, and let the drill restart cut the tail with its log line. Any `REFUSAL` instead fails the drill: investigate the named file and offset, do not hand-edit the segment into parseability. ## Support boundaries - Supported: hash/signature verification, scenario tests, saved-decision replay as restore acceptance. - Reference setting: store location, labels, schedule (created-example: daily sets, 30-day retention), and the service account are operator-policy-examples. - "Verified" in this article means only "hashes match and replays are clean" — it does not attest the normative correctness of the canon. ## Next step Schedule the drill cadence in the operations calendar; for undoing a bad switch rather than a lost host, see [Rollback](/operate/rollback/).