docs← Back to article

Markdown for LLMs

Backup and restore

The source Markdown for this article. Copy it into your assistant or download it as a text file.

Download this articlePlain text ↗
# Backup and restore

## Goal

Capture everything needed to rebuild a working deployment on a clean
machine, and prove the restore works before it is needed.

## Scope

Component `law` CLI deployment: project files, bundles, journals, keys,
and serve configuration; tool `0.1.1`. Scenario: scheduled backups plus
a restore drill on a clean machine. Paths and schedules are synthetic
created-examples.

## Prerequisites

- A backup store the deployment host can write to but cannot silently
  rewrite (operator-policy-example: append-only object prefix
  `s3://example-backups/law/` — created-example location).
- The production inventory: project dir, bundle files, `--journal` dir,
  signing/registry keys, serve invocation, trusted-key configuration.
- A clean drill machine with the pinned `law` version installed
  (see [Install and trust artifacts](/operate/install-trust-artifacts/)).

## Steps

1. Back up, as one labeled set (created-example label `bk-2026-10-03`):
   - project sources: `law.toml`, `law.lock`, `.law` files, queries,
     `pinning.toml` / `*.deontic-window.toml` inputs if used;
   - release bundles: every deployed `.arxo` file;
   - decision journals: the full `--journal` directories;
   - saved evaluations: `law ask --out` directories kept as replay
     corpus;
   - keys: signing key files and registry credentials — separately
     encrypted (see step 3);
   - serve configuration: the exact `law serve` invocation, port,
     `--token-file` reference, `trustedKeys`, registry `ID=LOCATION`
     map.
2. Copy consistency: journals are append-only during serve and each
   decision is served only after its record is fsync'd (real serve
   behavior). Quiesce writers with this ordered sequence — fencing
   the port alone is not enough, because already-accepted calls keep
   running and writing after accept stops:
   1. stop intake (disable the ingress route);
   2. stop the process and confirm the PID is gone (`systemctl stop`
      then check). There is no graceful drain: in-flight calls are
      killed, and only decisions completed before the kill have
      fsynced records — which is exactly the consistent set to copy;
   3. confirm no writer holds the journal open (`lsof +D <journal>`
      empty, or the process table shows no server);
   4. copy or snapshot the volume;
   5. resume serving.
   If the backup includes mutable case packages, their external
   writers (editors, authoring tools) must be quiesced the same way.
   Never back up a journal dir with a plain recursive copy while the
   server appends to it.
3. Data + secret protection: bundles and journals are integrity-checked
   by hash/signature at restore (see Result check); key files keep mode
   `0600` (real `law key new` behavior) and travel encrypted under a
   key the backup store never sees (operator-policy-example: age/PGP
   envelope — created-example tooling).
4. Record a manifest for the set: filename, bytes, `sha256` per file,
   `law version --json`, bundle `--verify` lines (created-example
   manifest format; the hashes themselves are the check) — plus a
   replay map: for each distinct (recorded pin, executor identity)
   pair found in the journal, which saved world directory AND
   which runtime executable replay it. The pin alone is NOT the
   group key: a runtime-only upgrade keeps the pin while the
   executable changes, so one pin routinely spans two
   executors (decision A on E1, decision B on E2) that must
   not share a replay row without a tested compatibility
   statement. Replay's guarantee is conditional by
   construction: the same world (pin), the same input, and a
   compatible executor (real, `replay.rs` header) — the tool
   pin-checks the world but does not verify executor
   compatibility itself. Each record carries its executor
   identity in `engine: {host, version, binaryDigest}` plus an
   execution receipt (`implementation:
   law-serve-core/<version>`, real, `record.rs`); copy that
   identity per (pin, executor) group into the replay map. The
   rule is explicit: a group replays only under its named
   executor (same version + binary digest), or under a version
   covered by a tested compatibility statement recorded in the
   manifest. Without either, keep the old binary with the set
   — "the current `law` replays everything" is neither
   assumed nor denied here. A shared journal normally spans
   several (pin, executor) pairs, so the manifest must say
   which worlds and which executors cover which records. It
   must also declare the contract: full-history replay (every
   pair has its world AND its executor saved) or
   current-generation-only (older pairs restore as evidence
   but have no saved world to replay against).
5. Restore drill on the clean machine: install `law` from pinned supply,
   restore the set, then verify in this order (all real commands):
   `law inspect <bundle> --verify --trust <key>`,
   `law test` over the restored project,
   `law eval <saved-calculation-dir>` for saved calculations,
   then the grouped journal replay from step 7 (never one
   unscoped `replay-record` over a multi-generation journal).
6. Start `law serve` on the drill machine with the restored world +
   journal copy and confirm `GET /healthz`, `GET /readyz`, and
   `GET /v1/world` answer as in production.
7. Saved-decision replay, grouped by (recorded pin, executor
   identity). `law replay-record` re-executes journal entries on
   the loaded world and compares `resultHash`/`outcomeHash` per
   entry (real) — but it pin-checks every record against that
   world first (`replay.rs`: record `/world/pin` versus the
   loaded pin), so a record from another generation reports
   `PIN_MISMATCH` instead of replaying. Keep the journal whole
   and split only the selection:
   1. list decision ids per (pin, executor) group from the
      restored segments (real record fields `decisionId`,
      `world.pin`, `engine`) with the torn-tail-safe kit
      helper — never a plain `jq` over the segments, which
      prints partial output and exits `5` on the crash state
      the reader itself tolerates:
      `kit/bin/arxo-list-decisions <journal-copy>` (stdout TSV:
      `decisionId`, `world`, `resolutionHash`, `engineVersion`,
      `binaryDigest`, in segment order).
      Every newline-terminated line must be a valid record or the
      helper refuses naming file and offset; trailing bytes without
      a newline are accepted ONLY in the last segment and reported
      as `TORN_TAIL` (file, offset, length). The helper is
      read-only; a server restart instead cuts the tail and logs
      it, so run the listing on the backup copy before any restart
      consumes the evidence;
   2. for each (pin, executor) group, pick the saved world AND the
      named executor binary the manifest replay map records for
      that pair and run `<that-law> replay-record --world
      <that-world> --journal <dir> --decision <id>…` for the
      group's ids. A group replayed under a different,
      undeclared executor proves nothing about the backup —
      executor identity is part of the acceptance, not a
      footnote. After a runtime-only upgrade this means two
      replay runs for one pin: pre-upgrade ids under the old
      binary, post-upgrade ids under the new one;
   3. confirm the groups cover every id the manifest claims for the
      contract (full history, or the current generation plus older
      evidence kept without replay worlds — see step 4). A clean
      grouped replay is the restore's acceptance proof.

## Expected result

- Every manifest hash matches; every check command named in the
  manifest exits `0`; served endpoints answer with the production pin.
- The drill machine can answer the replay corpus identically to the
  backup source.

## Result check

- Byte check: recompute `sha256` over restored files vs the manifest.
- Semantic check: the grouped replay from step 7 reports matching
  hashes for every entry of every covered (pin, executor) group,
  and the union of the groups equals the manifest's claimed
  coverage; spot-check `law case diff` between a production
  evaluation and the drill's re-execution — expect an empty diff.
- Any mismatch fails the drill; do not "accept with notes".

## Failures and diagnostics

- Hash mismatch on restore: storage or transfer corruption (or a
  tampered set); re-pull from the append-only store, investigate, do
  not serve from the mismatched copy.
- `law replay-record` mismatch on an entry: first distinguish
  `PIN_MISMATCH` (the entry belongs to another (pin, executor)
  group — replay it against the world AND the executor binary
  the manifest names for its recorded pair) from a hash mismatch
  on the right world, which means world drift (wrong bundle or
  lock), the wrong executor for a runtime-only split group, or
  journal truncation.
- Permission errors on keys (`--key`, `--token-file`): restore
  preserved the wrong mode/owner; keys must be readable only by the
  service account (operator-policy-example: dedicated `lawsvc` user —
  created-example name).
- Partial journal copy (server was appending): date the set as
  inconsistent and retake it quiesced; a torn journal is not repaired
  by retrying the copy.
- `TORN_TAIL` from the listing helper after a crash (not a quiesce
  failure): expected and acceptable — the tail bytes were never a
  confirmed decision. Keep the backup bytes as taken (evidence),
  record file/offset/length in the drill report, and let the drill
  restart cut the tail with its log line. Any `REFUSAL` instead
  fails the drill: investigate the named file and offset, do not
  hand-edit the segment into parseability.

## Support boundaries

- Supported: hash/signature verification, scenario tests, saved-decision
  replay as restore acceptance.
- Reference setting: store location, labels, schedule
  (created-example: daily sets, 30-day retention), and the service
  account are operator-policy-examples.
- "Verified" in this article means only "hashes match and replays are
  clean" — it does not attest the normative correctness of the canon.

## Next step

Schedule the drill cadence in the operations calendar; for undoing a
bad switch rather than a lost host, see [Rollback](/operate/rollback/).