Skip to content
docs
Arxo ↗

Batch: many cases, one loop

For LLMs7 sections

The engine answers one case per call: there is no native batch call that takes many cases at once. To evaluate many cases, the application loops: one case in, one answer out, errors kept per case. This page shows that loop on the deadline example.

Output
deadline-app-0.1.0.zip: src/batch.ts, test/fixtures/batch-cases.json

Use the loop when cases arrive together — a file of forms, a queue of submissions, a backlog to clear. Each case still travels the full data path (validate, facts, query, engine, read, view, capture); the loop adds no new layer and changes no module.

Run the shipped fixture — three good rows and one bad row:

Terminal
node --import tsx src/cli.ts batch test/fixtures/batch-cases.json /tmp/report.ndjson --captures /tmp/captures
Output
rows: 4, ok: 3, failed: 1
report: /tmp/report.ndjson

One NDJSON line per row, in input order. The bad row records its validation error and the run continues past it:

JSON
{"id": "calc-1", "ok": true, "headline": "The period ends on 2026-03-20", "capture": "case-0001.json"}
{"id": "verify-1", "ok": true, "headline": "Yes — the proposed date is established", "capture": "case-0002.json"}
{"id": "bad-1", "ok": false, "errors": ["eventDate must be a calendar date in YYYY-MM-DD form"]}
{"id": "verify-2", "ok": true, "headline": "Not established either way", "capture": "case-0004.json"}

(The detail field ships on every kept line; it is trimmed here for width. The full lines are in the report file the command writes.)

For each case, in order:

  1. Validate the raw fields with validateFormInput. A bad form records an error for that case and the loop moves on; no engine call runs.
  2. Build a fresh case input with toCaseInput(form). Fresh per case: nothing carries over from the previous iteration.
  3. Evaluate with evaluateCollect(form) — or the truth call when the row proposes a candidate — against the shared model handle (de.bgb.fristen@0.1.0, offline: true).
  4. Read the answer into value|claim|not-computed|unknown and shape it with toViewModel.
  5. Store a capture per kept answer under --captures/case-NNNN.json with CAPTURE_FORMAT deadline-app.capture/1.
src/batch.ts
export async function runBatch(
rows: BatchRow[],
opts: { model?: LawPackage; capturesDir: string; onLine: (line: BatchLine) => void },
): Promise<BatchCounts> {

The try/catch sits inside the loop, around one case: a failure records itself next to its case id and the run continues — including an engine-level throw, which a dedicated test pins with a failing model stub. The shared model handle is read-only across iterations; all per-case state is built fresh inside the iteration. Capture files are index-named (case-0004.json), never id-named: row ids are untrusted input and must not become path segments. The command exits 1 when any row failed, 0 when all passed.

The loop awaits each case before starting the next. Results are written as each case completes, so finished cases never accumulate — but the input array is held whole: runBatch(rows: BatchRow[], …) takes every row up front, and that array plus the working data of the active computation is what lives in memory. What sequential execution does not bound is the report itself either — the file grows with the number of cases, and that is expected. Order is the other reason: failures stay attributable, and the report reads in input order.

The fixture run ends with kept rows and error rows side by side:

RowCallOutcomeCapture
calc-1collectCOMPUTED 2026-03-20case-0001.json
verify-1truth on 2026-03-20TRUE_ONLYcase-0002.json
bad-1—validation error; the engine never runs—
verify-2truth on 2026-03-21NEITHER (the period ends on the 23rd)case-0004.json

Each row makes one call — collect when no candidate is proposed, truth when one is. A row that needs both answers is two rows, or a loop of your own over runBatch lines.

Every kept row reports via as local and carries the program, semantic, and result hashes; sources is empty in this canon build. To reproduce any row later, replay its capture with replayCapture under the same pins and the same SDK.

  • Not a native batch call: the engine exposes no multi-case entry point, and the loop does not emulate one — it calls the single-case functions once per case.
  • Not parallel: iterations run in order against one shared handle.
  • Not a merge: cases never combine into one case input, and answers never combine into one document. Each kept answer gets its own capture file, replayed by the SDK that saved it.
  • Not constant total memory: the input array is held whole and the report file grows with the run; only the per-case working set is bounded to the active computation. Streaming the input (AsyncIterable, chunked reads, backpressure on the sink) is a separate extension for large runs, not part of this recipe.
Terminal
node --import tsx --test test/batch.test.ts
Output
ok 1 - batch runs the fixture: two keeps, one validation error, one NEITHER
ok 2 - batch continues after an engine-level failure
# tests 2
# pass 2
# fail 0

Predict before you open each answer.

Row `bad-1` has a malformed `eventDate`. Does the engine run for it, and does the batch stop?

Neither. Validation rejects the row before any engine call, the error is recorded next to bad-1, and the loop moves on to verify-2. The command exits 1 at the end because one row failed.

Why is the capture for `verify-2` named `case-0004.json`, not `verify-2.json`?

Capture files are index-named: verify-2 is the fourth row. Row ids are untrusted input and must not become path segments.

A row proposes a candidate date. How many engine calls does it make?

One: a truth call on the candidate. Each row makes one call — collect when no candidate is proposed, truth when one is. A row that needs both answers is two rows.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.