Skip to content
docs
Arxo ↗

Alignment and concept bridges

For LLMs6 sections

Different editions, languages, and upstream sources describe the same ideas in different words. Alignment records those correspondences explicitly — without ever merging meanings by assumption. One revision in two languages is two editions with a declared correspondence, never one shared meaning.

Confirmation marks in parentheses, such as (S) or (I), follow the legend on the topic index.

Section titled “A worked mapping: the renamed holding link”

The lab tour renames the shared holding relation between two fixture states: owns in labparcels.iface at 0.1.0 becomes holds at 0.2.0, with parameters, kind, key, and label unchanged. The compiler treats those as two different nodes — a relation pin is not executed, so no annotation carries the old identity across. Saying “holds continues owns” is therefore a human-authored mapping, and a useful one only spells out its direction, its preserved conditions, and its status:

  • Direction. Old revision to new revision: rows recorded under owns read as rows under holds. The reverse is not claimed — nothing promises the old consumers accept the new name.
  • Preserved conditions. Same parameter shapes (Owner, Parcel), same kind (institutional), same key (parcel), same label text. Any later change to one side reopens the mapping.
  • Status. review while a second author checks the claim, accepted once recorded, revoked if the meanings drift apart. Revocation deletes no rows; it withdraws the permission to read across.

The report schema for the narrow profiles draws the decisive boundary for exactly this shape: a decision carries interpretation_basis declared_mapping_assumption, keeps all_source_conditions_retained true, and keeps semantic_equivalence_proved false — see the report format. The mapping below follows that shape as a team-kept illustration (T); no checkout command produces this report for the parcels fixtures, whose profiles cover other pairs of systems:

JSON
{
"source": "labparcels.iface@0.1.0 owns",
"target": "labparcels.iface@0.2.0 holds",
"direction": "source_to_target_under_interpretation",
"status": "accepted",
"conditions": {
"interpretation_basis": "declared_mapping_assumption",
"all_source_conditions_retained": true,
"semantic_equivalence_proved": false
}
}

Read the last line as the whole point: the record proves that a reviewed human said “these correspond under these conditions”, never that the two nodes mean the same thing.

The grammar offers an edition-level correspondence block between two references, with a relation name, a free-text status, and a reviewer:

Output
alignment_decl = "align", reference, "with", reference, "{",
{ alignment_item },
"}" ;
alignment_item = "relation", qualified_name, ";"
| "status", identifier, ";"
| "reviewer", expression, ";"
| metadata_item ;

The declaration is metadata only: it is stored on the document, never lowered into formulas or evidence. Status values are a per-package team convention, not a closed vocabulary. Four maturity facts, kept in sync with the vocabulary bridge notes: the declaration is supported grammar and compiles in real packages (S); no structural query reads align rows back — the links question covers five other link kinds (S); adoption is broad, with over two thousand blocks across dozens of source files observed in the corpus (S); and the semantic effect is nil (S).

For dictionary-scale correspondence, the established pattern is an ordinary package that maps shared vocabularies to outside records, handwritten and reviewed like any other corpus content. The reference example, arxo.jurisdiction_alignment, maps the currency, geography, country, and courts dictionaries to their outside records. A generic composition mechanism was considered for this job and rejected in favor of explicit packages. This is a team convention (T).

Shared dictionaries themselves (vocab.geo and its siblings) stay types-plus-constants packages with no rules and no sources, which is what makes them safe mapping endpoints. (S)

The concept catalog is a derived full-text index over compiled packages and import snapshots. It answers “which packages talk about this idea” and never merges semantics: a catalog hit is a pointer for a human or a policy to follow up, not an identity claim. Observed tool surface:

Output
$ python3 corpus/tools/concepts/catalog.py --help
usage: catalog.py [-h] --db DB {build,search,show,candidates} ...
Derived concept catalogue: pinned package/Import IR views, never a semantic
merger.

Build the index from pinned inputs, then search, show, or list candidates from the database file. The catalog tool requires a source checkout (I): it lives in the repository, not in the public release.

A worked route over the lab fixtures: an author needs a parcel concept for a new package and checks the catalog before minting one. The inputs are pinned in a manifest — two compiled dependency snapshots, nothing else:

JSON
{
"format": "arxo.concept-catalog.inputs/0.1",
"packages": ["../parcels-fees/deps/*.lawir.json"],
"imports": []
}

Build the index, then search for the idea:

Output
$ python3 corpus/tools/concepts/catalog.py --db /tmp/lab-work/catalog-route.db build docs/corpus/lab/fixtures/catalog-mini/manifest.json
{"concepts":13,"indexed":2,"removed":0,"reused":0,"sources":2}
$ python3 corpus/tools/concepts/catalog.py --db /tmp/lab-work/catalog-route.db search parcel | python3 -c "import json,sys; [print(r['concept']['native_id'],'|',r['concept']['project']) for r in json.load(sys.stdin)]"
urn:law:lab:parcels:iface#Parcel | labparcels.iface
urn:law:lab:parcels:registry#Parcel | labparcels.registry
urn:law:lab:parcels:registry#ParcelKind | labparcels.registry
urn:law:lab:parcels:iface#Owner | labparcels.iface
urn:law:lab:parcels:registry#registered | labparcels.registry
urn:law:lab:parcels:iface#owns | labparcels.iface
urn:law:lab:parcels:iface#ParcelKind | labparcels.iface
urn:law:lab:parcels:registry#kind_of | labparcels.registry
urn:law:lab:parcels:registry#transfer_request | labparcels.registry

Nine hits across two packages. Inspect the canonical-looking one:

Output
$ python3 corpus/tools/concepts/catalog.py --db /tmp/lab-work/catalog-route.db show 'urn:law:lab:parcels:iface#Parcel' | python3 -c "import json,sys; d=json.load(sys.stdin); print(d['name'],'|',d['family'],'|',d['native_id']); print('visibility:',d['raw']['visibility'],'| project:',d['project']); print('labels:',[l['text'] for l in d['labels']])"
Parcel | type | urn:law:lab:parcels:iface#Parcel
visibility: public | project: labparcels.iface
labels: ['registered parcel']

A public entity type with a matching label — a reuse candidate. Ask the catalog for correspondence candidates before deciding:

Output
$ python3 corpus/tools/concepts/catalog.py --db /tmp/lab-work/catalog-route.db candidates 'urn:law:lab:parcels:iface#Parcel' | python3 -c "import json,sys; d=json.load(sys.stdin); print('automatic_acceptance:',d['automatic_acceptance'],'| scope:',d['scope']); [print(c['target']['native_id'],'|',c['status'],'| score',c['score'],'|',','.join(c['reasons']),'| failures:',','.join(c['checks']['failures']) or 'none') for c in d['candidates']]" | head -3; echo ...
automatic_acceptance: False | scope: bounded identifier/label/signature retrieval; no equivalence proof
urn:law:lab:parcels:iface#Owner | candidate | score 5 | lexical_overlap | failures: none
urn:law:lab:parcels:registry#Parcel | incompatible | score 45 | lexical_overlap,same_label | failures: different_type_kind
...

Eight further rows follow, all incompatible (the full JSON also carries pending: semantic_correspondence_requires_review on every row). No automatic acceptance — and read the top-scoring row carefully: registry#Parcel is refused as a correspondence endpoint with different_type_kind, but that is a matcher limitation, not a verdict of different concepts. The matcher only compares declaration kinds (entity versus alias) and never follows the alias target; the compiled node itself says where it points:

Output
$ python3 -c "import json; d=json.load(open('fixtures/parcels-fees/deps/labparcels.registry.lawir.json')); n=next(n for n in d['nodes'] if n.get('name')=='Parcel'); print(n['typeKind'], '->', n['target']['name'])"
alias -> labparcels.iface#Parcel

registry#Parcel is a direct alias for the shared iface::Parcel — same concept, re-exported spelling. The reuse decision is the author’s, and here it is to reuse: import labparcels.iface at its pinned version and reference its public Parcel, instead of minting a third parcel type. The catalog row stays a pointer; the decision lives in the new package’s import pins.

A few bespoke bridges map specific outside systems into corpus form. Each is a hand-registered module for one pair of systems — physics units against a public unit ontology, order theory against two proof assistants — plus a positive-rules helper. Observed modules:

Output
$ ls corpus/tools/concepts/alignment_profiles/
__pycache__
isabelle_orders.py
lean_orders.py
lean_stacks.py
physics.py
positive_rules.py

Each profile ships its own policy and report formats, and the full-profile checks verify source links, replay, revocation, and impact. These are narrow experimental bridges, not a general mapping framework: adding a new pair means authoring a new module with its own policy. Experimental (X).

  • LawQL reference for the structural relations that alignment tooling queries.
  • Schema catalog for the alignment policy and report formats.
  • MCP tools for catalog-style discovery from clients.

Documentation for Arxo. Writings — blog.arxo.io.

Anonymous visit counts on stats.arxo.io, no cookies.