Legal-data conformance
Find coherence failures before full amendment replay exists.
A legal corpus can provide immediate value through identity, manifestation, version, date, language, relationship, and structure checks—even when its amendment semantics are not yet executable.
For legal-information institutes, publishers, and data teams
Make internal assumptions testable.
Identity and manifestations
Duplicate or unstable work IDs, conflicting expressions, missing official files, hash drift, and ambiguous Work–Expression–Manifestation relationships.
Versions and legal time
Broken predecessor/successor chains, overlapping validity intervals, future versions selected as current, missing expiry, and inconsistent effective dates.
Effects and references
Unresolved amendment links, impossible targets, dangling cross-references, unknown effect types, and effects without source manifestations.
Languages and structure
Missing required expressions, asymmetric publication checkpoints, unmatched structural units, unknown XML elements, and schema drift.
A practical first scope
Acquisition evidence can grow without being mistaken for replay.
The Japan, Korea, Poland, Switzerland, EU, and US federal lanes can expose corpus-quality and source-account results while their effect semantics remain blocked.
Inputs
Corpus inventory, source manifestations, identifiers, version metadata, language tags, effects/links where available, and the source-authority model.
Outputs
Typed conformance findings, completeness matrices, broken-link queues, schema inventories, source-account summaries, and exact non-claims.
Operational use
This work inspects corpus data without mutating legal-state output, so institutions can improve source quality and integration before deciding whether replay is feasible.
Next gate
Use the findings to select a source-complete transition subset for typed effect extraction and witnessed dry-run checks.