Conversation
added 3 commits
September 20, 2026 21:46
Install a schema-level invariant: every reference an export can resolve must join rows that live in the same knowledge base. A column foreign key proves the target exists, not that it is the same KB's — the exporter would otherwise mint local IRIs naming foreign rows, or silently drop vocabulary references that resolve to nothing. Three layers: a precondition scan that refuses the migration on a dirty ledger, row triggers on every edge (deferred constraint triggers on same-table self-references so COPY and multi-row inserts are judged at commit), and kb-ownership immutability on every owned table.
The export read model moves from "open a pool connection per page" to a caller-provided transaction: a route can now pin one REPEATABLE READ snapshot across the preflight check, the vocabulary reads, and every page. A mid-stream commit can no longer leak half a rule or a dangling wasGeneratedBy into a finished file. Integrity is checked on the rows that survive, not on a second look. Every page query selects the referenced row's kb atomically with the row itself; a foreign, dangling, or merged-out reference refuses the whole export rather than minting a local IRI that names another KB's row or silently dropping a vocabulary link. The scan covers the provenance chain, derivation premises, edge qualifiers, vocabulary references, and the open-statement / time-mention / binding edges the schema has grown since — the same families migration 0070 guards on the write side. Read-model additions carried by the same pages: statement qualifiers, time mentions, typed-fact sources, fact layer/phrase/validity grade, evidence quote offsets, chunk origin metadata, entity descriptions, document reader/time-context fields, and vocabulary updated_at.
Rules, attribute rules, document versions, chunks and evidence were read from the same snapshot but never serialized; their facts' prov:wasGeneratedBy, prov:wasDerivedFrom and locator references dangled in the file. Serialize them in snapshot order so every emitted reference resolves. New upstream ledger surfaces serialize too: class/relation updatedAt, entity descriptions, document reader/time-context, chunk origin metadata, evidence quote offsets, open-statement layer/phrase/ qualifiers/time mentions, typed-fact fromStatement provenance, validFromGrade, and evidenceOrigin.
This was referenced Sep 20, 2026
Contributor
Author
|
Closing pending author's final review of the PR set — will reopen once the series is finalized. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PROBLEM
Rules, attribute rules, document versions, chunks and evidence are read from the export snapshot but never serialized — their facts'
prov:wasGeneratedBy,prov:wasDerivedFromand locator references dangle in the finished file. Newer ledger surfaces are absent too: class/relationupdatedAt, entity descriptions, document reader/time context, chunk origin metadata, evidence quote offsets, open-statement layer/phrase/qualifiers/time mentions, typed-factfromStatementprovenance,validFromGrade, andevidenceOrigin.WHY THIS IS A GENERIC UTOPIA BUG OR CONTRACT GAP
The serializer was never extended as the ledger schema grew, so the export contract silently covers only a subset of what the read model returns — a file that claims to be the ledger but omits whole families and leaves references pointing at nodes that were never written.
FIX
Serialize every family in snapshot order: vocabulary (classes, relations), rules and attribute rules with their conditions and expression-predicate references, documents and document versions, chunks, entities, facts with qualifiers / time mentions / lifecycle anchors, evidence, and derivations with ordered premises. Emission fails closed — an unresolvable reference aborts the export rather than minting an IRI for a node that was never written.
REGRESSION EVIDENCE
New in-memory unit tests cover each emission family in both Turtle and JSON-LD: rule conclusions and conditions, expression predicates, document-version locators, chunk/evidence links, open-statement ledgers, qualifiers (literal and entity), lifecycle marks, imported vocabularies, and fail-closed paths (unresolvable qualifier type, merged entity reference, foreign IRI minting). 32/32
rdf::tests pass.COMPATIBILITY RISK
Additive only: exports grow by the previously missing families; no existing triple is removed or renamed.
Stacked on #832 → #833 — this PR's diff includes those commits until they merge.