fix(GATE-ISSUE-INDEX-TABLE-SHAPE): record checkers report every failure in one run, not the first (rebased onto the frozen-evidence stack) - #3353
Merged
mudler-agent merged 4 commits intoSep 29, 2026
Conversation
…en-evidence records still cite .agents/completed/issue-index.md is the frozen archive every Frozen archive evidence block quotes, and it is not frozen in the way the name claims. 2e84a07 deleted the file wholesale -- 913 lines -- alongside its legitimate spec addition. e3539d9 then restored a hand-picked copy, 886 content lines: 27 archived rows came back missing. Those 27 dropped rows are exactly what 43 frozen-evidence records cite: 24 cite a line past the restored file's end, 19 cite a line that now holds a different row. No gate sees any of that today. The frozen-evidence comparison runs only in the _intake branch of validate_issue_record, whose 9 records all cite surviving lines, so all 43 stale quotes are row-owned records that validate untouched. But every one of those 43 quotes is byte-true to the pre-deletion archive, so no edit to any record can repair it against a truncated archive; the defect is in the archive, not in the quoting. The restore is a splice, verified byte for byte: re-inserting the 27 lines at the positions difflib reports leaves the file byte-identical to the 2e84a07~1 archive except line 406, which keeps e3539d9's own deliberate edit there (a space inserted so `[](` does not parse as an empty Markdown link). 28 insertions, 0 deletions, 0 link re-points -- every restored target already resolves from .agents/completed/, so e3539d9's re-pointing rule needs nothing further. The one record that cites line 406 (ISSUE-mudlerGH-1256) is byte-equal against the restored line. MEASURED over the 831 records that carry a frozen-evidence block, against the committed LF blob: before the restore 350 are byte-equal and 458 fail this base's byte-only comparison (481 under the mudler#3350 link-rebase comparison); after the restore 373 are byte-equal, and under mudler#3350's comparison all 831 resolve, 0 failing. check-agent-record stays green with unchanged row counts (ENGINE=179 MODEL=384 QUANT=87 KERNEL=60 BACKEND=90). The agent-issue-index failure set is proven IDENTICAL before and after -- one pre-existing invalid record in both, nothing greened, nothing newly red; this PR greens nothing on its own and is the data half of a pair whose checker half enforces the comparison outside _intake on top of mudler#3350. The row's spec lands here too (.agents/specs/gate-issue-archive-restore.md), which is what makes the row canonical for the record; the row is unplaced, so no matrix row is touched. ISSUE-LOCAL-01M3NR65JDEV0NH28W5ZPQDPYC FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
…ve link, and nothing else An _intake record's Problem must quote its frozen archive line byte for byte, and check-agent-record's check_links requires every link in a record to resolve from the record's OWN directory. Those two demands cannot both hold on one string. The archive lives at .agents/completed/issue-index.md, one level under .agents, so a link to a spec reads ../specs/x.md there; the record that QUOTES the row lives at .agents/issues/<owner>/, two levels down, so the same link must read ../../specs/x.md from there. e3539d9 re-pointed the ISSUE-mudlerGH-1033 quote to the record-relative spelling to clear the dangling link, which is correct for its own gate and broke the byte comparison: first divergence at column 2191 of archive line 350, the archive line 2237 characters against 2240 in the evidence. The comparison now resolves each side's relative link targets to the file they denote, and only then. Byte equality is still tried first and still answers an unmoved archive, so no record that passes today changes how it passes. When the bytes do differ, the rebase admits exactly the spelling a MOVE forces and nothing else: remote and root-absolute targets are never rewritten, so a swapped URL still fails byte for byte, and a rebase to a different file, a different fragment, a changed title, a changed kind, a changed row cell, or a different line all still fail. The resolution is scoped to the record's own directory, which validate_issue_record derives from the owner directory it already checked, and valid_intake_archive_evidence keeps byte-only behaviour for a caller that cannot say where the quote lives rather than guessing a base. MEASURED over the 831 records that carry a Frozen archive evidence block: 350 are byte-equal, 438 differ from their cited line ONLY by a relative link rebase, and 43 cite a line the archive no longer has -- 24 past its 886 lines, since 2e84a07 deleted the file and e3539d9 re-added a shorter one, and 19 naming a row that moved. This change does not extend the check to row-owned records, so it greens nothing on its own; it removes the contradiction that made those 438 unfixable-as-written, and it is the prerequisite for enforcing the block outside _intake at all. The 43 stale citations are a separate question, because their quotes are historically true and the archive is not frozen in the way the name claims. On the ISSUE-mudlerGH-1033 record specifically: byte-equal False, modulo link rebase True, and the record validates once the check is reachable for a row-owned quote. Its own state is UNKNOWN with Availability METADATA_ONLY, so the evidence branch that runs this comparison is the _intake one and does not apply to it today; that is precisely why the drift went unnoticed, and why this is filed as the prerequisite rather than as a green gate. TESTS: six cases in tests/scripts/test_issue_records.py. One accepts a quote whose only difference is a rebase that resolves to the same file; four hold the rejections (a rebase to a different file, a different fragment, a remote target, a root-absolute target) so the relaxation cannot widen; one holds that a caller with no record base still gets byte equality only. RED-first proven: the two acceptance cases fail against the pristine checker (one on the comparison, one on the new argument) and pass with it. test_issue_records 114 passed, test_agent_issue_index 13 passed, check-agent-record rc=0 (ENGINE=179 MODEL=384 QUANT=87 KERNEL=60 BACKEND=90), invalid records 4 before and 4 after against the committed archive. ISSUE-LOCAL-01M3NH23E78PEXCJF4HW1XHQQ3 FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
… wherever a record quotes the archive validate_issue_record ran the frozen-evidence comparison -- the quoted archive line must equal the line the record declares, and must be about this record -- only in the _intake branch. Row-owned and _owed records that carry a Frozen archive evidence block were never compared: any of the 822 non-intake blocks could drift from the archive, swap a URL, or quote another issue's row, and every gate would stay green. Intake was an island; the block's name promised more than the code delivered. The comparison and the identity requirement are hoisted into _quoted_evidence_errors and applied in the _owed and row-owned branches as well. Absence of the block stays legal everywhere except _intake (456 row-owned records legitimately have none); presence is not. The identity half reuses _archive_row_owner, the same check intake already had, so the rule is now one rule in one place: quote the line you declare, modulo the relative-link rebase your directory forces (mudler#3350), and quote YOUR row. MEASURED on the restored archive (mudler#3351): 831 records carry a block, 831 resolve under the mudler#3350 comparison with the record directory as base, 831 quote a line carrying their own GitHub number, 0 violations. The ratchet adds no new red today while closing the drift door. The comparison also strips a trailing CR from the archived line before comparing. The committed blob is LF; a Windows working copy checks the archive out CRLF, and without this the same tree answered green under CI and red on a Windows checkout -- even for the 9 intake records CI greens. Byte equality must answer the CONTENT of the line, not which checkout read it. TESTS: nine cases in TestFrozenEvidenceEnforcement, red-first -- before the change six fail (the row-owned acceptance, the owed acceptance, the no-base strictness, the CRLF case) and the three rejections hold, so the rejections are not new. test_issue_records 123 passed, test_agent_issue_index 13 passed, check-agent-record rc=0 (ENGINE=179 MODEL=384 QUANT=87 KERNEL=60 BACKEND=90). Stacked on mudler#3351 (the 27 restored rows) and mudler#3350 (the rebase comparison): on a truncated archive this change would red 43 records, which is why the data half lands first. ISSUE-LOCAL-01M3NSTDJSHREP8HCX83K5KXV0 FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
…re in one run, not the first Both record validators stopped at the first invalid input, so one run revealed one defect and each fix exposed the next only after another run. The ORPHAN-MODEL-ROWS repair met exactly this: three separate defects (two one-cell-short rows and an ownerless SPIKE row) hid behind one another, and the tool output never named more than the first. The loader's first file alphabetically is an _intake record nobody is editing, which made the pattern worse: the error a maintainer saw rarely belonged to the row they touched. Two changes, one per validator. scripts/agent-issue-index.py load_local_files now collects every record that fails parse or validation and raises one IssueRecordError naming every file with its reason, instead of raising at the first. A file that fails to read joins the same list. The record collection is only returned when every file is valid, so callers guarded by the exception see no new empty-collection state. check-agent-record.py parse_claim_rows now REPORTS a malformed row and still parses it, instead of dropping it: the ratchet counts what is on disk, duplicate detection sees the duplicate, and the per-row contract checks can add their own findings to the same run. A malformed row parses with an empty state so no state-conditional contract fires on it; the reported defects stay the shape ones. Both consumers of parse_claim_rows stay correct: audit-live-rows and check-gate-commands collect parse errors and raise before classifying rows, so their censuses only become more complete, and claim-view already rejects on any collected error. Tests: LoadLocalFilesReportsEveryInvalidFile in tests/scripts/test_agent_issue_index.py holds the two-files-one-run contract and the clean-load path. Two cases in tests/scripts/test_agent_record.py hold that a malformed row no longer corrupts the ratchet count and that a malformed duplicate is reported as both malformed AND duplicate in one run. MEASURED on this tree: check-agent-record rc=0 (ENGINE=179 MODEL=384 QUANT=87 KERNEL=60 BACKEND=90); test_agent_issue_index 13/13; test_audit_live_rows 58/58; claim-view --check rc=0. On a clean checkout of d15b1cc the test_agent_record suite fails identically (26 failures, 32 errors, same case set), so this change introduces no regression there; those are the deleted-helper class the RECORD-ANCHOR-UNWIRED branch repairs. ISSUE-LOCAL-01M3NC14GE995V9E6F7GTYSQJ3 FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fix(GATE-ISSUE-INDEX-TABLE-SHAPE): record checkers report every failure in one run, not the first
Both record validators stopped at the first invalid input, so one
run revealed one defect and each fix exposed the next only after
another run. The ORPHAN-MODEL-ROWS repair met exactly this: three
separate defects (two one-cell-short rows and an ownerless SPIKE
row) hid behind one another, and the tool output never named more
than the first. The loader's first file alphabetically is an
_intake record nobody is editing, which made the pattern worse:
the error a maintainer saw rarely belonged to the row they touched.
Two changes, one per validator.
scripts/agent-issue-index.py load_local_files now collects every
record that fails parse or validation and raises one IssueRecordError
naming every file with its reason, instead of raising at the first.
A file that fails to read joins the same list. The record collection
is only returned when every file is valid, so callers guarded by the
exception see no new empty-collection state.
check-agent-record.py parse_claim_rows now REPORTS a malformed row
and still parses it, instead of dropping it: the ratchet counts what
is on disk, duplicate detection sees the duplicate, and the per-row
contract checks can add their own findings to the same run. A
malformed row parses with an empty state so no state-conditional
contract fires on it; the reported defects stay the shape ones.
Both consumers of parse_claim_rows stay correct: audit-live-rows
and check-gate-commands collect parse errors and raise before
classifying rows, so their censuses only become more complete, and
claim-view already rejects on any collected error.
Tests: LoadLocalFilesReportsEveryInvalidFile in
tests/scripts/test_agent_issue_index.py holds the two-files-one-run
contract and the clean-load path. Two cases in
tests/scripts/test_agent_record.py hold that a malformed row no
longer corrupts the ratchet count and that a malformed duplicate is
reported as both malformed AND duplicate in one run.
MEASURED on THIS tree (the stack below plus this commit):
check-agent-record rc=0 (ENGINE=179 MODEL=384 QUANT=87 KERNEL=60
BACKEND=90); test_issue_records 123 passed; test_agent_issue_index 11
passed, 3 subtests passed; test_audit_live_rows 58/58; claim-view
--check rc=0; commit-style and commit-trailers rc=0 over the range.
agent-issue-index --check now reports ALL FOUR pre-existing invalid
records in one run (three non-canonical-row records and one State:
PARTIAL), where every prior tree stopped at the first. test_agent_record
is 74 failed / 82 passed here against 73 / 82 at pristine d15b1cc
(pytest 9.1.1; the original session's unittest runner counted the same
case set as 26 failures + 32 errors). The one-case delta is
SUBFAILED(.agents/completed/issue-index.md start=1973): the restored
archive row for #2317 cites check-agent-record.py:1973, a citation
ALREADY dangling at pristine -- the tracked record ISSUE-GH-2317 cites
the same moved line and fails the same test at d15b1cc -- so the
restore surfaces a second copy of an existing dangling citation rather
than breaking a live one; INDEX_PREAMBLE no longer exists in the
checker. Editing that citation means editing a frozen archive row and
its quoting record in lockstep, which is a records fix for the
ENG-RECORD-CONFLICT-SURFACES owner, not a checker change; deliberately
out of scope here. #3347's own checker edits add ZERO new failures:
enforce-tree vs rebase-tree fail sets are identical, and the suite's
citation-floor test confirms the parse_claim_rows padding kept every
tracked anchor stable.
THIS PR IS A STACK OF FOUR COMMITS, squashed by the merge. The landed
message below describes the report-all commit; the stack carries the
frozen-evidence work with it, in order: the 27 dropped archive rows
restored (#3351, 617730b), the relative-link rebase comparison
(#3350, 957be3a), the enforcement of the frozen-evidence contract
outside _intake (#3352's tip, 9e17af3), and this report-all change
(41519cc). If #3352 merges first, this branch reduces to the single
report-all commit and rebases clean; the tree here already contains
the whole stack, and the four PRs are mutually consistent.
ISSUE-LOCAL-01M3NC14GE995V9E6F7GTYSQJ3
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]