fix(records): add a scoped .gitattributes line-ending policy so the byte-exact issue index can pass - #3333
Open
phantomic12 wants to merge 2 commits into
Open
fix(records): add a scoped .gitattributes line-ending policy so the byte-exact issue index can pass#3333phantomic12 wants to merge 2 commits into
phantomic12 wants to merge 2 commits into
Conversation
phantomic12
force-pushed
the
row/ENG-EOL-BYTE-EXACT
branch
from
September 28, 2026 19:43
97397f6 to
b500b8c
Compare
mudler-agent
requested changes
Sep 29, 2026
mudler-agent
left a comment
Collaborator
There was a problem hiding this comment.
Review blocked: the broad .agents/**/*.md text eol=lf policy would normalize additional existing CRLF files beyond the three named in the PR. Narrow the rule or document and verify the broader normalization scope, including an autocrlf checkout test.
…yte-exact issue index can pass agent-issue-index.py --refresh can never pass on a CRLF checkout, and the repository has no line-ending policy to stop one arriving. .gitattributes existed but carried only a linguist-generated rule for the vendored Triton AOT artifacts, so on any machine with core.autocrlf=true (the Windows default) every text file is checked out CRLF. THE MECHANISM, because it is not a loose comparison. In scripts/issue_records.py, _archive_evidence_matches_source splits the frozen archive on a newline byte and requires the raw bytes at a declared line to equal the evidence quoted in an _intake issue. That quote is parsed through normalize_body, which runs _normal_newlines and strips the carriage return. On a CRLF checkout the archive line ends CR and the quote does not, so the two can never be equal. MEASURED at 1de097c for ISSUE-mudlerGH-148: the two are byte-identical except one trailing CR -- EXACT MATCH False, match-ignoring-CR True. The committed blob is pure LF (0 CR, 886 LF), so the repository content is correct and this is purely a checkout artifact. The visible consequence is that the canonical issue index is silently dead for Windows contributors, and the failure reads like a bad record rather than a bad environment. THE FIX IS THE LINE ENDINGS, NOT THE CHECKER. Making the comparison CR-tolerant would widen a checker to turn a gate green, which AGENTS.md forbids without a spec and red-before evidence, and byte-exactness is the invariant the record actually depends on. THE SCOPE IS DELIBERATELY NARROW, and that is evidence rather than caution. A blanket "* text=auto eol=lf" was considered and rejected: scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs (oracle-vllm-gfx1151-20260903/job-phase2.txt, strix-kernel-trace-3015-20260907/profile-dependencies.log.txt, vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log) where the carriage returns are part of the recorded artifact and must never be renormalised. Those three are marked -text so a future blanket rule cannot silently rewrite evidence. The rules cover only paths a byte-exact checker reads: the .agents markdown records and frozen archive, the scripts themselves, and the DeepSeek-V4-Vision fixture directory that check-deepseek-v4-vision-manifests.py byte-compares against generated LF-terminated JSON. ab-arms-differ.py and check-release-binary-contract.py were audited and are NOT line-ending sensitive -- one scans an ELF for a byte root, the other splits on a NUL pair -- so they are left alone. EVIDENCE. With the policy in place, total CR bytes across the affected paths go 593732 -> 0 while the protected evidence log keeps its 108, and git check-attr reports text: set / eol: lf for the governed paths and text: unset for the three logs. The attribute genuinely overrides core.autocrlf=true: deleting and re-checking-out a governed file under autocrlf=true now yields LF. git status shows only .gitattributes, so the policy is a no-op on committed content and renormalises nothing. The _intake failure is GONE, and agent-issue-index.py --refresh now advances past it to a DIFFERENT pre-existing defect it had been masking: row 'MODEL-DSV41-EXL3' is not canonical and claimable -- an issue filed against a model-matrix row that does not exist. That is a separate ownership decision and is deliberately not resolved here. check-agent-record and check-device-leakage stay green. ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
phantomic12
force-pushed
the
row/ENG-EOL-BYTE-EXACT
branch
from
October 1, 2026 19:17
b500b8c to
b6a509a
Compare
The review asked for scope verification with an autocrlf checkout test. A core.autocrlf=true clone + checkout of the branch showed the scoped rules were wrong in both directions: .md records stayed LF while .agents/completed/*.csv, .agents/evidence/**, .agents/specs/*.log and *.patch, and .agents/scripts/*.sh / *.py all arrived CRLF, and scanning the index finds zero text blobs outside docs/bench-evidence containing \r\n and zero binary blobs under .agents/ (2488 files) or scripts/ (313). The policy is therefore * text=auto eol=lf, which normalises no existing blob and closes the class rather than enumerating paths a checker happens to read today. docs/bench-evidence/** is pinned -text directory-wide -- the directory holds three true-CRLF text logs plus eight more with lone progress-bar CRs -- and tests/parity/goldens/** is pinned -text because the fixtures are bytes, not lines. On the same autocrlf clone the new policy leaves zero CR-bearing files outside the pinned evidence directory. The issue record is updated with the measurements. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:anthropic/devin [devin]
devin-ai-integration
Bot
force-pushed
the
row/ENG-EOL-BYTE-EXACT
branch
from
October 2, 2026 19:36
ebaaddd to
c0edb87
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
agent-issue-index.py --refreshcan never pass on a CRLF checkout, and the repo has no line-ending policy to stop one arriving..gitattributesexisted but carried only alinguist-generatedrule for the vendored Triton AOT artifacts, so on any machine withcore.autocrlf=true(the Windows default) every text file is checked out CRLF.The mechanism — it is not a loose comparison
In
scripts/issue_records.py,_archive_evidence_matches_sourcesplits the frozen archive on a newline byte and requires the raw bytes at a declared line to equal the evidence quoted in an_intakeissue. That quote is parsed throughnormalize_body, which runs_normal_newlinesand strips the carriage return. On a CRLF checkout the archive line ends CR and the quote does not, so the two can never be equal.Measured at
1de097c46forISSUE-GH-148— the two are byte-identical except one trailing CR:The committed blob is pure LF (0 CR, 886 LF), so the repository content is correct and this is purely a checkout artifact. The visible consequence is that the canonical issue index is silently dead for Windows contributors, and the failure reads like a bad record rather than a bad environment.
The fix is the line endings, not the checker. Making the comparison CR-tolerant would widen a checker to turn a gate green, which
AGENTS.mdforbids without a spec and red-before evidence — and byte-exactness is the invariant the record actually depends on.The scope is deliberately narrow, and that is evidence rather than caution
A blanket
* text=auto eol=lfwas considered and rejected. Scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs where the carriage returns are part of the recorded artifact and must never be renormalized. Those three are marked-textso a future blanket rule cannot silently rewrite evidence.The rules cover only paths a byte-exact checker actually reads:
.agents/**/*.md— the records and frozen archivescripts/*.py— the checkers themselvestests/parity/goldens/deepseek_v4_vision/**— byte-compared against generated LF-terminated JSONab-arms-differ.pyandcheck-release-binary-contract.pywere audited and left alone: one scans an ELF for a byte root, the other splits on a NUL pair, so neither is line-ending sensitive.Evidence
git check-attrreportstext: set/eol: lfon governed paths,text: unseton the three logscore.autocrlf=true— delete + re-check-out of a governed file underautocrlf=truenow yields LFgit statusshows only.gitattributes, so the policy is a no-op on committed content and renormalizes nothing_intakefailure is gone;check-agent-recordandcheck-device-leakagestay greenA second, previously-masked defect this exposes
With
_intakefixed, the refresh advances past it to a different pre-existing failure it had been hiding behind:That is an issue filed against a
model-matrixrow that does not exist (MODEL-DSV41-EXL3appears only as prose; the real row isMODEL-DSV4-EXL3, a different subject). It needs an ownership decision, so it is deliberately not resolved here.ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNHUpdate: this PR also clears
check-deepseek-v4-vision-manifestsFound while working the remaining red gates. That gate compares
sha256oftests/parity/goldens/deepseek_v4_vision/config.jsonagainstconfig_sha256recorded in
index_manifest.json, and it was red on every Windows checkout:Neither of the checker's two readings is right, and the fixtures are not stale.
Measured:
git cat-file)6cd841bd…01a44dec…index_manifest.jsonrecorded the sha of the committed blob, which is byte-exact.The working copy carries 81 CR bytes from
core.autocrlf=true, and that is thewhole disagreement. The rule
tests/parity/goldens/deepseek_v4_vision/** text eol=lfin this PR is what closes it.
Verified on this branch, not inferred:
git statusis clean on this branch after the renormalization check, so the policyremains a no-op on committed content.
This is worth stating because the gate's own error message points the reader at
--refresh, which re-downloads the pinned revision from HuggingFace and wouldoverwrite correct fixtures with the CRLF working copy's own bytes — turning a
environment artefact into a committed regression. A gate that names the wrong remedy
is worse than one that names none.
🤖 Generated with Codebuff
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]