Skip to content

fix(records): add a scoped .gitattributes line-ending policy so the byte-exact issue index can pass - #3333

Open
phantomic12 wants to merge 2 commits into
mudler:mainfrom
phantomic12:row/ENG-EOL-BYTE-EXACT
Open

phantomic12 wants to merge 2 commits into
mudler:mainfrom
phantomic12:row/ENG-EOL-BYTE-EXACT

Conversation

@phantomic12

@phantomic12 phantomic12 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

agent-issue-index.py --refresh can never pass on a CRLF checkout, and the repo has no line-ending policy to stop one arriving.

.gitattributes existed but carried only a linguist-generated rule for the vendored Triton AOT artifacts, so on any machine with core.autocrlf=true (the Windows default) every text file is checked out CRLF.

The mechanism — it is not a loose comparison

In scripts/issue_records.py, _archive_evidence_matches_source splits the frozen archive on a newline byte and requires the raw bytes at a declared line to equal the evidence quoted in an _intake issue. That quote is parsed through normalize_body, which runs _normal_newlines and strips the carriage return. On a CRLF checkout the archive line ends CR and the quote does not, so the two can never be equal.

Measured at 1de097c46 for ISSUE-GH-148 — the two are byte-identical except one trailing CR:

EXACT MATCH:            False
match ignoring CR:      True

The committed blob is pure LF (0 CR, 886 LF), so the repository content is correct and this is purely a checkout artifact. The visible consequence is that the canonical issue index is silently dead for Windows contributors, and the failure reads like a bad record rather than a bad environment.

The fix is the line endings, not the checker. Making the comparison CR-tolerant would widen a checker to turn a gate green, which AGENTS.md forbids without a spec and red-before evidence — and byte-exactness is the invariant the record actually depends on.

The scope is deliberately narrow, and that is evidence rather than caution

A blanket * text=auto eol=lf was considered and rejected. Scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs where the carriage returns are part of the recorded artifact and must never be renormalized. Those three are marked -text so a future blanket rule cannot silently rewrite evidence.

The rules cover only paths a byte-exact checker actually reads:

  • .agents/**/*.md — the records and frozen archive
  • scripts/*.py — the checkers themselves
  • tests/parity/goldens/deepseek_v4_vision/** — byte-compared against generated LF-terminated JSON

ab-arms-differ.py and check-release-binary-contract.py were audited and left alone: one scans an ELF for a byte root, the other splits on a NUL pair, so neither is line-ending sensitive.

Evidence

  • total CR bytes across affected paths: 593732 → 0; the protected evidence log keeps its 108
  • git check-attr reports text: set / eol: lf on governed paths, text: unset on the three logs
  • the attribute genuinely overrides core.autocrlf=true — delete + re-check-out of a governed file under autocrlf=true now yields LF
  • git status shows only .gitattributes, so the policy is a no-op on committed content and renormalizes nothing
  • the _intake failure is gone; check-agent-record and check-device-leakage stay green

A second, previously-masked defect this exposes

With _intake fixed, the refresh advances past it to a different pre-existing failure it had been hiding behind:

row 'MODEL-DSV41-EXL3' is not canonical and claimable

That is an issue filed against a model-matrix row that does not exist (MODEL-DSV41-EXL3 appears only as prose; the real row is MODEL-DSV4-EXL3, a different subject). It needs an ownership decision, so it is deliberately not resolved here.

ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH


Update: this PR also clears check-deepseek-v4-vision-manifests

Found while working the remaining red gates. That gate compares sha256 of
tests/parity/goldens/deepseek_v4_vision/config.json against config_sha256
recorded in index_manifest.json, and it was red on every Windows checkout:

index manifest: config_sha256 is '6cd841bdd6702f5e2ac34671bc78047ed80817102465525ae2a41c502abbcd75',
derived         '01a44deca47a04531674210a8dae5ce68396820e517783f3337562fb7a9e060c'
1 manifest disagreement(s). Either the derivation changed and the fixtures are stale
(rerun ... --refresh against the pinned revision), or the derivation is wrong.

Neither of the checker's two readings is right, and the fixtures are not stale.
Measured:

bytes sha256 CR matches the manifest
committed blob (git cat-file) 6cd841bd… 0 yes
Windows working copy 01a44dec… 81 no

index_manifest.json recorded the sha of the committed blob, which is byte-exact.
The working copy carries 81 CR bytes from core.autocrlf=true, and that is the
whole disagreement. The rule tests/parity/goldens/deepseek_v4_vision/** text eol=lf
in this PR is what closes it.

Verified on this branch, not inferred:

$ python3 scripts/check-deepseek-v4-vision-manifests.py
ok deepseek-ai/DeepSeek-V4-Flash-Vision-Exp@86f746b36186: 72633 tensors over 48 shards,
   267 of them vision, 932786176 vision payload bytes; no network, no weight bytes read

git status is clean on this branch after the renormalization check, so the policy
remains a no-op on committed content.

This is worth stating because the gate's own error message points the reader at
--refresh, which re-downloads the pinned revision from HuggingFace and would
overwrite correct fixtures with the CRLF working copy's own bytes — turning a
environment artefact into a committed regression. A gate that names the wrong remedy
is worse than one that names none.

🤖 Generated with Codebuff

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]

@mudler-agent mudler-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review blocked: the broad .agents/**/*.md text eol=lf policy would normalize additional existing CRLF files beyond the three named in the PR. Narrow the rule or document and verify the broader normalization scope, including an autocrlf checkout test.

…yte-exact issue index can pass

agent-issue-index.py --refresh can never pass on a CRLF checkout, and the
repository has no line-ending policy to stop one arriving. .gitattributes
existed but carried only a linguist-generated rule for the vendored Triton
AOT artifacts, so on any machine with core.autocrlf=true (the Windows
default) every text file is checked out CRLF.

THE MECHANISM, because it is not a loose comparison. In
scripts/issue_records.py, _archive_evidence_matches_source splits the
frozen archive on a newline byte and requires the raw bytes at a declared
line to equal the evidence quoted in an _intake issue. That quote is
parsed through normalize_body, which runs _normal_newlines and strips the
carriage return. On a CRLF checkout the archive line ends CR and the quote
does not, so the two can never be equal. MEASURED at 1de097c for
ISSUE-mudlerGH-148: the two are byte-identical except one trailing CR --
EXACT MATCH False, match-ignoring-CR True. The committed blob is pure LF
(0 CR, 886 LF), so the repository content is correct and this is purely a
checkout artifact. The visible consequence is that the canonical issue
index is silently dead for Windows contributors, and the failure reads
like a bad record rather than a bad environment.

THE FIX IS THE LINE ENDINGS, NOT THE CHECKER. Making the comparison
CR-tolerant would widen a checker to turn a gate green, which AGENTS.md
forbids without a spec and red-before evidence, and byte-exactness is the
invariant the record actually depends on.

THE SCOPE IS DELIBERATELY NARROW, and that is evidence rather than
caution. A blanket "* text=auto eol=lf" was considered and rejected:
scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and
all three are captured bench-evidence logs
(oracle-vllm-gfx1151-20260903/job-phase2.txt,
strix-kernel-trace-3015-20260907/profile-dependencies.log.txt,
vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log) where the
carriage returns are part of the recorded artifact and must never be
renormalised. Those three are marked -text so a future blanket rule cannot
silently rewrite evidence. The rules cover only paths a byte-exact checker
reads: the .agents markdown records and frozen archive, the scripts
themselves, and the DeepSeek-V4-Vision fixture directory that
check-deepseek-v4-vision-manifests.py byte-compares against generated
LF-terminated JSON. ab-arms-differ.py and check-release-binary-contract.py
were audited and are NOT line-ending sensitive -- one scans an ELF for a
byte root, the other splits on a NUL pair -- so they are left alone.

EVIDENCE. With the policy in place, total CR bytes across the affected
paths go 593732 -> 0 while the protected evidence log keeps its 108, and
git check-attr reports text: set / eol: lf for the governed paths and
text: unset for the three logs. The attribute genuinely overrides
core.autocrlf=true: deleting and re-checking-out a governed file under
autocrlf=true now yields LF. git status shows only .gitattributes, so the
policy is a no-op on committed content and renormalises nothing. The
_intake failure is GONE, and agent-issue-index.py --refresh now advances
past it to a DIFFERENT pre-existing defect it had been masking:
row 'MODEL-DSV41-EXL3' is not canonical and claimable -- an issue filed
against a model-matrix row that does not exist. That is a separate
ownership decision and is deliberately not resolved here.

check-agent-record and check-device-leakage stay green.

ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]
@phantomic12
phantomic12 force-pushed the row/ENG-EOL-BYTE-EXACT branch from b500b8c to b6a509a Compare October 1, 2026 19:17
The review asked for scope verification with an autocrlf checkout test.
A core.autocrlf=true clone + checkout of the branch showed the scoped
rules were wrong in both directions: .md records stayed LF while
.agents/completed/*.csv, .agents/evidence/**, .agents/specs/*.log and
*.patch, and .agents/scripts/*.sh / *.py all arrived CRLF, and scanning
the index finds zero text blobs outside docs/bench-evidence containing
\r\n and zero binary blobs under .agents/ (2488 files) or scripts/
(313). The policy is therefore * text=auto eol=lf, which normalises no
existing blob and closes the class rather than enumerating paths a
checker happens to read today. docs/bench-evidence/** is pinned -text
directory-wide -- the directory holds three true-CRLF text logs plus
eight more with lone progress-bar CRs -- and tests/parity/goldens/**
is pinned -text because the fixtures are bytes, not lines. On the same
autocrlf clone the new policy leaves zero CR-bearing files outside the
pinned evidence directory. The issue record is updated with the
measurements.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:anthropic/devin [devin]
@devin-ai-integration
devin-ai-integration Bot force-pushed the row/ENG-EOL-BYTE-EXACT branch from ebaaddd to c0edb87 Compare October 2, 2026 19:36

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants