From b6a509a65c9ef1e3feed327a3e28e0e1e73e7135 Mon Sep 17 00:00:00 2001 From: Yoav Date: Sun, 27 Sep 2026 19:19:45 -0400 Subject: [PATCH 1/2] fix(records): add a scoped .gitattributes line-ending policy so the byte-exact issue index can pass agent-issue-index.py --refresh can never pass on a CRLF checkout, and the repository has no line-ending policy to stop one arriving. .gitattributes existed but carried only a linguist-generated rule for the vendored Triton AOT artifacts, so on any machine with core.autocrlf=true (the Windows default) every text file is checked out CRLF. THE MECHANISM, because it is not a loose comparison. In scripts/issue_records.py, _archive_evidence_matches_source splits the frozen archive on a newline byte and requires the raw bytes at a declared line to equal the evidence quoted in an _intake issue. That quote is parsed through normalize_body, which runs _normal_newlines and strips the carriage return. On a CRLF checkout the archive line ends CR and the quote does not, so the two can never be equal. MEASURED at 1de097c46 for ISSUE-GH-148: the two are byte-identical except one trailing CR -- EXACT MATCH False, match-ignoring-CR True. The committed blob is pure LF (0 CR, 886 LF), so the repository content is correct and this is purely a checkout artifact. The visible consequence is that the canonical issue index is silently dead for Windows contributors, and the failure reads like a bad record rather than a bad environment. THE FIX IS THE LINE ENDINGS, NOT THE CHECKER. Making the comparison CR-tolerant would widen a checker to turn a gate green, which AGENTS.md forbids without a spec and red-before evidence, and byte-exactness is the invariant the record actually depends on. THE SCOPE IS DELIBERATELY NARROW, and that is evidence rather than caution. A blanket "* text=auto eol=lf" was considered and rejected: scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs (oracle-vllm-gfx1151-20260903/job-phase2.txt, strix-kernel-trace-3015-20260907/profile-dependencies.log.txt, vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log) where the carriage returns are part of the recorded artifact and must never be renormalised. Those three are marked -text so a future blanket rule cannot silently rewrite evidence. The rules cover only paths a byte-exact checker reads: the .agents markdown records and frozen archive, the scripts themselves, and the DeepSeek-V4-Vision fixture directory that check-deepseek-v4-vision-manifests.py byte-compares against generated LF-terminated JSON. ab-arms-differ.py and check-release-binary-contract.py were audited and are NOT line-ending sensitive -- one scans an ELF for a byte root, the other splits on a NUL pair -- so they are left alone. EVIDENCE. With the policy in place, total CR bytes across the affected paths go 593732 -> 0 while the protected evidence log keeps its 108, and git check-attr reports text: set / eol: lf for the governed paths and text: unset for the three logs. The attribute genuinely overrides core.autocrlf=true: deleting and re-checking-out a governed file under autocrlf=true now yields LF. git status shows only .gitattributes, so the policy is a no-op on committed content and renormalises nothing. The _intake failure is GONE, and agent-issue-index.py --refresh now advances past it to a DIFFERENT pre-existing defect it had been masking: row 'MODEL-DSV41-EXL3' is not canonical and claimable -- an issue filed against a model-matrix row that does not exist. That is a separate ownership decision and is deliberately not resolved here. check-agent-record and check-device-leakage stay green. ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff] --- .../ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md | 29 +++++++++++++++++++ .gitattributes | 29 +++++++++++++++++++ 2 files changed, 58 insertions(+) create mode 100644 .agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md diff --git a/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md b/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md new file mode 100644 index 0000000000..8efdfadb67 --- /dev/null +++ b/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md @@ -0,0 +1,29 @@ +ID: ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH +Title: No .gitattributes line-ending policy: agent-issue-index.py --refresh can never pass on a CRLF checkout +Row: BACKEND-VULKAN +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-27 +Updated: 2026-09-27 +Closed: - + +## Problem + +The repository has NO line-ending policy. .gitattributes exists but contains only a linguist-generated rule for the vendored Triton AOT artifacts, so on any contributor machine with core.autocrlf=true (the default on Windows) every text file is checked out CRLF. + +Several tools compare file content as EXACT BYTES, and one of them is now structurally unable to pass on such a checkout. scripts/issue_records.py _archive_evidence_matches_source splits the frozen archive on b newline and requires the raw bytes at a declared line to equal the evidence quoted in an _intake issue. parse_issue_file normalises the quoted evidence through normalize_body, which runs _normal_newlines and strips the carriage return. So on a CRLF checkout the archive line ends CR and the quote does not, and they can never be equal. + +MEASURED 2026-09-27 on a Windows checkout at 1de097c46, for ISSUE-GH-148: the archive line and the quoted evidence are byte-identical except for one trailing CR, EXACT MATCH False, match-ignoring-CR True. The committed blob is pure LF (0 CR, 886 LF in .agents/completed/issue-index.md), so the repository content is correct and this is purely a checkout artifact. Consequence: scripts/agent-issue-index.py --refresh fails permanently for Windows contributors, which silently disables the canonical issue index, and the failure reads like a record defect rather than an environment one. + +The fix is the line endings, NOT the checker. Making the comparison CR-tolerant would widen a checker to make a gate green, which AGENTS.md forbids without a spec and red-before evidence, and byte-exactness is the invariant the record actually depends on. + +SCOPE IS DELIBERATELY NARROW, and that is an evidence-backed choice rather than caution. A blanket * text=auto eol=lf was considered and rejected: scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs (docs/bench-evidence/oracle-vllm-gfx1151-20260903/job-phase2.txt, docs/bench-evidence/strix-kernel-trace-3015-20260907/profile-dependencies.log.txt, docs/bench-evidence/vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log) where the carriage returns are part of the recorded artifact and must never be renormalised. Those are marked -text so a future blanket rule cannot silently mutate evidence. + +The rules cover only the paths a byte-exact checker reads: the .agents markdown records and frozen archive, the scripts themselves (whose literals are compared), and the DeepSeek-V4-Vision fixture directory that check-deepseek-v4-vision-manifests.py byte-compares against generated LF-terminated JSON. ab-arms-differ.py and check-release-binary-contract.py were audited and are NOT line-ending sensitive (one scans an ELF for a byte root, the other splits on a NUL pair), so they are left alone. + +## Resolution + +- diff --git a/.gitattributes b/.gitattributes index d67e7e7212..62cc29df16 100644 --- a/.gitattributes +++ b/.gitattributes @@ -1,3 +1,32 @@ # Vendored Triton AOT artifacts are generated code (embedded cubins): collapse # in diffs/review and exclude from language stats. Regen: scripts/regen-triton-aot.sh src/vt/cuda/triton_aot_vendored/** linguist-generated=true + +# --------------------------------------------------------------------------- +# Line endings. Scoped, NOT a blanket `* text=auto eol=lf`. +# +# Several checkers compare file content as EXACT BYTES, so a CRLF checkout +# makes them fail on content that is correct. The worst case is +# `_archive_evidence_matches_source` in scripts/issue_records.py: it splits the +# frozen archive on a newline byte and requires the raw bytes at a declared +# line to equal evidence quoted in an _intake issue, but that quote is +# normalised through `normalize_body`, which strips the carriage return. On a +# CRLF checkout the two differ by exactly one CR, so +# `agent-issue-index.py --refresh` can never pass -- permanently, and for a +# reason that reads like a bad record rather than a bad environment. +# +# These rules cover only the paths a byte-exact checker actually reads. +# `ab-arms-differ.py` and `check-release-binary-contract.py` were audited and +# are NOT line-ending sensitive (one scans an ELF for a byte root, the other +# splits on a NUL pair), so they are deliberately left alone. +.agents/**/*.md text eol=lf +scripts/*.py text eol=lf +tests/parity/goldens/deepseek_v4_vision/** text eol=lf + +# Captured bench evidence, where a CR is part of the recorded artifact and must +# survive verbatim. These are the ONLY tracked text blobs in the tree that +# contain CRLF (3 of 7114, measured at 1de097c46); marking them -text means a +# future blanket policy cannot silently rewrite evidence. +docs/bench-evidence/oracle-vllm-gfx1151-20260903/job-phase2.txt -text +docs/bench-evidence/strix-kernel-trace-3015-20260907/profile-dependencies.log.txt -text +docs/bench-evidence/vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log -text From c0edb8767181e6fa8f6018e67edd4788bed195c2 Mon Sep 17 00:00:00 2001 From: ULTRAPHANTOM102 Date: Fri, 2 Oct 2026 19:35:59 +0000 Subject: [PATCH 2/2] fix(records): widen the line-ending policy to the measured CRLF surface The review asked for scope verification with an autocrlf checkout test. A core.autocrlf=true clone + checkout of the branch showed the scoped rules were wrong in both directions: .md records stayed LF while .agents/completed/*.csv, .agents/evidence/**, .agents/specs/*.log and *.patch, and .agents/scripts/*.sh / *.py all arrived CRLF, and scanning the index finds zero text blobs outside docs/bench-evidence containing \r\n and zero binary blobs under .agents/ (2488 files) or scripts/ (313). The policy is therefore * text=auto eol=lf, which normalises no existing blob and closes the class rather than enumerating paths a checker happens to read today. docs/bench-evidence/** is pinned -text directory-wide -- the directory holds three true-CRLF text logs plus eight more with lone progress-bar CRs -- and tests/parity/goldens/** is pinned -text because the fixtures are bytes, not lines. On the same autocrlf clone the new policy leaves zero CR-bearing files outside the pinned evidence directory. The issue record is updated with the measurements. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:anthropic/devin [devin] --- .../ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md | 9 +++-- .gitattributes | 34 +++++++++++-------- 2 files changed, 26 insertions(+), 17 deletions(-) diff --git a/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md b/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md index 8efdfadb67..db7e10decf 100644 --- a/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md +++ b/.agents/issues/BACKEND-VULKAN/ISSUE-LOCAL-01M3JJ9CE2WH0KE640HQFGACNH.md @@ -7,7 +7,7 @@ GitHub: - Mirror: PENDING Availability: FULL Created: 2026-09-27 -Updated: 2026-09-27 +Updated: 2026-10-02 Closed: - ## Problem @@ -20,9 +20,12 @@ MEASURED 2026-09-27 on a Windows checkout at 1de097c46, for ISSUE-GH-148: the ar The fix is the line endings, NOT the checker. Making the comparison CR-tolerant would widen a checker to make a gate green, which AGENTS.md forbids without a spec and red-before evidence, and byte-exactness is the invariant the record actually depends on. -SCOPE IS DELIBERATELY NARROW, and that is an evidence-backed choice rather than caution. A blanket * text=auto eol=lf was considered and rejected: scanning all 7114 tracked text blobs finds exactly 3 containing CRLF, and all three are captured bench-evidence logs (docs/bench-evidence/oracle-vllm-gfx1151-20260903/job-phase2.txt, docs/bench-evidence/strix-kernel-trace-3015-20260907/profile-dependencies.log.txt, docs/bench-evidence/vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log) where the carriage returns are part of the recorded artifact and must never be renormalised. Those are marked -text so a future blanket rule cannot silently mutate evidence. +SCOPE REVISED 2026-10-02 after review asked for the checkout test. The first cut covered only .agents/**/*.md, scripts/*.py and the DeepSeek-V4-Vision fixture dir, reasoning that those were the paths a byte-exact checker reads. A core.autocrlf=true clone + checkout of the branch showed that framing was wrong in BOTH directions: -The rules cover only the paths a byte-exact checker reads: the .agents markdown records and frozen archive, the scripts themselves (whose literals are compared), and the DeepSeek-V4-Vision fixture directory that check-deepseek-v4-vision-manifests.py byte-compares against generated LF-terminated JSON. ab-arms-differ.py and check-release-binary-contract.py were audited and are NOT line-ending sensitive (one scans an ELF for a byte root, the other splits on a NUL pair), so they are left alone. +- The record surface is wider than *.md. On the autocrlf checkout the .md rows stayed LF, but .agents/completed/*.csv manifests, .agents/evidence/** JSON/stdout captures, .agents/specs/*.log and *.patch, and .agents/scripts/*.sh / *.py all arrived CRLF -- the same byte-exact hazard one directory over. +- The policy normalises nothing that exists. Scanning the index finds ZERO text blobs outside docs/bench-evidence containing \r\n, and ZERO binary blobs under .agents/ (2488 files) or scripts/ (313), so a whole-tree * text=auto eol=lf is safe and rewrites no stored content -- it only stops the class from arriving. + +The policy is therefore `* text=auto eol=lf` with two byte-preserving pins. docs/bench-evidence/** is pinned -text directory-wide, not file-by-file: the directory carries three true-CRLF text logs (job-phase2.txt, profile-dependencies.log.txt, gen-20260903T012806Z.log) PLUS eight more text logs with lone progress-bar CRs (limb3-strict-gate gen-*.log, job-phase3.txt, q4km-neartie harness-stdout-*.txt), and any of them committed from an autocrlf checkout would be rewritten. tests/parity/goldens/** is pinned -text because the .npy/.i32/.raw fixtures and generated manifests are bytes, not lines. After the change, the same autocrlf clone produces zero CR-bearing files outside the pinned evidence directory. ## Resolution diff --git a/.gitattributes b/.gitattributes index 62cc29df16..31842d22b3 100644 --- a/.gitattributes +++ b/.gitattributes @@ -3,7 +3,7 @@ src/vt/cuda/triton_aot_vendored/** linguist-generated=true # --------------------------------------------------------------------------- -# Line endings. Scoped, NOT a blanket `* text=auto eol=lf`. +# Line endings. # # Several checkers compare file content as EXACT BYTES, so a CRLF checkout # makes them fail on content that is correct. The worst case is @@ -15,18 +15,24 @@ src/vt/cuda/triton_aot_vendored/** linguist-generated=true # `agent-issue-index.py --refresh` can never pass -- permanently, and for a # reason that reads like a bad record rather than a bad environment. # -# These rules cover only the paths a byte-exact checker actually reads. -# `ab-arms-differ.py` and `check-release-binary-contract.py` were audited and -# are NOT line-ending sensitive (one scans an ELF for a byte root, the other -# splits on a NUL pair), so they are deliberately left alone. -.agents/**/*.md text eol=lf -scripts/*.py text eol=lf -tests/parity/goldens/deepseek_v4_vision/** text eol=lf +# Measured scope, not a guess: a `core.autocrlf=true` clone + checkout of this +# revision was compared against the LF worktree. The record surface is wider +# than *.md -- .agents/completed/*.csv manifests, .agents/evidence/** JSON and +# stdout captures, .agents/specs/*.log and *.patch, and .agents/scripts/*.sh / +# *.py all arrived CRLF while the .md rows stayed LF. There are ZERO binary +# blobs under .agents/ or scripts/ (2488 and 313 files scanned), so a +# directory-wide text rule is safe; and there are ZERO tracked text blobs +# outside docs/bench-evidence that contain \r\n, so this policy normalises +# nothing that exists today -- it only stops the class from arriving. +* text=auto eol=lf # Captured bench evidence, where a CR is part of the recorded artifact and must -# survive verbatim. These are the ONLY tracked text blobs in the tree that -# contain CRLF (3 of 7114, measured at 1de097c46); marking them -text means a -# future blanket policy cannot silently rewrite evidence. -docs/bench-evidence/oracle-vllm-gfx1151-20260903/job-phase2.txt -text -docs/bench-evidence/strix-kernel-trace-3015-20260907/profile-dependencies.log.txt -text -docs/bench-evidence/vllm-gguf-plugin-thor-20260903/gen-20260903T012806Z.log -text +# survive verbatim. Pinned directory-wide rather than file-by-file: the tree +# carries more CR-bearing text logs than the three files that contain \r\n +# (progress-bar carriage returns in limb3-strict-gate and q4km-neartie logs), +# and any of them committed from an autocrlf checkout would be rewritten. +docs/bench-evidence/** -text + +# Parity goldens are fixture bytes -- .npy/.i32/.raw tensors and generated +# manifests -- never line-ending-normalised. +tests/parity/goldens/** -text