From 198c2a7018a50b9af966d1ae21e9c631d461d989 Mon Sep 17 00:00:00 2001 From: amrali-eg <32075105+amrali-eg@users.noreply.github.com> Date: Thu, 17 Sep 2026 22:47:36 +0300 Subject: [PATCH] Bump version to 1.7.0 and write release notes Bumps LineEndingNormalizer.csproj and app.manifest 1.6.0 -> 1.7.0. Adds RELEASE-NOTES-v1.7.0.md covering the UTF-32 fix, -Report preflight and atomic write, the new directory-coverage summary lines, and the repository-hygiene changes since v1.6.0. Also corrects RELEASE-NOTES-v1.6.0.md and SAFETY-AUDIT.md's v1.6.0 entry: both described an opposite-byte-order UTF-32 ambiguity test as this repository's shipped policy. Checked against git history (fix/utf32-unprovable-refusal landed as c944307, confirmed NOT an ancestor of the v1.6.0 tag via git merge-base) - v1.6.0 as actually published genuinely ships that insufficient test, not the corrected unconditional refusal. v1.7.0 is the first release containing the fix; both files now say so plainly rather than leaving a shipped defect undocumented. Local gates: build 0 warnings, 316/316 tests, detector parity clean across EC/LEN/CorpusTesters. Co-Authored-By: Claude Sonnet 5 --- docs/RELEASE-NOTES-v1.6.0.md | 12 ++ docs/RELEASE-NOTES-v1.7.0.md | 105 ++++++++++++++++++ docs/SAFETY-AUDIT.md | 6 + .../LineEndingNormalizer.csproj | 2 +- sources/LineEndingNormalizer/app.manifest | 2 +- 5 files changed, 125 insertions(+), 2 deletions(-) create mode 100644 docs/RELEASE-NOTES-v1.7.0.md diff --git a/docs/RELEASE-NOTES-v1.6.0.md b/docs/RELEASE-NOTES-v1.6.0.md index 20d1de3..0abdbbb 100644 --- a/docs/RELEASE-NOTES-v1.6.0.md +++ b/docs/RELEASE-NOTES-v1.6.0.md @@ -1,5 +1,17 @@ # LineEndingNormalizer v1.6.0 +**Correction (added in v1.7.0):** the Safety section below describes and this +version genuinely ships an opposite-byte-order ambiguity test for BOM-less +UTF-32, reporting `AmbiguousBomlessUtf32`. That test is insufficient: a +BOM-less UTF-16 file can misdetect as UTF-32 and pass this check, because the +opposite UTF-32 order need not fail to decode a file that is not really +UTF-32 at all. **v1.6.0 as published can convert such a file, silently +corrupting it.** v1.7.0 replaces this with an unconditional refusal +(`UnprovableBomlessUtf32`) that closes the gap; see +[SAFETY.md](SAFETY.md#bom-less-utf-32) and +[RELEASE-NOTES-v1.7.0.md](RELEASE-NOTES-v1.7.0.md). This file is left +otherwise unedited as a historical record of what was intended at the time. + Adds one narrow new safety refusal, closes a test-coverage gap in the existing backup verification, updates the legacy-detection dependency, and reorganizes the repository into the same `sources/` layout EncodingChecker already uses. diff --git a/docs/RELEASE-NOTES-v1.7.0.md b/docs/RELEASE-NOTES-v1.7.0.md new file mode 100644 index 0000000..34ea3fa --- /dev/null +++ b/docs/RELEASE-NOTES-v1.7.0.md @@ -0,0 +1,105 @@ +# LineEndingNormalizer v1.7.0 + +Fixes a real BOM-less UTF-32 misdetection gap left open by v1.6.0's guard, adds +two coverage/safety checks around `-Report`, and adds two new coverage +counters to the run summary. Everything else is repository hygiene. + +It is a minor release rather than a patch because the CLI's summary output +gains two new lines and a new exit-3 refusal path, even though the UTF-32 fix +alone would have justified at least a patch on its own. + +## Safety + +- **BOM-less UTF-32 is now refused unconditionally**, closing a real gap in + v1.6.0's guard. That guard refused UTF-32 only when the *opposite* byte + order also strictly decoded the whole file — sufficient for UTF-16, but not + for UTF-32: a BOM-less UTF-16 file with one character per line puts a C0 + control in every second code unit, so each resulting four-byte group is an + in-range, unassigned UTF-32 scalar and the file misdetects as UTF-32LE, + while the *opposite* UTF-32 order correctly fails to decode that same file. + v1.6.0's opposite-order test therefore called such a file unambiguous and + would have converted it, rewriting real UTF-16 content as if it were + UTF-32 — silent corruption, not merely a wrong refusal. Every BOM-less + UTF-32 file is now refused with reason code `UnprovableBomlessUtf32` + (replacing `AmbiguousBomlessUtf32`), matching EncodingChecker's own + historically-hardened policy for the same reason. See + [SAFETY.md](SAFETY.md#bom-less-utf-32) for the full explanation, and + [RELEASE-NOTES-v1.6.0.md](RELEASE-NOTES-v1.6.0.md) for the correction to + that release's notes. + + **The cost is real:** ordinary, unambiguous BOM-less UTF-32 is refused too, + with no override — LEN has no source-encoding flag, so there is no + documented way through this refusal short of adding a byte-order mark. + +- **An unwritable `-Report` path is now caught before any file is converted.** + Previously, a missing parent directory or a path naming an existing + directory was discovered only after every source file had already been + converted, leaving the requested report absent regardless. This is now + checked up front and returns exit code 3 without touching any file. + +- **`-Report` is now written atomically.** The report file was previously + truncated in place; an interrupted write could leave it empty or + half-written and destroy whatever prior report was already there. It is now + written to a temporary file and installed only once complete, reusing the + same atomic-replace mechanism as file conversion. + +## Reporting + +- **Two new run-summary lines, shown only when nonzero:** `Dirs unreadable` + (a directory LEN could not list) and `Dirs skipped` (a directory excluded by + reserved name — `.git`, `bin`, `obj`, and similar). Previously an unreadable + directory produced only a stderr warning and no counter, and a + reserved-name skip was entirely silent — a run that missed part of the tree + could report a clean summary and exit 0 with no trace of the coverage loss. + `-DetectOnly` is unchanged; it has no summary to add these to. + +## Internal + +- Removed the unused `coverlet.collector` dependency (nothing in this repo + invokes coverage collection) and bumped `Microsoft.NET.Test.Sdk` (17.14.1 → + 18.10.1) and `xunit.runner.visualstudio` (3.1.4 → 4.0.0, still xUnit v2 + compatible), matching the same bump already done in EncodingChecker. + Test-tooling only. + +## Repository hygiene + +- Added `.editorconfig` (previously missing), matching EncodingChecker's: + `.cs` is UTF-8 without a BOM, CRLF. All 14 tracked `.cs` files carried a BOM + before this — LEN's actual standing convention, just undocumented — and + were re-saved BOM-less to match the new rule. Each file lost exactly the + 3-byte BOM and nothing else. +- Fixed two stale/missing `.gitignore` rules found by comparing against EC: + the publish-profile un-ignore paths still pointed at the pre-`sources/`-reorg + location, and `**/.claude/settings.local.json` (machine-specific, never + shared) was missing entirely. +- Added `AGENTS.md`, matching EncodingChecker's convention of keeping the same + project instructions available to Codex and other generic agent tooling + under both filenames. + +## Unchanged + +Conversion, backup, and replacement behavior for every case other than the +UTF-32 guard and `-Report` handling above are exactly what they were in +v1.6.0. All exit codes other than the new `-Report` preflight path, reason +codes other than the UTF-32 one, and CLI options are unchanged. + +## Known limits + +- No restore command exists. A `.bak` is independently hash-verified recovery + data, not a restore feature. +- The backup is created before the source is replaced, so a later failure can + leave a valid `.bak` beside an unchanged source. This is deliberate. +- The final destination check narrows, but does not eliminate, the race between + checking a file and replacing it. +- CSV fields are RFC 4180 quoted but are not neutralized against spreadsheet + formula interpretation. +- Establishing BOM-less UTF-16 or UTF-32 safety costs a second complete read + of the file. +- `-DetectOnly` has no `-Report` preflight check and no directory-coverage + counters; both are specific to the normalize/validate/`-WhatIf` path. +- The detector still evaluates entropy before honoring Unicode BOMs. Changing + that order is deferred pending corpus testing across all three repositories. +- LEN's own `detector-parity.yml` compares `UnicodeDetector.cs` against + EncodingChecker only; it does not compare `TextValidation.cs` and does not + include CorpusTesters. A future divergence in `TextValidation.cs` alone + would not be caught by this repository's CI. diff --git a/docs/SAFETY-AUDIT.md b/docs/SAFETY-AUDIT.md index ab73fee..377b16f 100644 --- a/docs/SAFETY-AUDIT.md +++ b/docs/SAFETY-AUDIT.md @@ -96,6 +96,12 @@ mirroring the existing UTF-16 guard) and closes a test-coverage gap in the backup hash-verification check; both rest on the regression suite, not on corpus measurement. +**Correction (added in v1.7.0):** the UTF-32 guard this version ships is +insufficient — a BOM-less UTF-16 file can misdetect as UTF-32 and pass the +opposite-order test unchanged, so v1.6.0 as published can silently corrupt +such a file. v1.7.0 replaces it with an unconditional refusal. See that +version's entry below and [RELEASE-NOTES-v1.7.0.md](RELEASE-NOTES-v1.7.0.md). + ``` commit 31dd78ac4c28a9e8713c09c49a13d2f7f8733add (annotated tag v1.6.0) project 1.6.0 manifest 1.6.0.0 binary reports 1.6.0 diff --git a/sources/LineEndingNormalizer/LineEndingNormalizer.csproj b/sources/LineEndingNormalizer/LineEndingNormalizer.csproj index 7177e06..05f2c2e 100644 --- a/sources/LineEndingNormalizer/LineEndingNormalizer.csproj +++ b/sources/LineEndingNormalizer/LineEndingNormalizer.csproj @@ -10,7 +10,7 @@ - 1.6.0 + 1.7.0 diff --git a/sources/LineEndingNormalizer/app.manifest b/sources/LineEndingNormalizer/app.manifest index 611a0a5..baac38e 100644 --- a/sources/LineEndingNormalizer/app.manifest +++ b/sources/LineEndingNormalizer/app.manifest @@ -1,6 +1,6 @@ - +