Skip to content

Commit b67bb72

Browse files
jun0claude
andcommitted
[rustjava-2026-09-23-stale-next-pointer-and-euc-kr-boundary-adopt-p0] refactor(rustjava-runtime): charset 보류 판단을 Charset 의 exhaustive match 로 옮긴다
InputStreamReader::read() 가 멀티바이트 경계 보류 바이트 수를 charset 이름 문자열 비교로 정하던 것을 Charset::bytes_to_hold_back(&[u8]) -> usize 로 옮겼다. match 에 wildcard 가 없어 새 변종은 컴파일이 막는다. 동작 불변 — UTF-8 역주사·EUC-KR 전방 주사 본문은 그대로 옮겼다. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 parent c654ae2 commit b67bb72

6 files changed

Lines changed: 83 additions & 31 deletions

File tree

‎REPORT.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,9 @@
11
# REPORT
2+
## [2026-09-23] charset 보류 판단을 `Charset` 의 exhaustive match 로 옮겼다 (rustjava-2026-09-23-stale-next-pointer-and-euc-kr-boundary-adopt-p0)
3+
- 무엇을: `InputStreamReader::read()` 의 charset 이름 문자열 비교 2곳을 `Charset::bytes_to_hold_back` 로 옮겼다(wildcard 없는 match · 동작 불변).
4+
- 왜: 채택 제안 `2026-09-23-stale-next-pointer-and-euc-kr-boundary#p0` — 새 charset 을 더하면 그 if 사슬은 조용히 빠졌다. 이제 컴파일이 막는다.
5+
- 사용자 영향: 없음(동작 불변). 다음에 멀티바이트 charset 을 더할 때 읽기 경계에서 글자가 사라지는 결함을 «잊어서» 만들 수 없다. 후속 추천 0건. 상세 = `docs/worklog/2026-09-23-charset-hold-back-exhaustive-match.{md,json}`.
6+
27
## [2026-09-23] `## 다음` 이 세 번째로 낡았고, 그 밑에서 EUC-KR 한 글자가 조용히 사라지고 있었다 (rustjava-next-slice-and-stale-next-pointer)
38
- 무엇을: `STATE.md` `## 다음` 의 모든 「다음 후보」를 `git merge-base --is-ancestor` 로 **전수 재측**해 **8건 중 7건이 이미 닫혀 있음**을 확인하고 사료로 접었으며, 살아 있는 후보만 새 `⓪` 블록에 올렸다. 재측 과정에서 유일하게 «열려 있던» 항목(④-3 `InputStreamReader` 디코더)이 실제 결함임이 드러나 **그 자리에서 고쳤다**.
49
- 왜: `LANE_IDLE rustjava`(2026-09-23 · ★**1088분** 조용 · 큐 0 · running 0). 근인은 워커가 아니라 **발권**이었고, 발권이 멈춘 이유는 `## 다음` 의 최우선 항목이 **이미 끝난 일**을 가리켰기 때문이다. ★**이 절이 그 병을 스스로 두 번 기록해 놓고 세 번째를 냈다** — 2026-09-11 에 「다음 실작업 = ③의 null-guard」로 고쳐 쓴 그 null-guard 가 **같은 날 이미 닫혀 있었다**(`6da7d66f`).

‎STATE.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@
77
(둘 다 이것보다 오래됐고 MERGEABLE/CONFLICTING 처분이 이미 걸려 있다). 겹침은 전부 **append 형 합집합**이라 해소는 기계적이다)
88

99
## 완료
10+
- [rustjava-2026-09-23-stale-next-pointer-and-euc-kr-boundary-adopt-p0] charset 보류 판단을 `Charset::bytes_to_hold_back`(wildcard 없는 match)로 옮김 · `read()` 이름 비교 0 · 동작 불변 · 변이 양방향 확인. 채택 `2026-09-23-stale-next-pointer-and-euc-kr-boundary#p0`.
1011
- [rustjava-prune-declined-followup-proposals-2026-09-21] ★**추천 후속작업 2건 기각** — 운영자 지시(2026-09-21 우선순위 정리). ★제품 코드 **0줄** · 새 제안 **0** · 검사기/CI 신설 **0**.
1112
★**닫은 둘**(검사기 다듬기 축): `2026-09-19-nonliteral-blind-spot-is-reported-not-gated#p0` · `2026-09-20-lock-script-output-order#p0`.
1213
★**남긴 둘**(런타임 축): `2026-09-20-string-array-hiding-overflows-stack#p0` · `2026-09-20-name-the-missing-bootstrap-class#p0`.
Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,17 @@
1+
{
2+
"schema": "worklog/v1",
3+
"date": "2026-09-23",
4+
"taskId": "rustjava-2026-09-23-stale-next-pointer-and-euc-kr-boundary-adopt-p0",
5+
"summary": "Behaviour-preserving refactor. InputStreamReader::read() chose how many trailing bytes to withhold at a multibyte boundary by comparing the charset name string (\"UTF-8\" / \"EUC-KR\"). That decision now lives in Charset::bytes_to_hold_back, a match over Charset with no wildcard arm, so adding a Charset variant fails to compile until someone decides its boundary rule instead of silently dropping characters.",
6+
"changes": [
7+
"rustjava-runtime/src/charset.rs: new Charset::bytes_to_hold_back(&[u8]) -> usize; UTF-8 backward scan and EUC-KR forward walk moved verbatim; Iso8859_1 | UsAscii => 0 explicitly",
8+
"rustjava-runtime/src/classes/java/io/input_stream_reader.rs: read() calls it; name-string branches removed; stale comment at the default ctor reworded"
9+
],
10+
"verification": [
11+
"DoD 10 commands rc 0; cargo test --all 592 passed 0 failed",
12+
"temp Charset variant -> E0004 non-exhaustive in bytes_to_hold_back (reverted)",
13+
"EUC-KR hold forced to 0 -> #90 boundary test red (reverted); restored -> green"
14+
],
15+
"proposals": [],
16+
"adoptedProposals": ["2026-09-23-stale-next-pointer-and-euc-kr-boundary#p0"]
17+
}
Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,12 @@
1+
# 2026-09-23 — charset 보류 술어를 `Charset` 쪽으로 (rustjava-2026-09-23-stale-next-pointer-and-euc-kr-boundary-adopt-p0)
2+
3+
채택 제안 `2026-09-23-stale-next-pointer-and-euc-kr-boundary#p0`.
4+
5+
- **무엇을**: `read()` 의 `charset == "UTF-8"` / `charset == "EUC-KR"` 분기를 `Charset::bytes_to_hold_back(&[u8]) -> usize` 로 옮겼다.
6+
match 에 `_ =>` 가 없다 — 단일바이트 charset 은 `Iso8859_1 | UsAscii => 0` 으로 명시.
7+
- **반증 ⒝**: `new_stream_decoder` 는 exhaustive 지만 UTF-8·EUC-KR 이 같은 `CharsetStreamDecoder::EncodingRs` 로 접혀
8+
디코더만으로는 보류 규칙을 가를 수 없다 ⇒ 합칠 곳이 없어 `Charset` 메서드로 뒀다.
9+
- **동작 불변**: 두 술어 본문은 글자 그대로 옮겼다(UTF-8 의 `decode_length > 0` 가드는 `is_empty()` 조기 반환으로).
10+
- **양방향 변이**: 임시 변종 → `bytes_to_hold_back` 에서 non-exhaustive 컴파일 오류 · EUC-KR 보류 0 → #90 경계 테스트 red · 원상 green.
11+
- **한계(제안 자신의 tradeoff 그대로)**: match 가 막는 것은 «잊음»이지 «틀린 술어»가 아니다.
12+
- 후속 추천 0건.

‎rustjava-runtime/src/charset.rs‎

Lines changed: 43 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -59,6 +59,49 @@ impl Charset {
5959
}
6060
}
6161

62+
// How many trailing bytes of `bytes` to withhold from a non-final decode because they open a character the buffer
63+
// does not yet complete. No wildcard arm on purpose: a new charset must decide this, or its characters split at
64+
// read boundaries vanish silently.
65+
pub fn bytes_to_hold_back(&self, bytes: &[u8]) -> usize {
66+
match self {
67+
Self::Utf8 => {
68+
if bytes.is_empty() {
69+
return 0;
70+
}
71+
let mut lead_index = bytes.len() - 1;
72+
while lead_index > 0 && bytes[lead_index] & 0xc0 == 0x80 {
73+
lead_index -= 1;
74+
}
75+
let expected_length = match bytes[lead_index] {
76+
0xc0..=0xdf => 2,
77+
0xe0..=0xef => 3,
78+
0xf0..=0xf7 => 4,
79+
0x00..=0xbf | 0xf8..=0xff => 1,
80+
};
81+
if bytes.len() - lead_index < expected_length {
82+
bytes.len() - lead_index
83+
} else {
84+
0
85+
}
86+
}
87+
Self::EucKr => {
88+
// Unlike UTF-8, EUC-KR trail bytes overlap the lead range (0x81..=0xfe), so the last
89+
// byte alone cannot say whether the final pair is complete — a whole pair ends in a
90+
// byte that looks exactly like a lead. Withholding it there strands the pair's own
91+
// lead byte, which the decoder then swallows into state this read is about to drop.
92+
// The buffer always begins on a character boundary, so walk it forward instead: a byte
93+
// >= 0x81 opens a pair, anything else stands alone. Only a lead byte that overruns
94+
// the buffer is held back.
95+
let mut index = 0;
96+
while index < bytes.len() {
97+
index += if bytes[index] >= 0x81 { 2 } else { 1 };
98+
}
99+
if index > bytes.len() { 1 } else { 0 }
100+
}
101+
Self::Iso8859_1 | Self::UsAscii => 0,
102+
}
103+
}
104+
62105
pub fn new_stream_decoder(&self) -> CharsetStreamDecoder {
63106
match self {
64107
Self::Utf8 => CharsetStreamDecoder::EncodingRs(encoding_rs::UTF_8.new_decoder_without_bom_handling()),

‎rustjava-runtime/src/classes/java/io/input_stream_reader.rs‎

Lines changed: 5 additions & 31 deletions
Original file line numberDiff line numberDiff line change
@@ -56,7 +56,7 @@ impl InputStreamReader {
5656
// Unlike the (InputStream, String) ctor, JDK's single-argument ctor is not declared to throw
5757
// UnsupportedEncodingException, so an unusable default encoding must surface at read() instead.
5858
let charset = System::get_charset(jvm).await?;
59-
// Canonicalize when we recognize it so read()'s multibyte-boundary checks see "UTF-8"/"EUC-KR";
59+
// Canonicalize when we recognize it so the stored charset field holds the canonical name;
6060
// pass an unknown name through unchanged so read() reports it verbatim.
6161
let charset_name = Charset::from_name(&charset).map_or(charset.as_str(), |x| x.canonical_name());
6262
Self::init_fields(jvm, this, r#in, charset_name).await
@@ -161,40 +161,14 @@ impl InputStreamReader {
161161

162162
let charset_ref = jvm.get_field(&this, "charset", "Ljava/lang/String;").await?;
163163
let charset = JavaLangString::to_rust_string(jvm, &charset_ref).await?;
164-
let mut decoder = Charset::resolve(jvm, &charset).await?.new_stream_decoder();
164+
let charset = Charset::resolve(jvm, &charset).await?;
165+
let mut decoder = charset.new_stream_decoder();
165166

166167
let read_buf_data: Vec<u8> = cast_vec(read_buf_data);
167168
let end_of_input: bool = jvm.get_field(&this, "endOfInput", "Z").await?;
168169
let mut decode_length = read_buf_data.len();
169-
if !end_of_input && charset == "UTF-8" && decode_length > 0 {
170-
let mut lead_index = decode_length - 1;
171-
while lead_index > 0 && read_buf_data[lead_index] & 0xc0 == 0x80 {
172-
lead_index -= 1;
173-
}
174-
let expected_length = match read_buf_data[lead_index] {
175-
0xc0..=0xdf => 2,
176-
0xe0..=0xef => 3,
177-
0xf0..=0xf7 => 4,
178-
_ => 1,
179-
};
180-
if decode_length - lead_index < expected_length {
181-
decode_length = lead_index;
182-
}
183-
} else if !end_of_input && charset == "EUC-KR" {
184-
// Unlike UTF-8, EUC-KR trail bytes overlap the lead range (0x81..=0xfe), so the last
185-
// byte alone cannot say whether the final pair is complete — a whole pair ends in a
186-
// byte that looks exactly like a lead. Withholding it there strands the pair's own
187-
// lead byte, which the decoder then swallows into state this read is about to drop.
188-
// readBuf always begins on a character boundary, so walk it forward instead: a byte
189-
// >= 0x81 opens a pair, anything else stands alone. Only a lead byte that overruns
190-
// the buffer is held back.
191-
let mut index = 0;
192-
while index < decode_length {
193-
index += if read_buf_data[index] >= 0x81 { 2 } else { 1 };
194-
}
195-
if index > decode_length {
196-
decode_length -= 1;
197-
}
170+
if !end_of_input {
171+
decode_length -= charset.bytes_to_hold_back(&read_buf_data);
198172
}
199173

200174
let mut decoded = vec![0; BUF_SIZE * 3];

0 commit comments

Comments
 (0)