Skip to content

fix: decode only percent-encoded escapes in decodeURIComponent (#25672) (CP: 24.10) - #26102

Merged
mcollovati merged 1 commit into
24.10from
cherry-pick-25672-to-24.10-1790838818186
Oct 1, 2026
Merged

mcollovati merged 1 commit into
24.10from
cherry-pick-25672-to-24.10-1790838818186

Conversation

@totally-not-ai

@totally-not-ai totally-not-ai Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

This PR cherry-picks changes from the original PR #25672 to branch 24.10.

Original PR description

Summary

UrlUtil.decodeURIComponent treated every non-ASCII character as a raw UTF-8 byte, even when it was never percent-encoded. Because of that, paths that already contain literal characters like ü or 日 were corrupted into �, so routes with non-ASCII segments did not match and wildcard parameters lost their text. Now only real %XX escapes are decoded and all other characters are left untouched.

What changed

Behavior change: UrlUtil.decodeURIComponent no longer rewrites non-ASCII characters that are not percent-encoded. This affects anyone who passes an already decoded (or partly decoded) string: before it came back mangled, now it comes back unchanged. Strings that only contain %XX escapes decode exactly as before, so normal encoded input is unaffected.

  • decodeURIComponent now collects consecutive %XX escapes into a byte sequence and decodes that sequence as UTF-8. Text between escapes is copied as-is, so a multi-byte character split over several escapes still decodes to one character.
  • Input without any escape is returned directly.
  • Javadoc now states that unescaped characters are kept as they are.

Why this matters in practice: a servlet container decodes the path info, so the first server-side navigation sees literal characters. Static route segments with a non-ASCII character never matched, and @WildcardParameter values were corrupted. Jar URLs are also not required to be percent-encoded, so ResourceFolderUtil silently found no resources in a folder whose entry name contains a non-ASCII character.

No public or protected API was added, removed, or changed.

Fixes #25671

Test summary

# Status What the test verifies Why it matters
1 ✅ A string with literal non-ASCII characters (grüße, 日本, an emoji) is returned unchanged This is the bug: such input used to become �
2 ✅ A string mixing %XX escapes and literal characters decodes to grüße-ü-äxö Both forms must work in the same string; also pins that multi-byte escapes still decode
3 ✅ A static route grüße matches both the literal and the percent-encoded location Server-side navigation sees literal text, client-side sees encoded text
4 ✅ A @WildcardParameter value keeps grüße for literal input and decodes it for encoded input Corrupted parameter values were the visible symptom for apps
5 ✅ PathUtil.getSegmentsListWithDecoding keeps literal UTF-8 segments and splits them correctly Route resolution is built on this splitting step
6 ✅ ResourceFolderUtil.visitFiles finds files in a jar folder named thèmes/ Resources in such folders were silently skipped
7 ❗ gap Behaviour for a malformed or truncated escape (for example a lone %C3) Invalid input should degrade predictably, not throw
  • UrlUtilTest.decodeURIComponent_literalNonAsciiCharacters_returnedUnchanged → 1
  • UrlUtilTest.decodeURIComponent_literalAndEncodedNonAsciiCharacters_bothDecoded → 2
  • RouterTest.static_route_with_non_ascii_character → 3
  • RouterTest.wildcard_parameter_with_non_ascii_characters → 4
  • PathUtilTest.getSegmentsListWithDecoding_handlesUtf8Characters (extended) → 5
  • ResourceFolderUtilTest.folderPathContainsLiteralNonAsciiCharacter_filesInTheJarAreVisited → 6

Deliberately not tested: the new private appendDecoded helper, which is covered through the public method, and plain ASCII or %2F decoding, which existing tests in UrlUtilTest and PathUtilTest already pin.

## Summary
`UrlUtil.decodeURIComponent` treated every non-ASCII character as a raw
UTF-8 byte, even when it was never percent-encoded. Because of that,
paths that already contain literal characters like `ü` or `日` were
corrupted into `�`, so routes with non-ASCII segments did not match and
wildcard parameters lost their text. Now only real `%XX` escapes are
decoded and all other characters are left untouched.

## What changed
**Behavior change:** `UrlUtil.decodeURIComponent` no longer rewrites
non-ASCII characters that are not percent-encoded. This affects anyone
who passes an already decoded (or partly decoded) string: before it came
back mangled, now it comes back unchanged. Strings that only contain
`%XX` escapes decode exactly as before, so normal encoded input is
unaffected.

- `decodeURIComponent` now collects consecutive `%XX` escapes into a
byte sequence and decodes that sequence as UTF-8. Text between escapes
is copied as-is, so a multi-byte character split over several escapes
still decodes to one character.
- Input without any escape is returned directly.
- Javadoc now states that unescaped characters are kept as they are.

Why this matters in practice: a servlet container decodes the path info,
so the first server-side navigation sees literal characters. Static
route segments with a non-ASCII character never matched, and
`@WildcardParameter` values were corrupted. Jar URLs are also not
required to be percent-encoded, so `ResourceFolderUtil` silently found
no resources in a folder whose entry name contains a non-ASCII
character.

No public or protected API was added, removed, or changed.

Fixes #25671

## Test summary

| # | Status | What the test verifies | Why it matters |
|---|--------|------------------------|----------------|
| 1 | ✅ | A string with literal non-ASCII characters (`grüße`, `日本`, an
emoji) is returned unchanged | This is the bug: such input used to
become `�` |
| 2 | ✅ | A string mixing `%XX` escapes and literal characters decodes
to `grüße-ü-äxö` | Both forms must work in the same string; also pins
that multi-byte escapes still decode |
| 3 | ✅ | A static route `grüße` matches both the literal and the
percent-encoded location | Server-side navigation sees literal text,
client-side sees encoded text |
| 4 | ✅ | A `@WildcardParameter` value keeps `grüße` for literal input
and decodes it for encoded input | Corrupted parameter values were the
visible symptom for apps |
| 5 | ✅ | `PathUtil.getSegmentsListWithDecoding` keeps literal UTF-8
segments and splits them correctly | Route resolution is built on this
splitting step |
| 6 | ✅ | `ResourceFolderUtil.visitFiles` finds files in a jar folder
named `thèmes/` | Resources in such folders were silently skipped |
| 7 | ❗ **gap** | Behaviour for a malformed or truncated escape (for
example a lone `%C3`) | Invalid input should degrade predictably, not
throw |

-
`UrlUtilTest.decodeURIComponent_literalNonAsciiCharacters_returnedUnchanged`
→ 1
-
`UrlUtilTest.decodeURIComponent_literalAndEncodedNonAsciiCharacters_bothDecoded`
→ 2
- `RouterTest.static_route_with_non_ascii_character` → 3
- `RouterTest.wildcard_parameter_with_non_ascii_characters` → 4
- `PathUtilTest.getSegmentsListWithDecoding_handlesUtf8Characters`
(extended) → 5
-
`ResourceFolderUtilTest.folderPathContainsLiteralNonAsciiCharacter_filesInTheJarAreVisited`
→ 6

Deliberately not tested: the new private `appendDecoded` helper, which
is covered through the public method, and plain ASCII or `%2F` decoding,
which existing tests in `UrlUtilTest` and `PathUtilTest` already pin.

---------

Co-authored-by: totally-not-ai[bot] <290682512+totally-not-ai[bot]@users.noreply.github.com>
Co-authored-by: Artur Signell <artur@vaadin.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Test Results

1 195 files  ± 0  1 195 suites  ±0   1h 4m 18s ⏱️ + 5m 19s
8 968 tests + 4  8 906 ✅ + 4  62 💤 ±0  0 ❌ ±0 
9 146 runs   - 19  9 083 ✅  - 18  63 💤  - 1  0 ❌ ±0 

Results for commit c22972a. ± Comparison against base commit 404da9b.

@mcollovati
mcollovati merged commit 59f9cb2 into 24.10 Oct 1, 2026
32 checks passed
@mcollovati
mcollovati deleted the cherry-pick-25672-to-24.10-1790838818186 branch October 1, 2026 09:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant