Skip to content

fix: decode only percent-encoded escapes in decodeURIComponent (#25672) (CP: 25.3) - #25703

Merged
vaadin-bot merged 1 commit into
25.3from
cherry-pick-25672-to-25.3-1789382990434
Sep 14, 2026
Merged

vaadin-bot merged 1 commit into
25.3from
cherry-pick-25672-to-25.3-1789382990434

Conversation

@vaadin-bot

Copy link
Copy Markdown
Collaborator

This PR cherry-picks changes from the original PR #25672 to branch 25.3.

Original PR description

Summary

UrlUtil.decodeURIComponent treated every non-ASCII character as a raw UTF-8 byte, even when it was never percent-encoded. Because of that, paths that already contain literal characters like ü or 日 were corrupted into �, so routes with non-ASCII segments did not match and wildcard parameters lost their text. Now only real %XX escapes are decoded and all other characters are left untouched.

What changed

Behavior change: UrlUtil.decodeURIComponent no longer rewrites non-ASCII characters that are not percent-encoded. This affects anyone who passes an already decoded (or partly decoded) string: before it came back mangled, now it comes back unchanged. Strings that only contain %XX escapes decode exactly as before, so normal encoded input is unaffected.

  • decodeURIComponent now collects consecutive %XX escapes into a byte sequence and decodes that sequence as UTF-8. Text between escapes is copied as-is, so a multi-byte character split over several escapes still decodes to one character.
  • Input without any escape is returned directly.
  • Javadoc now states that unescaped characters are kept as they are.

Why this matters in practice: a servlet container decodes the path info, so the first server-side navigation sees literal characters. Static route segments with a non-ASCII character never matched, and @WildcardParameter values were corrupted. Jar URLs are also not required to be percent-encoded, so ResourceFolderUtil silently found no resources in a folder whose entry name contains a non-ASCII character.

No public or protected API was added, removed, or changed.

Fixes #25671

Test summary

# Status What the test verifies Why it matters
1 ✅ A string with literal non-ASCII characters (grüße, 日本, an emoji) is returned unchanged This is the bug: such input used to become �
2 ✅ A string mixing %XX escapes and literal characters decodes to grüße-ü-äxö Both forms must work in the same string; also pins that multi-byte escapes still decode
3 ✅ A static route grüße matches both the literal and the percent-encoded location Server-side navigation sees literal text, client-side sees encoded text
4 ✅ A @WildcardParameter value keeps grüße for literal input and decodes it for encoded input Corrupted parameter values were the visible symptom for apps
5 ✅ PathUtil.getSegmentsListWithDecoding keeps literal UTF-8 segments and splits them correctly Route resolution is built on this splitting step
6 ✅ ResourceFolderUtil.visitFiles finds files in a jar folder named thèmes/ Resources in such folders were silently skipped
7 ❗ gap Behaviour for a malformed or truncated escape (for example a lone %C3) Invalid input should degrade predictably, not throw
  • UrlUtilTest.decodeURIComponent_literalNonAsciiCharacters_returnedUnchanged → 1
  • UrlUtilTest.decodeURIComponent_literalAndEncodedNonAsciiCharacters_bothDecoded → 2
  • RouterTest.static_route_with_non_ascii_character → 3
  • RouterTest.wildcard_parameter_with_non_ascii_characters → 4
  • PathUtilTest.getSegmentsListWithDecoding_handlesUtf8Characters (extended) → 5
  • ResourceFolderUtilTest.folderPathContainsLiteralNonAsciiCharacter_filesInTheJarAreVisited → 6

Deliberately not tested: the new private appendDecoded helper, which is covered through the public method, and plain ASCII or %2F decoding, which existing tests in UrlUtilTest and PathUtilTest already pin.

## Summary
`UrlUtil.decodeURIComponent` treated every non-ASCII character as a raw
UTF-8 byte, even when it was never percent-encoded. Because of that,
paths that already contain literal characters like `ü` or `日` were
corrupted into `�`, so routes with non-ASCII segments did not match and
wildcard parameters lost their text. Now only real `%XX` escapes are
decoded and all other characters are left untouched.

## What changed
**Behavior change:** `UrlUtil.decodeURIComponent` no longer rewrites
non-ASCII characters that are not percent-encoded. This affects anyone
who passes an already decoded (or partly decoded) string: before it came
back mangled, now it comes back unchanged. Strings that only contain
`%XX` escapes decode exactly as before, so normal encoded input is
unaffected.

- `decodeURIComponent` now collects consecutive `%XX` escapes into a
byte sequence and decodes that sequence as UTF-8. Text between escapes
is copied as-is, so a multi-byte character split over several escapes
still decodes to one character.
- Input without any escape is returned directly.
- Javadoc now states that unescaped characters are kept as they are.

Why this matters in practice: a servlet container decodes the path info,
so the first server-side navigation sees literal characters. Static
route segments with a non-ASCII character never matched, and
`@WildcardParameter` values were corrupted. Jar URLs are also not
required to be percent-encoded, so `ResourceFolderUtil` silently found
no resources in a folder whose entry name contains a non-ASCII
character.

No public or protected API was added, removed, or changed.

Fixes #25671

## Test summary

| # | Status | What the test verifies | Why it matters |
|---|--------|------------------------|----------------|
| 1 | ✅ | A string with literal non-ASCII characters (`grüße`, `日本`, an
emoji) is returned unchanged | This is the bug: such input used to
become `�` |
| 2 | ✅ | A string mixing `%XX` escapes and literal characters decodes
to `grüße-ü-äxö` | Both forms must work in the same string; also pins
that multi-byte escapes still decode |
| 3 | ✅ | A static route `grüße` matches both the literal and the
percent-encoded location | Server-side navigation sees literal text,
client-side sees encoded text |
| 4 | ✅ | A `@WildcardParameter` value keeps `grüße` for literal input
and decodes it for encoded input | Corrupted parameter values were the
visible symptom for apps |
| 5 | ✅ | `PathUtil.getSegmentsListWithDecoding` keeps literal UTF-8
segments and splits them correctly | Route resolution is built on this
splitting step |
| 6 | ✅ | `ResourceFolderUtil.visitFiles` finds files in a jar folder
named `thèmes/` | Resources in such folders were silently skipped |
| 7 | ❗ **gap** | Behaviour for a malformed or truncated escape (for
example a lone `%C3`) | Invalid input should degrade predictably, not
throw |

-
`UrlUtilTest.decodeURIComponent_literalNonAsciiCharacters_returnedUnchanged`
→ 1
-
`UrlUtilTest.decodeURIComponent_literalAndEncodedNonAsciiCharacters_bothDecoded`
→ 2
- `RouterTest.static_route_with_non_ascii_character` → 3
- `RouterTest.wildcard_parameter_with_non_ascii_characters` → 4
- `PathUtilTest.getSegmentsListWithDecoding_handlesUtf8Characters`
(extended) → 5
-
`ResourceFolderUtilTest.folderPathContainsLiteralNonAsciiCharacter_filesInTheJarAreVisited`
→ 6

Deliberately not tested: the new private `appendDecoded` helper, which
is covered through the public method, and plain ASCII or `%2F` decoding,
which existing tests in `UrlUtilTest` and `PathUtilTest` already pin.

---------

Co-authored-by: totally-not-ai[bot] <290682512+totally-not-ai[bot]@users.noreply.github.com>
Co-authored-by: Artur Signell <artur@vaadin.com>
@vaadin-bot

Copy link
Copy Markdown
Collaborator Author

This PR is eligible for auto-merging policy, so it has been approved automatically. If there are pending conditions, auto merge (with 'squash' method) has been enabled for this PR [Message is sent from bot]

@vaadin-bot
vaadin-bot enabled auto-merge (squash) September 14, 2026 11:02
@sonarqubecloud

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown
Contributor

Test Results

 1 439 files  ±0   1 523 suites  ±0   1h 36m 58s ⏱️ + 4m 58s
12 051 tests +5  11 983 ✅ +5  68 💤 ±0  0 ❌ ±0 
12 369 runs  +5  12 301 ✅ +5  68 💤 ±0  0 ❌ ±0 

Results for commit 0e16297. ± Comparison against base commit 361f4ae.

@vaadin-bot
vaadin-bot merged commit 357fffc into 25.3 Sep 14, 2026
42 checks passed
@vaadin-bot
vaadin-bot deleted the cherry-pick-25672-to-25.3-1789382990434 branch September 14, 2026 11:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants