Skip to content

Commit f7586aa

Browse files
gh-157675: Add *limit* argument to encodings.idna.nameprep (GH-157682)
Add a *limit* argument to nameprep to allow ToASCII and ToUnicode to reject extremely large input early. This replaces the check in #99092, while allowing any number of harmless "characters mapped to nothing" (RFC 3454 §3.1). The default stays unlimited, to not change behaviour for users that call nameprep manually (and don't necessarily follow up with punycode). Co-authored-by: Stan Ulbrych <stan@python.org>
1 parent 2095eab commit f7586aa

5 files changed

Lines changed: 92 additions & 31 deletions

File tree

‎Doc/library/codecs.rst‎

Lines changed: 21 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -1395,15 +1395,6 @@ encodings.
13951395
| | | :mod:`encodings.idna`. |
13961396
| | | Only ``errors='strict'`` |
13971397
| | | is supported. |
1398-
| | | |
1399-
| | | .. warning:: |
1400-
| | | |
1401-
| | | This codec builds on |
1402-
| | | ``punycode``, whose |
1403-
| | | algorithms scale |
1404-
| | | poorly, so limit the |
1405-
| | | length of untrusted |
1406-
| | | input. |
14071398
+--------------------+---------+---------------------------+
14081399
| mbcs | ansi, | Windows only: Encode the |
14091400
| | dbcs | operand according to the |
@@ -1655,11 +1646,6 @@ Applications) and :rfc:`3492` (Nameprep: A Stringprep Profile for
16551646
Internationalized Domain Names (IDN)). It builds upon the ``punycode`` encoding
16561647
and :mod:`stringprep`.
16571648

1658-
.. warning::
1659-
1660-
This module builds on ``punycode``, whose algorithms scale poorly, so limit
1661-
the length of untrusted input.
1662-
16631649
If you need the IDNA 2008 standard from :rfc:`5891` and :rfc:`5895`, use the
16641650
third-party :pypi:`idna` module.
16651651

@@ -1697,11 +1683,31 @@ international domain names, and to unify similar characters. The nameprep
16971683
functions can be used directly if desired.
16981684

16991685

1700-
.. function:: nameprep(label)
1686+
.. function:: nameprep(label, *, limit=None)
17011687

17021688
Return the nameprepped version of *label*. The implementation currently assumes
17031689
query strings, so ``AllowUnassigned`` is true.
17041690

1691+
Raise :exc:`UnicodeEncodeError` if the nameprep algorithm emits an error.
1692+
1693+
If the *limit* argument is given, it should be set to the maximum size
1694+
of an encoded A-label (that is, 63 for IDNA).
1695+
:func:`!nameprep` will raise :exc:`UnicodeEncodeError` if the label is
1696+
**much** larger than *limit*.
1697+
Note that this is only a rough check meant to skip expensive processing
1698+
of extremely large input; the caller should check any exact
1699+
limits separately.
1700+
1701+
.. warning::
1702+
1703+
For backwards compatibility, label size is unlimited by default.
1704+
This may cause issues when processing the result with the
1705+
``punycode`` encoding, whose algorithms scale poorly.
1706+
1707+
.. versionchanged:: next
1708+
1709+
Added the *limit* parameter.
1710+
17051711

17061712
.. function:: ToASCII(label)
17071713

‎Doc/whatsnew/3.16.rst‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -397,6 +397,10 @@ encodings
397397
used for international IMAP4 mailbox names (:rfc:`3501`).
398398
(Contributed by Serhiy Storchaka in :gh:`66788`.)
399399

400+
* :func:`encodings.idna.nameprep` now takes a *limit* argument that allows
401+
rejecting extremely large input early.
402+
(Contributed by Petr Viktorin in :gh:`157675`.)
403+
400404

401405
gzip
402406
----

‎Lib/encodings/idna.py‎

Lines changed: 31 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
# This module implements the RFCs 3490 (IDNA) and 3491 (Nameprep)
22

3+
import sys
34
import stringprep, re, codecs
45
from unicodedata import ucd_3_2_0 as unicodedata
56

@@ -11,14 +12,36 @@
1112
sace_prefix = "xn--"
1213

1314
# This assumes query strings, so AllowUnassigned is true
14-
def nameprep(label): # type: (str) -> str
15+
def nameprep(label, *, limit=None): # type: (str) -> str
16+
if limit is None:
17+
limit = sys.maxsize
18+
else:
19+
# Protection from gh-98433 and gh-157675 (passing unbounded input to
20+
# the quadratic-complexity punycode algorithm).
21+
# While the "map" step can remove characters, later steps (in ToASCII
22+
# and FromASCII) will not shorten the result *drastically*.
23+
# (NFKC normalization can compress e.g. '\u03c9\u0314\u0300\u0345'
24+
# to '\u1fa3' -- a 4-fold reduction. Non-ASCII labels then get
25+
# longer via prefixing & punycode).
26+
# We bail if the number of non-ignored input characters exceeds 8 times
27+
# the limit, which gives ample room for future Unicode versions to
28+
# include long normalizations, while still preventing us from wasting
29+
# time decoding a big thing that'll just hit the actual <= 63 limit in
30+
# ToASCII.
31+
limit *= 8
32+
1533
# Map
1634
newlabel = []
1735
for c in label:
1836
if stringprep.in_table_b1(c):
1937
# Map to nothing
2038
continue
2139
newlabel.append(stringprep.map_table_b2(c))
40+
41+
if len(newlabel) > limit:
42+
raise UnicodeEncodeError("idna", label, 0, len(label),
43+
"label way too long")
44+
2245
label = "".join(newlabel)
2346

2447
# Normalize
@@ -80,7 +103,7 @@ def ToASCII(label): # type: (str) -> bytes
80103
raise UnicodeEncodeError("idna", label, 0, len(label), "label too long")
81104

82105
# Step 2: nameprep
83-
label = nameprep(label)
106+
label = nameprep(label, limit=63)
84107

85108
# Step 3: UseSTD3ASCIIRules is false
86109
# Step 4: try ASCII
@@ -115,18 +138,6 @@ def ToASCII(label): # type: (str) -> bytes
115138
raise UnicodeEncodeError("idna", label, 0, len(label), "label too long")
116139

117140
def ToUnicode(label):
118-
if len(label) > 1024:
119-
# Protection from https://github.com/python/cpython/issues/98433.
120-
# https://datatracker.ietf.org/doc/html/rfc5894#section-6
121-
# doesn't specify a label size limit prior to NAMEPREP. But having
122-
# one makes practical sense.
123-
# This leaves ample room for nameprep() to remove Nothing characters
124-
# per https://www.rfc-editor.org/rfc/rfc3454#section-3.1 while still
125-
# preventing us from wasting time decoding a big thing that'll just
126-
# hit the actual <= 63 length limit in Step 6.
127-
if isinstance(label, str):
128-
label = label.encode("utf-8", errors="backslashreplace")
129-
raise UnicodeDecodeError("idna", label, 0, len(label), "label way too long")
130141
# Step 1: Check for ASCII
131142
if isinstance(label, bytes):
132143
pure_ascii = True
@@ -139,7 +150,7 @@ def ToUnicode(label):
139150
if not pure_ascii:
140151
assert isinstance(label, str)
141152
# Step 2: Perform nameprep
142-
label = nameprep(label)
153+
label = nameprep(label, limit=63)
143154
# It doesn't say this, but apparently, it should be ASCII now
144155
try:
145156
label = label.encode("ascii")
@@ -151,6 +162,11 @@ def ToUnicode(label):
151162
if not label.lower().startswith(ace_prefix):
152163
return str(label, "ascii")
153164

165+
# Below in steps 6-7, `label` must match the result of `ToASCII`, so it's
166+
# limited to 63 chars. Check before the expensive punycode decode.
167+
if len(label) >= 64:
168+
raise UnicodeDecodeError("idna", label, 0, len(label), "label too long")
169+
154170
# Step 4: Remove ACE prefix
155171
label1 = label[len(ace_prefix):]
156172

‎Lib/test/test_codecs.py‎

Lines changed: 31 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1646,6 +1646,14 @@ def test_nameprep(self):
16461646
except Exception as e:
16471647
raise support.TestFailed("Test 3.%d: %s" % (pos+1, str(e)))
16481648

1649+
def test_long_input(self):
1650+
from encodings.idna import nameprep
1651+
self.assertEqual(nameprep("x" + "\N{ZWJ}" * 10_000 + "y"), 'xy')
1652+
self.assertEqual(nameprep("x" + "\N{ZWJ}" * 10_000 + "y", limit=2),
1653+
'xy')
1654+
with self.assertRaises(UnicodeEncodeError):
1655+
nameprep("x" + "\N{SNAKE}" * 10_000 + "y", limit=10)
1656+
16491657

16501658
class IDNACodecTest(unittest.TestCase):
16511659

@@ -1716,11 +1724,33 @@ def test_builtin_encode_invalid(self):
17161724
self.assertEqual(exc.end, expected.end)
17171725

17181726
def test_builtin_decode_length_limit(self):
1719-
with self.assertRaisesRegex(UnicodeDecodeError, "way too long"):
1727+
with self.assertRaisesRegex(UnicodeDecodeError, "too long"):
17201728
(b"xn--016c"+b"a"*1100).decode("idna")
17211729
with self.assertRaisesRegex(UnicodeDecodeError, "too long"):
17221730
(b"xn--016c"+b"a"*70).decode("idna")
17231731

1732+
def test_builtin_encode_length_limit(self):
1733+
with self.assertRaisesRegex(UnicodeEncodeError, "too long"):
1734+
("x" * 64).encode("idna")
1735+
with self.assertRaisesRegex(UnicodeEncodeError, "too long"):
1736+
("short." + "x" * 64).encode("idna")
1737+
1738+
# Test at both sides of the limit (<64 bytes)
1739+
self.assertEqual(len(("\N{SNAKE}" * 56).encode("idna")), 63)
1740+
with self.assertRaisesRegex(UnicodeEncodeError, "too long"):
1741+
("\N{SNAKE}" * 57).encode("idna")
1742+
with self.assertRaisesRegex(UnicodeEncodeError, "too long"):
1743+
("short." + "\N{SNAKE}" * 57).encode("idna")
1744+
1745+
# Very long names are handled
1746+
with self.assertRaisesRegex(UnicodeEncodeError, "way too long"):
1747+
("\N{SNAKE}"*50_000).encode("idna")
1748+
1749+
# The limit doesn't apply to ignored characters
1750+
self.assertEqual(('a' + "\N{ZWSP}"*50_000 + 'b').encode('idna'), b'ab')
1751+
self.assertEqual(('a' + "\N{ZWSP}"*50_000 + '\N{SNAKE}').encode('idna'),
1752+
b'xn--a-012s')
1753+
17241754
def test_stream(self):
17251755
r = codecs.getreader("idna")(io.BytesIO(b"abc"))
17261756
r.read(3)
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
:func:`encodings.idna.nameprep` now takes a *limit* argument that allows
2+
rejecting extremely large input early. The ``idna`` encoding and the
3+
:func:`!encodings.idna.ToASCII` and :func:`!encodings.idna.ToUnicode`
4+
functions use this to avoid passing unbounded input to quadratic-time
5+
algorithms.

0 commit comments

Comments
 (0)