Skip to content

email.parser.BytesParser.parse() cannot handle binary data that include \x0d \x0a correctly. #128949

Description

@tnakamot

Bug report

Bug description:

I would like to extract a binary file in a multipart MIME file using email.parser.BytesParser, but a byte sequence "0x0d 0x0a" (CR + LF) in the binary file is replaced by "0x0a" (LF). Below is a minimal reproducible example.

from email.parser import BytesParser
from email.policy import default
from io import BytesIO

mime_file_byte_array = b'MIME-Version: 1.0\r\nContent-Type: multipart/mixed; boundary="MIME\
_boundary-1";\r\n\r\n--MIME_boundary-1\r\nContent-Type: application/octet-stream\r\nContent\
-Location: test.bin\r\n\r\na\r\nb\r\n--MIME_boundary-1--\r\n\r\n'
fp = BytesIO(mime_file_byte_array)
parser = BytesParser(policy=default)
msg = parser.parse(fp)

parts = [part for part in msg.walk()]
binary_data = parts[1].get_payload(decode=True)

print('===== Beginning of Original MIME File =====')
print(mime_file_byte_array.decode())
print('===== End of Original MIME File =====')
print('')
print('===== test.bin after parse =====')
print(binary_data)
print('===== test.bin after parse =====')

As can be seen in the fifth line, the multipart MIME file includes a binary file "test.bin". The contents of the binary file is b"a\r\nb".
Therefore, the variable binary_data is supposed to contain b"a\r\nb", but it was actually b"a\nb".

It is probably because TextIOWrapper in BytesParser.parse() translates CR+LF to LF on Linux.

fp = TextIOWrapper(fp, encoding='ascii', errors='surrogateescape')

When I replaced the above line with the line below, this problem was fixed. However, this fix may have a side effect which I cannot foresee.

fp = TextIOWrapper(fp, encoding='ascii', errors='surrogateescape', newline='')

CPython versions tested on:

3.10

Operating systems tested on:

Linux

Linked PRs

Activity

  1. RanKKI commented on Feb 16, 2025

    @RanKKI
    Contributor

    Using the parameter newline='' resolves this issue. By default (when newline is None), the wrapper translates CRLF/LF to os.linesep. When newline is set to an empty string, the wrapper does not modify CRLF/LF.

    cpython/Modules/_io/textio.c

    Lines 1085 to 1089 in a7d41a8

    * On output, if newline is None, any '\n' characters written are
    translated to the system default line separator, os.linesep. If
    newline is '' or '\n', no translation takes place. If newline is any
    of the other legal values, any '\n' characters written are translated
    to the given string.

    But,

    according to RFC 2046#4.1.1, text/* types must always use CRLF line endings

    The canonical form of any MIME "text" subtype MUST always represent a
    line break as a CRLF sequence.

    The current implementation has a bug (I guess), and I don't see an easy solution to fix it.

  2. tnakamot commented on Feb 16, 2025

    @tnakamot
    Author

    Thank you for your comment, @RanKKI . Yes, it is understandable if this translation from CRLF to os.linesep is activated only for the text/* types. I guess that some python applications that use the email.parser module may be relying on this feature of the line break translation. So, in order to maintain the backward compatibility, a possible solution is to deactivate this translation if the content type is not text/*.

  3. aeurielesn commented on Jul 20, 2025

    @aeurielesn
    Contributor

    It does not seem inconceivable to add the newline parameter to BytesParser.parse as it seems to have been done in other cases where disabling universal newline conversion is desirable or, alternatively, using BytesParser.parsebytes in some cases.

  4. medmunds commented on Sep 5, 2026

    @medmunds
    Contributor

    This issue results in corrupt images and attachments when parsing messages that use Content-Transfer-Encoding: 8bit. In case that's not obvious from the example in the original report, here's another:

    from email import policy
    from email.parser import BytesParser
    from io import BytesIO
    
    # A 10x2573 GIF image (some data omitted)
    gif = b"GIF89a\x0a\x00\x0d\x0a...\x3b"
    
    raw_message = b"""\
    Subject: Here's a GIF\r
    Content-Type: image/gif\r
    Content-Transfer-Encoding: 8bit\r
    \r
    """ + gif
    
    msg1 = BytesParser(policy=policy.default).parsebytes(raw_message)
    print(repr(msg1.get_content()))  # b'GIF89a\n\x00\r\n...;'
    assert msg1.get_content() == gif  # Works correctly
    
    msg2 = BytesParser(policy=policy.default).parse(BytesIO(raw_message))
    print(repr(msg2.get_content()))  # b'GIF89a\n\x00\n...;
    assert msg2.get_content() == gif  # AssertionError (corrupted image)

    A related problem is that text parts will have inconsistent line endings, depending which BytesParser method was used to parse the message.

    Disabling newline conversion where BytesParser.parse() uses a TextIOWrapper seems workable, and a minimal fix.

    Another option would be switching to a BytesFeedParser within BytesParser, rather than trying to force bytes-as-text intoParser. (Eventually everything ends up in a FeedParser anyway; this would just change the intermediary and avoid the TextIOWrapper. The distinction between bytes and non-bytes parsers is kind of left over from Python 2.)

    I'd defer to @bitdancer on the better approach.

  5. added a commit that references this issue on Sep 18, 2026
  6. added a commit that references this issue on Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    stdlibStandard Library Python modules in the Lib/ directorytopic-emailtype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions