Skip to content

BUG: Type the xref stream as StreamObject rather than ContentStream - #3972

Merged
stefan6419846 merged 2 commits into
py-pdf:mainfrom
RavSinghChandan:fix-xref-stream-type
Aug 24, 2026
Merged

stefan6419846 merged 2 commits into
py-pdf:mainfrom
RavSinghChandan:fix-xref-stream-type

Conversation

@RavSinghChandan

@RavSinghChandan RavSinghChandan commented Aug 16, 2026 •

Copy link
Copy Markdown
Contributor

_read_pdf15_xref_stream casts the result of read_object to
ContentStream:

xref_stream = cast(ContentStream, read_object(stream, self))

A compressed cross-reference stream is read as an EncodedStreamObject,
which is a sibling of ContentStream rather than a subclass - they share
StreamObject as their base but neither is the other. The cast asserted
something untrue, and _sanitize_pdf15_xref_stream_index_pairs declared
the same wrong type for its xref_stream parameter, so a runtime protocol
check rejects the object that is actually passed.

Both now use StreamObject. That is the common base of the three stream
types this function can return, and it still provides get_data, which is
the only method either place calls on the value.

The return annotation listed the three siblings but not their base, so it
is collapsed to StreamObject as well. The single caller passes the result
to _process_xref_stream, which takes a DictionaryObject, and StreamObject
extends DictionaryObject. Removing the union left the ContentStream and
DecodedStreamObject imports unused, so those are dropped too.

Under make testtype this is 138 failures in test_writer alone; they go
to 0. mypy reports the same 11 pre-existing errors before and after.

Added an offline test that builds a minimal PDF with a FlateDecode xref
stream and asserts the result is a StreamObject but not a ContentStream.
It fails on main.

`_read_pdf15_xref_stream` casts the result of `read_object` to
ContentStream, but a compressed cross-reference stream is read as an
EncodedStreamObject, which is a sibling of ContentStream rather than a
subclass. The cast hid that, and
`_sanitize_pdf15_xref_stream_index_pairs` declared the same wrong type
for its parameter, so a runtime protocol check rejects the real object.

Both now use StreamObject, the common base that actually covers the
three stream types the function can return, and which still provides the
only method used here, `get_data`. The return annotation listed the three
siblings but not their base, so it is collapsed to StreamObject as well;
the single caller passes the result to `_process_xref_stream`, which
takes a DictionaryObject, and StreamObject extends that.

Under `make testtype` this is 138 failures in test_writer alone; they go
to 0. mypy reports the same 11 pre-existing errors before and after.
@codecov

codecov Bot commented Aug 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.98%. Comparing base (1bce7a7) to head (fe35c42).
⚠️ Report is 18 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3972      +/-   ##
==========================================
+ Coverage   97.94%   97.98%   +0.03%     
==========================================
  Files          57       57              
  Lines       11031    11098      +67     
  Branches     2065     2078      +13     
==========================================
+ Hits        10804    10874      +70     
+ Misses        126      125       -1     
+ Partials      101       99       -2     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@stefan6419846

Copy link
Copy Markdown
Collaborator

_sanitize_pdf15_xref_stream_index_pairs declared the same wrong type

This is because the corresponding method stems from a complexity refactoring with the original types.

Added an offline test that builds a minimal PDF with a FlateDecode xref stream and asserts the result is a StreamObject but not a ContentStream.

I am not sure whether we need to build the PDF file explicitly and do not have a basic file in the resources or sample-files directory which we could use directly. Additionally, having typing tested in this manner might not be required at all.

The change is annotations only, and the existing cross-reference stream tests
already cover this path with real files.
@RavSinghChandan

Copy link
Copy Markdown
Contributor Author

Both fair. I've dropped the test entirely — the change is annotations only, so an isinstance assertion tests the type checker rather than the library, and the existing cross-reference stream tests already cover this path with real files. tests/test_reader.py is now untouched and the PR is just the three annotations in _reader.py.

Thanks for the context on _sanitize_pdf15_xref_stream_index_pairs — that explains why it carried the same type.

@stefan6419846
stefan6419846 merged commit 1bd549d into py-pdf:main Aug 24, 2026
60 of 76 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants