BUG: Type the xref stream as StreamObject rather than ContentStream - #3972
Conversation
`_read_pdf15_xref_stream` casts the result of `read_object` to ContentStream, but a compressed cross-reference stream is read as an EncodedStreamObject, which is a sibling of ContentStream rather than a subclass. The cast hid that, and `_sanitize_pdf15_xref_stream_index_pairs` declared the same wrong type for its parameter, so a runtime protocol check rejects the real object. Both now use StreamObject, the common base that actually covers the three stream types the function can return, and which still provides the only method used here, `get_data`. The return annotation listed the three siblings but not their base, so it is collapsed to StreamObject as well; the single caller passes the result to `_process_xref_stream`, which takes a DictionaryObject, and StreamObject extends that. Under `make testtype` this is 138 failures in test_writer alone; they go to 0. mypy reports the same 11 pre-existing errors before and after.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #3972 +/- ##
==========================================
+ Coverage 97.94% 97.98% +0.03%
==========================================
Files 57 57
Lines 11031 11098 +67
Branches 2065 2078 +13
==========================================
+ Hits 10804 10874 +70
+ Misses 126 125 -1
+ Partials 101 99 -2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This is because the corresponding method stems from a complexity refactoring with the original types.
I am not sure whether we need to build the PDF file explicitly and do not have a basic file in the |
The change is annotations only, and the existing cross-reference stream tests already cover this path with real files.
|
Both fair. I've dropped the test entirely — the change is annotations only, so an isinstance assertion tests the type checker rather than the library, and the existing cross-reference stream tests already cover this path with real files. tests/test_reader.py is now untouched and the PR is just the three annotations in _reader.py. Thanks for the context on _sanitize_pdf15_xref_stream_index_pairs — that explains why it carried the same type. |
_read_pdf15_xref_streamcasts the result ofread_objecttoContentStream:
A compressed cross-reference stream is read as an EncodedStreamObject,
which is a sibling of ContentStream rather than a subclass - they share
StreamObject as their base but neither is the other. The cast asserted
something untrue, and
_sanitize_pdf15_xref_stream_index_pairsdeclaredthe same wrong type for its
xref_streamparameter, so a runtime protocolcheck rejects the object that is actually passed.
Both now use StreamObject. That is the common base of the three stream
types this function can return, and it still provides
get_data, which isthe only method either place calls on the value.
The return annotation listed the three siblings but not their base, so it
is collapsed to StreamObject as well. The single caller passes the result
to
_process_xref_stream, which takes a DictionaryObject, and StreamObjectextends DictionaryObject. Removing the union left the ContentStream and
DecodedStreamObject imports unused, so those are dropped too.
Under
make testtypethis is 138 failures in test_writer alone; they goto 0. mypy reports the same 11 pre-existing errors before and after.
Added an offline test that builds a minimal PDF with a FlateDecode xref
stream and asserts the result is a StreamObject but not a ContentStream.
It fails on main.