Repository navigation
Base the tokenizer API on source offsets #153569
Copy link
Copy link
Open
Labels
interpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
Description
Activity
Hi @pablogsal, I'd like to work on this issue. Can you please assign it to me?
Hi @pablogsal, I'd like to work on this issue. Can you please assign it to me?
Thanks for the help @Arbaaz123676! Unfortunately this is something I am working on right now as a series of commits and is very delicate work. If you want to help reviewing the PRs that would be good!
Reacted by Arbaaz Ahmed- addedtype-featureA feature request or enhancementA feature request or enhancementinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jul 12, 2026 - added 5 commits that reference this issue
on Aug 25, 2026 220 remaining items
Load more actions- removedtype-featureA feature request or enhancementA feature request or enhancement
on Oct 5, 2026 - added 12 commits that reference this issue
on Oct 5, 2026
Metadata
Metadata
Assignees
Labels
interpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
The tokenizer currently stores many source positions as pointers into buffers that can move. Reallocating the input means rebasing all these pointers, and missing one is very easy, especially with incremental input and f-strings.
I think tokenizer positions should be offsets into the decoded source instead:
The main ideas would be:
_tokenizeask the source for a view or copy of a span.tokenize(readline)can eventually use reclaimable chunks.For incremental tokenization, new decoded input would be appended only when the cursor needs more data. Existing offsets would remain valid even if the underlying storage moves. Once
tokenizehas returned copied token and line strings, old input could be discarded when no active token, cursor, f-string frame or error still refers to it.This is very nice because it separates source storage from tokenizer state, removes pointer rebasing, and means consumers no longer need to access tokenizer internals.
Linked PRs