Skip to content

gh-158880: Fix asyncio ps showing tasks from only one interpreter - #158900

Open
deadlovelll wants to merge 3 commits into
python:mainfrom
deadlovelll:gh-158880-interp
Open

deadlovelll wants to merge 3 commits into
python:mainfrom
deadlovelll:gh-158880-interp

Conversation

@deadlovelll

Copy link
Copy Markdown
Contributor

Fix asyncio ps showing tasks from only one interpreter

For details see gh-158880

Comment thread Lib/test/test_external_inspection.py Outdated
loop.run_in_executor(pool, _asyncio_in_subinterpreter)
task = asyncio.create_task(main_worker(), name="main_worker")
self.addCleanup(task.cancel)
await asyncio.sleep(1)

@maurycy maurycy Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.sleep feels like a flaky risk. What do you think about using busy_retry instead?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I fixed it


// Process all threads
if (iterate_threads(self, process_thread_for_async_stack_trace, result) < 0) {
if (iterate_threads(self, self->interpreter_addr,

@maurycy maurycy Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless I'm missing something, get_async_stack_trace() still has the same problem. Maybe it should be a follow-up, though.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, it has the problem, i planned to open a separate issue for this to avoid "big-bang" pr

return NULL;
}
if (refresh_generation_caches_for_interpreter(self, self->interpreter_addr) < 0) {
PyObject *seen = PySet_New(NULL);

@maurycy maurycy Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is just a list:

uintptr_t current_interpreter = self->interpreter_addr;
while (current_interpreter != 0) {

current_interpreter = GET_MEMBER(uintptr_t, interp_state_buffer,
self->debug_offsets.interpreter_state.next);

const size_t MAX_INTERPRETERS = 256;
size_t interp_count = 0;
while (current_interp != 0 && interp_count < MAX_INTERPRETERS) {

Is there a reproducible race scenario where we'd see a cycle? (It was quiet easy to reproduce issues like ABA.)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I couldn't reproduce a cycle. I thought the interpreter walk could loop
because of the previous PR, that's why I added the set. I agree with your suggestion and will rework it soon

Comment on lines +523 to +524
for info in RemoteUnwinder(
os.getpid()).get_all_awaited_by()

@maurycy maurycy Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One way to silence https://github.com/python/cpython/actions/runs/37499870518/job/112393935495#step:6:1023 is by wrapping this in except TRANSIENT_ERRORS:

Perhaps a better approach, similar to #158801

from _queue import SimpleQueue
from test import support
go = threading.Lock()
stop = threading.Lock()
go.acquire()
stop.acquire()
ready = SimpleQueue()
def leaf():
ready.put(None)
stop.acquire()
def start_leaf():
ready.put(None)
go.acquire()
leaf()
def park():
ready.put(None)
stop.acquire()

# SimpleQueue.put() and Lock.acquire() do not push Python frames.
# Once notified, the worker's stack stays stable until go is released.
threading.Thread(target=leaf, daemon=True).start()
ready.get(timeout=support.SHORT_TIMEOUT)
for _ in range(16):
threading.Thread(target=park, daemon=True).start()
ready.get(timeout=support.SHORT_TIMEOUT)
threading.Thread(target=start_leaf, daemon=True).start()
ready.get(timeout=support.SHORT_TIMEOUT)
u = RemoteUnwinder(os.getpid(), all_threads=True, cache_frames=False)
assert leaf_count(u) == 1
go.release()
ready.get(timeout=support.SHORT_TIMEOUT)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants