Skip to content

Commit f4a13b1

Browse files
committed
gh-158790: Improve performance for list.insert() and del list[i] in free-threaded builds
`ins1()`(used by `list.insert()`) and `list_ass_item_lock_held()` shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than `memmove()` for large shifts. I propose using the existing `ptr_wise_atomic_memmove()` helper here. It uses `memmove()` as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe. The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks. pyperf, free-threaded release build (--disable-gil, -O3): del l[0] 2.26x faster del l[len(l)//2] 2.14x faster l.insert(0, x) 1.95x faster l.insert(len(l)//2, x) 1.86x faster insert(0)+del[0], n=1000 2.86x faster sliding window, n=1000 2.46x faster Compiling both versions in GIL mode produces the same code for `ins1()` and `list_ass_item_lock_held()`, so the default build is not affected. This continues gh-129069, which introduced `ptr_wise_atomic_memmove()`.
1 parent 9d22a53 commit f4a13b1

2 files changed

Lines changed: 22 additions & 2 deletions

File tree

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
Speed up ``list.insert()`` and ``del list[i]`` in free-threaded builds by using ``memmove()`` to shift elements when the list is not shared.

‎Objects/listobject.c‎

Lines changed: 21 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -478,10 +478,13 @@ end:;
478478
return ret;
479479
}
480480

481+
static void ptr_wise_atomic_memmove(PyListObject *a, PyObject **dest,
482+
PyObject **src, Py_ssize_t n);
483+
481484
static int
482485
ins1(PyListObject *self, Py_ssize_t where, PyObject *v)
483486
{
484-
Py_ssize_t i, n = Py_SIZE(self);
487+
Py_ssize_t n = Py_SIZE(self);
485488
PyObject **items;
486489
if (v == NULL) {
487490
PyErr_BadInternalCall();
@@ -500,8 +503,17 @@ ins1(PyListObject *self, Py_ssize_t where, PyObject *v)
500503
if (where > n)
501504
where = n;
502505
items = self->ob_item;
503-
for (i = n; --i >= where; )
506+
#ifdef Py_GIL_DISABLED
507+
if (where < n) {
508+
ptr_wise_atomic_memmove(self, &items[where + 1], &items[where],
509+
n - where);
510+
}
511+
#else
512+
/* The compiler expands this loop inline. This is faster than a
513+
memmove() call for short lists. */
514+
for (Py_ssize_t i = n; --i >= where; )
504515
FT_ATOMIC_STORE_PTR_RELEASE(items[i+1], items[i]);
516+
#endif
505517
FT_ATOMIC_STORE_PTR_RELEASE(items[where], Py_NewRef(v));
506518
return 0;
507519
}
@@ -1145,9 +1157,16 @@ list_ass_item_lock_held(PyListObject *a, Py_ssize_t i, PyObject *v)
11451157
PyObject *tmp = a->ob_item[i];
11461158
if (v == NULL) {
11471159
Py_ssize_t size = Py_SIZE(a);
1160+
#ifdef Py_GIL_DISABLED
1161+
if (i < size - 1) {
1162+
ptr_wise_atomic_memmove(a, &a->ob_item[i], &a->ob_item[i + 1],
1163+
size - 1 - i);
1164+
}
1165+
#else
11481166
for (Py_ssize_t idx = i; idx < size - 1; idx++) {
11491167
FT_ATOMIC_STORE_PTR_RELEASE(a->ob_item[idx], a->ob_item[idx + 1]);
11501168
}
1169+
#endif
11511170
Py_SET_SIZE(a, size - 1);
11521171
}
11531172
else {

0 commit comments

Comments
 (0)