Skip to content

Commit 00a523f

Browse files
dschoGit for Windows Build Agent
authored andcommitted
parallel-checkout: fix stack buffer overflow in Windows poll() with many workers (#6395)
### Symptom On Windows, `git checkout` and `git reset --hard` can abort with ``` *** stack smashing detected ***: terminated ``` and exit code `0xC0000409` (`STATUS_STACK_BUFFER_OVERRUN`). This is memory corruption, not a normal error. The process dies before Trace2 writes its log, so nothing shows up in a trace. A `.git/index.lock` is left behind. It happens when `checkout.workers` is large, or when it is `0` (meaning "use `online_cpus()`") on a machine with many logical processors. ### Mechanism `gather_results_from_workers()` in `parallel-checkout.c` polls one pipe per checkout worker: ```c CALLOC_ARRAY(pfds, num_workers); ... poll(pfds, num_workers, -1); ``` Windows has no native `poll()`, so `compat/poll/poll.c` emulates it with `MsgWaitForMultipleObjects()`. It collects one handle per polled descriptor in a fixed stack array: ```c HANDLE h, handle_array[FD_SETSIZE + 2]; /* 64 + 2 = 66 entries */ ... handle_array[nhandles++] = h; /* no bounds check */ ... handle_array[nhandles] = NULL; /* sentinel, no bounds check */ ``` `FD_SETSIZE` is the Winsock default 64, and nothing in the build overrides it. `run_parallel_checkout()` clamps `num_workers` only against the number of files, never against the array size or the Windows wait limit. A high worker count therefore writes past the end of the array and smashes the stack. Sockets are not involved: they are multiplexed onto a single event through `WSAEventSelect`, so only non-socket descriptors consume a slot. ### Why 62, and where the limit lives Two of the wait slots are never available for descriptors: * `compat/poll` uses index 0 for its own event object. * `QS_ALLINPUT` adds the thread message queue as an implicit wait object. The code confirms this, because it reports the message queue as `WAIT_OBJECT_0 + nhandles`. So `nhandles + 1 <= MAXIMUM_WAIT_OBJECTS`, which gives at most `MAXIMUM_WAIT_OBJECTS - 2` = 62 descriptors. That value is defined once, as `POLL_MAX_DESCRIPTORS` in `compat/poll/poll.h` alongside the `poll()` declaration it constrains, and both callers clamp against it. `compat/posix.h` defines it to `INT_MAX` where a native `poll()` is used, so the callers need no `#ifdef`. ### Why the array cannot simply be enlarged `MAXIMUM_WAIT_OBJECTS` is a kernel limit, not a header convenience. Passing more handles fails with `ERROR_INVALID_PARAMETER`. Growing the array would only turn memory corruption into a functional failure. Support for more descriptors needs a wait tree (helper threads each waiting on at most 62 handles) or completion ports, which is out of scope here. ### Why it surfaced in 2.54 `parallel-checkout.c` and `compat/poll/poll.c` are unchanged between 2.53 and 2.54. Only `online_cpus()` changed: | Version | API | Result | | --- | --- | --- | | 2.53 | `GetSystemInfo()` | processors in the current processor group only; a group holds at most 64 | | 2.54+ | `GetLogicalProcessorInformationEx()` | true system-wide logical processor count | The old API could never report more than 64, so the array always fit. That ceiling was accidental, not deliberate. The `online_cpus()` change is correct and must stay; it only exposed a latent bug. ### The changes 1. **`compat/poll: do not collect more handles than the wait supports`** — defines `POLL_MAX_DESCRIPTORS` (62) next to the `poll()` declaration and refuses to collect beyond it, returning `EINVAL` instead of appending past the end of the array. Two preprocessor assertions tie the constant to `MAXIMUM_WAIT_OBJECTS` and to the size of `handle_array`, so they cannot drift apart. `poll()` is now memory-safe for every input. The error path also undoes the `WSAEventSelect()` registrations made earlier in the same call. The loop that normally does that runs after the wait, so returning early would otherwise leave those sockets bound to `poll()`'s static event object and let later socket activity disturb an unrelated `poll()`. The limit is on the handles actually collected, **not** on `nfd`. Those are different: a descriptor only takes a handle when it is non-negative, is not a socket, and has no events pending yet. Callers routinely pass sparse arrays — `run_processes_parallel()` sizes its `pollfd` array to the configured job count and leaves the unused slots at `fd = -1`. An earlier revision of this PR rejected a large `nfd` instead, which broke `t7406` (`submodule.fetchJobs 67`, with only a handful of live pipes) with `fatal: poll: Invalid argument`. On platforms with a native `poll()` there is no such limit, so `POLL_MAX_DESCRIPTORS` is `INT_MAX` and callers can clamp against it unconditionally. 2. **`parallel-checkout: limit worker count to what poll() can wait on`** — clamps `num_workers` to `POLL_MAX_DESCRIPTORS` in `run_parallel_checkout()`, the single choke point before the workers start. The clamp is silent: fewer workers is correct, and a warning would fire on every checkout on a large machine. Also documents the cap, since `checkout.workers` was described as using one worker per logical core with no upper bound. 3. **`run-command: limit concurrent children to what poll() can wait on`** — the same limit applied to the other `poll()` fan-out. `pp_buffer_io()` polls one pipe per child sending output and a second per child being fed on stdin, and `fetch.parallel` / `submodule.fetchJobs` / `hook.jobs` all accept a high value (or `0`, meaning `online_cpus()`). Without this, change 1 would convert the old stack smash on that path into a hard `die_errno("poll")`. Only *concurrency* is limited; the configured maximum still sizes the arrays and is still reported by the trace, so the total number of tasks run is unchanged. ### Reproduction No clone, no special hardware, about 10 seconds. A many-core machine is not required: a positive `checkout.workers` is used verbatim, and `online_cpus()` is consulted only when the value is `0` or less. ```powershell # Use a NEW directory every attempt (see the note on timing below). $repo = "C:\tmp\poll-repro-$(Get-Random)" New-Item -ItemType Directory -Force -Path $repo | Out-Null Set-Location $repo git init -q -b main . git config user.email repro@example.com git config user.name repro git config checkout.workers 200 git config checkout.thresholdForParallelism 1 New-Item -ItemType Directory -Force -Path dir | Out-Null 1..400 | ForEach-Object { Set-Content -Path "dir\f$_.txt" -Value "base $_" -NoNewline } git add -A; git commit -qm base git checkout -qb other 1..400 | ForEach-Object { Set-Content -Path "dir\f$_.txt" -Value "changed $_ padding padding padding" -NoNewline } git commit -qam changed git checkout -q main Write-Host "exit=$LASTEXITCODE" ``` Before the fix, on a 12-core machine, this crashed 3 out of 3 runs with `exit=-1073740791` (`0xC0000409`) and left `.git/index.lock` behind. After the fix it exits 0 on 3 out of 3 runs, with the files correctly updated. A `checkout.workers 16` checkout still works, as before. ### The crash is not deterministic `poll()` only appends a descriptor when the worker's pipe has no data ready yet, so `nhandles` reflects the workers pending at that instant, not the workers spawned. On warm cache, pipes answer immediately and few workers stay pending. Measured before the fix: | workers | result | | --- | --- | | 16, 64, 65, 70, 72, 74, 76, 78 | pass | | 80 | crashed once, then passed 3 times | | 200, repeated checkouts in the same repo | passed 4 times | | 200, fresh repository each run | crashed 3 of 3 | The first out-of-bounds write happens at 65 descriptors by arithmetic, but the corruption does not reliably reach the stack cookie until well past that. The corruption is real from 65 onward whether or not it crashes. That is why the fix targets the contract (62), not the observed crash point. For the same reason, the added test asserts the **effective worker count** rather than a crash. `test_checkout_workers` counts the workers actually spawned from a TRACE2 log, so the test verifies the clamp took effect (62) instead of merely checking that the checkout did not crash. It is `MINGW`-gated, because that is the only platform where the cap applies. ### Testing * New test in `t/t2080-parallel-checkout-basics.sh`. `t2080`, `t2081`, `t2082`, `t0061`, `t5526` and `t7406` all pass on Windows. * Manual verification with the reproduction above, plus a low-worker-count regression check. ### Workaround for affected users ``` git config checkout.workers 16 ``` Any value at or below 62 avoids the overflow. No downgrade is needed.
2 parents 1657875 + 860dd91 commit 00a523f

17 files changed

Lines changed: 410 additions & 3 deletions

‎Documentation/config/checkout.adoc‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -30,6 +30,11 @@ commands or functionality in the future.
3030
all commands that perform checkout. E.g. checkout, clone, reset,
3131
sparse-checkout, etc.
3232
+
33+
On Windows the number of workers is capped at 62, because the `poll()`
34+
emulation cannot wait on more worker pipes than that. A higher configured
35+
value, including the logical core count on a machine with many cores, is
36+
silently reduced to the cap.
37+
+
3338
NOTE: Parallel checkout usually delivers better performance for repositories
3439
located on SSDs or over NFS. For repositories on spinning disks and/or machines
3540
with a small number of cores, the default sequential checkout often performs

‎Makefile‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1553,6 +1553,7 @@ CLAR_TEST_SUITES += u-odb-inmemory
15531553
CLAR_TEST_SUITES += u-oid-array
15541554
CLAR_TEST_SUITES += u-oidmap
15551555
CLAR_TEST_SUITES += u-oidtree
1556+
CLAR_TEST_SUITES += u-poll
15561557
CLAR_TEST_SUITES += u-prio-queue
15571558
CLAR_TEST_SUITES += u-reftable-basics
15581559
CLAR_TEST_SUITES += u-reftable-block

‎builtin/fetch.c‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2351,6 +2351,7 @@ static int fetch_multiple(struct string_list *list, int max_children,
23512351
.tr2_label = "parallel/fetch",
23522352

23532353
.processes = max_children,
2354+
.no_stdin_pipe = 1,
23542355

23552356
.get_next_task = &fetch_next_remote,
23562357
.start_failure = &fetch_failed_to_start,

‎builtin/submodule--helper.c‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2919,6 +2919,7 @@ static int update_submodules(struct update_data *update_data)
29192919
.tr2_label = "parallel/update",
29202920

29212921
.processes = update_data->max_jobs,
2922+
.no_stdin_pipe = 1,
29222923

29232924
.get_next_task = update_clone_get_next_task,
29242925
.start_failure = update_clone_start_failure,

‎compat/poll/poll.c‎

Lines changed: 47 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -303,6 +303,43 @@ compute_revents (int fd, int sought, fd_set *rfds, fd_set *wfds, fd_set *efds)
303303
}
304304
#endif /* !MinGW */
305305

306+
#ifdef WIN32_NATIVE
307+
/* POLL_MAX_DESCRIPTORS descriptors, plus hEvent and the QS_ALLINPUT message
308+
queue, must fit in one MsgWaitForMultipleObjects call, and the collected
309+
handles plus the NULL sentinel must fit in handle_array. */
310+
#if POLL_MAX_DESCRIPTORS + 2 > MAXIMUM_WAIT_OBJECTS
311+
#error POLL_MAX_DESCRIPTORS exceeds MAXIMUM_WAIT_OBJECTS
312+
#endif
313+
#if POLL_MAX_DESCRIPTORS + 2 > FD_SETSIZE + 2
314+
#error POLL_MAX_DESCRIPTORS does not fit in handle_array
315+
#endif
316+
317+
/* Undo the WSAEventSelect() calls made for the first NFD descriptors. */
318+
static void
319+
reset_socket_events (struct pollfd *pfd, nfds_t nfd)
320+
{
321+
nfds_t i;
322+
323+
for (i = 0; i < nfd; i++)
324+
{
325+
HANDLE h;
326+
327+
if (pfd[i].fd < 0)
328+
continue;
329+
if (!(pfd[i].events & (POLLIN | POLLRDNORM | POLLOUT | POLLWRNORM |
330+
POLLWRBAND | POLLPRI | POLLRDBAND)))
331+
continue;
332+
333+
h = (HANDLE) _get_osfhandle (pfd[i].fd);
334+
if (h == NULL || h == INVALID_HANDLE_VALUE)
335+
continue;
336+
337+
if (IsSocketHandle (h))
338+
WSAEventSelect ((SOCKET) h, NULL, 0);
339+
}
340+
}
341+
#endif
342+
306343
int
307344
poll (struct pollfd *pfd, nfds_t nfd, int timeout)
308345
{
@@ -504,7 +541,16 @@ poll (struct pollfd *pfd, nfds_t nfd, int timeout)
504541
bits for the "wrong" direction. */
505542
pfd[i].revents = win32_compute_revents (h, &sought);
506543
if (sought)
507-
handle_array[nhandles++] = h;
544+
{
545+
/* hEvent occupies handle_array[0]. See POLL_MAX_DESCRIPTORS. */
546+
if (nhandles > POLL_MAX_DESCRIPTORS)
547+
{
548+
reset_socket_events (pfd, i);
549+
errno = EINVAL;
550+
return -1;
551+
}
552+
handle_array[nhandles++] = h;
553+
}
508554
if (pfd[i].revents)
509555
timeout = 0;
510556
}

‎compat/poll/poll.h‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -59,6 +59,22 @@ typedef unsigned long nfds_t;
5959

6060
extern int poll (struct pollfd *pfd, nfds_t nfd, int timeout);
6161

62+
#if (defined _WIN32 || defined __WIN32__) && ! defined __CYGWIN__
63+
/*
64+
* This poll() is emulated with MsgWaitForMultipleObjects(), which waits on at
65+
* most MAXIMUM_WAIT_OBJECTS (64) objects. Two of those are never available for
66+
* polled descriptors: poll() waits on its own event object, and QS_ALLINPUT
67+
* adds the thread message queue. Sockets do not count, because they are all
68+
* multiplexed onto that one event object; every other descriptor takes a wait
69+
* slot of its own.
70+
*
71+
* Callers that poll one or more descriptors per child must keep the number of
72+
* simultaneously live descriptors within this limit. Exceeding it fails with
73+
* EINVAL.
74+
*/
75+
#define POLL_MAX_DESCRIPTORS 62
76+
#endif
77+
6278
/* Define INFTIM only if doing so conforms to POSIX. */
6379
#if !defined (_POSIX_C_SOURCE) && !defined (_XOPEN_SOURCE)
6480
#define INFTIM (-1)

‎compat/posix.h‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -133,6 +133,16 @@
133133
/* Pull the compat stuff */
134134
#include <poll.h>
135135
#endif
136+
137+
/*
138+
* compat/poll defines POLL_MAX_DESCRIPTORS to the largest number of
139+
* descriptors its poll() emulation can wait on. A native poll() has no such
140+
* limit, so callers that fan out one descriptor per child can clamp against
141+
* this unconditionally.
142+
*/
143+
#ifndef POLL_MAX_DESCRIPTORS
144+
#define POLL_MAX_DESCRIPTORS INT_MAX
145+
#endif
136146
#ifdef HAVE_BSD_SYSCTL
137147
#include <sys/sysctl.h>
138148
#endif

‎hook.c‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -798,6 +798,7 @@ int run_hooks_opt(struct repository *r, const char *hook_name,
798798

799799
.processes = jobs,
800800
.ungroup = jobs == 1,
801+
.no_stdin_pipe = !options->feed_pipe,
801802

802803
.get_next_task = pick_next_hook,
803804
.start_failure = notify_start_failure,

‎parallel-checkout.c‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -671,6 +671,13 @@ int run_parallel_checkout(struct checkout *state, int num_workers, int threshold
671671
if (parallel_checkout.nr < num_workers)
672672
num_workers = parallel_checkout.nr;
673673

674+
/*
675+
* gather_results_from_workers() polls one pipe per worker, so the
676+
* worker count must stay within what poll() can wait on.
677+
*/
678+
if (num_workers > POLL_MAX_DESCRIPTORS)
679+
num_workers = POLL_MAX_DESCRIPTORS;
680+
674681
if (num_workers <= 1 || parallel_checkout.nr < threshold) {
675682
write_items_sequentially(state);
676683
} else {

‎run-command.c‎

Lines changed: 19 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1658,6 +1658,9 @@ static int pp_start_one(struct parallel_processes *pp,
16581658
}
16591659
return 1;
16601660
}
1661+
if (opts->no_stdin_pipe && pp->children[i].process.in < 0)
1662+
BUG("get_next_task requested a stdin pipe despite "
1663+
"no_stdin_pipe");
16611664
if (!opts->ungroup) {
16621665
pp->children[i].process.err = -1;
16631666
pp->children[i].process.stdout_to_stderr = 1;
@@ -1893,6 +1896,7 @@ void run_processes_parallel(const struct run_process_parallel_opts *opts)
18931896
int i, code;
18941897
int timeout = 100;
18951898
int spawn_cap = 4;
1899+
size_t max_live;
18961900
struct parallel_processes_for_signal pp_sig;
18971901
struct parallel_processes pp = {
18981902
.buffered_output = STRBUF_INIT,
@@ -1902,6 +1906,20 @@ void run_processes_parallel(const struct run_process_parallel_opts *opts)
19021906
const char *tr2_label = opts->tr2_label;
19031907
const int do_trace2 = tr2_category && tr2_label;
19041908

1909+
/*
1910+
* Unless the caller handles its own output, pp_buffer_io() polls one
1911+
* output pipe per child and, unless excluded by no_stdin_pipe, may also
1912+
* poll an input pipe. Limit the number of live children so that all of
1913+
* their descriptors fit in one poll() call.
1914+
*/
1915+
max_live = opts->processes;
1916+
if (!opts->ungroup) {
1917+
size_t fds_per_process = opts->no_stdin_pipe ? 1 : 2;
1918+
1919+
if (max_live > POLL_MAX_DESCRIPTORS / fds_per_process)
1920+
max_live = POLL_MAX_DESCRIPTORS / fds_per_process;
1921+
}
1922+
19051923
if (do_trace2)
19061924
trace2_region_enter_printf(tr2_category, tr2_label, NULL,
19071925
"max:%"PRIuMAX,
@@ -1923,7 +1941,7 @@ void run_processes_parallel(const struct run_process_parallel_opts *opts)
19231941
while (1) {
19241942
for (i = 0;
19251943
i < spawn_cap && !pp.shutdown &&
1926-
pp.nr_processes < opts->processes;
1944+
pp.nr_processes < max_live;
19271945
i++) {
19281946
code = pp_start_one(&pp, opts);
19291947
if (!code)

0 commit comments

Comments
 (0)