From 0ac5047099019769adf0091d5c58fdf8e37cc572 Mon Sep 17 00:00:00 2001 From: Iliyan Velichkov Date: Mon, 5 Oct 2026 16:02:09 +0300 Subject: [PATCH] ci: split the api IT shard into two hash partitions so it fits its 90-minute budget The HTTP-level "api" shard (!ui & !slow) has grown to ~135 classes and ~83 min of test time, and since 2026-10-01 it has been cancelled at the 90-minute job cap on nearly every master push and nightly, on both databases, with every class it reached green. Its classes cost a near-uniform 40-60 s each, so no tag can move a long pole out of it. -Dit.partition=k/n now keeps the selected classes whose top-level class name hashes into partition k of n: ClassPartitionFilter, a JUnit PostDiscoveryFilter registered through META-INF/services and applied after the tag filter. The partitions are disjoint and complete by construction - a new IT lands in exactly one of them without being tagged - and nested classes stay with their enclosing class. Without the property every class is kept, so the PR smoke gate and local runs are unchanged. build.yml and nightly.yml run "api" as api-1 and api-2 (1/2, 2/2) on both the H2 and the PostgreSQL leg. Co-Authored-By: Claude Opus 5.5 --- .claude/docs/ci.md | 4 +- .github/workflows/build.yml | 33 +++-- .github/workflows/nightly.yml | 22 ++- pom.xml | 4 + tests/tests-integrations/pom.xml | 11 ++ .../tests/support/ClassPartitionFilter.java | 128 ++++++++++++++++++ ...unit.platform.launcher.PostDiscoveryFilter | 1 + .../support/ClassPartitionFilterTest.java | 109 +++++++++++++++ 8 files changed, 294 insertions(+), 18 deletions(-) create mode 100644 tests/tests-integrations/src/main/java/org/eclipse/dirigible/integration/tests/support/ClassPartitionFilter.java create mode 100644 tests/tests-integrations/src/main/resources/META-INF/services/org.junit.platform.launcher.PostDiscoveryFilter create mode 100644 tests/tests-integrations/src/test/java/org/eclipse/dirigible/integration/tests/support/ClassPartitionFilterTest.java diff --git a/.claude/docs/ci.md b/.claude/docs/ci.md index e22a60c3eba..6e67dc20173 100644 --- a/.claude/docs/ci.md +++ b/.claude/docs/ci.md @@ -4,7 +4,7 @@ - `code-style`: `mvn -T 1C formatter:validate` - `tests` (ubuntu + windows matrix): `mvn clean install -P unit-tests` -- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow`), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below). +- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow & !upgrade`, itself run as two hash partitions `api-1` / `api-2`, see the budget note below), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below). - `build-deploy`: `mvn clean install -P quick-build` then Docker buildx multi-arch image push to `dirigiblelabs/dirigible` ### PR gate vs full suite (smoke / nightly split) @@ -17,7 +17,7 @@ The full Selenide UI suite takes ~1.5h per DB, so it does **not** run on every P **The job budget, and how a shard that runs out of it reads (#7283).** `timeout-minutes` is **90** on every IT job (150 on the PR `smoke-tests` job) - comfortably above the measured times above, because a job cancelled at its cap reads as infrastructure rather than as a test result. At 60 the PostgreSQL `api` shard of the nightly was killed **two classes short of finishing** and that looked exactly like a hang. Two things to know before concluding "wedged IT" from a cancelled shard: - **Check the last `[INFO] Running org.eclipse...` line against the cancellation timestamp.** In the run that raised the issue the JVM was healthy: it kept starting one IT class per ~40 s until 10 minutes before the cancel, and the class suspected of hanging (`EdmModelRoundTripIT`) had passed in 50.51 s. Searching the downloaded log for a class name lands on its *last* mention, which is its own completion line - everything after it belongs to the next classes. -- **`api` cannot be rebalanced with `@Tag("slow")`.** Its 83 classes cost a near-uniform ~40 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows ~40 s per IT class added and has no long pole to move out. Splitting it needs a partition mechanism, not a tag - and a name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. +- **`api` cannot be rebalanced with `@Tag("slow")`.** Its classes cost a near-uniform ~40-60 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows by that much per IT class added and has no long pole to move out. A name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. So it is split by a **partition**: by October 2026 it had grown to ~133 classes / ~83 min and was cancelled at the 90-minute cap on every master and nightly run, and it now runs as `api-1` / `api-2`, each `-Dit.groups="!ui & !slow & !upgrade" -Dit.partition=k/2`. `ClassPartitionFilter` (tests-integrations `support/`, a JUnit `PostDiscoveryFilter` registered through `META-INF/services`, applied after the tag filter) keeps the classes whose **top-level** class name hashes into partition `k` of `n` - disjoint and complete by construction, nested classes stay with their enclosing class, nothing to tag. Without `-Dit.partition` it keeps everything. When a partition nears the cap again, go to `/3` rather than reshuffling tags; `-Djunit.platform.execution.dryRun.enabled=true` counts what each partition selects in about a minute. **A hang costs one class, not the shard.** `IntegrationTest` carries a class-level **`@Timeout(15, MINUTES)`** (inherited; per test *and* per lifecycle method - the Spring context start is an extension callback and is not counted), so a method that never returns fails its own class and the run continues and reports. The slowest measured method in the suite is 456 s (`CreateNewFileIT`, one `@Test`), so 15 minutes is roughly twice the worst legitimate case; a class needing a different bound declares its own `@Timeout`, which overrides the inherited one (`JavaLspIT`, `JavaDebugIT`, `SchemaTemplateForeignKeyIT` each declare a tighter one). There is deliberately **no `forkedProcessTimeoutInSeconds`**: it bounds the whole fork, and the jobs sharing this configuration range from 13 to 91 minutes, so one value would either be useless or kill the longest job. diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index 23619a66441..4fb3c972c7e 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -162,7 +162,8 @@ jobs: # The full IT suite is split ("sharded") into four disjoint slices selected by JUnit 5 tags, so # the wall-clock time of a run is the slowest shard instead of the whole suite (~2h30 serial): - # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles + # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles, + # run as two jobs (api-1, api-2) - see the partition note below # ui - the browser-driven ITs except the sample-project, camel and long-pole families # samples - the sample-project clone/publish ITs ("sample") + the camel journeys still driven # through the browser ("camel"), which is now only the two template starters - the @@ -181,24 +182,34 @@ jobs: # samples 13 / 13. "api" is the critical path and the "slow" tag cannot rebalance it: its 83 # classes cost a near-uniform ~40 s each, almost all of it the Spring context each class boots # (@DirtiesContext AFTER_EACH_TEST_METHOD), so the shard grows by ~40 s per IT class added and - # has no long pole to move out. #7283 raised the budget rather than split it; splitting "api" - # needs a partition mechanism, not a tag. + # has no long pole to move out. #7283 raised the budget rather than split it, and by October 2026 + # the shard (~133 classes, ~83 min of test time) was cancelled at the 90-minute cap on every run. + # So "api" is split by a partition, not by a tag: -Dit.partition=k/n keeps the classes whose + # top-level class name hashes into partition k of n (ClassPartitionFilter in tests-integrations, a + # JUnit PostDiscoveryFilter applied after the tag filter). The partitions are disjoint and complete + # by construction - a new IT lands in exactly one of them without being tagged - and, the class + # costs being near-uniform, about even (on the 2026-10-05 h2 shard log: 64 classes / 38.5 min + # against 69 / 44.7). Add a third partition (1/3, 2/3, 3/3) when one of them nears the cap again. integration-tests-h2: runs-on: ubuntu-latest strategy: fail-fast: false matrix: shard: - - name: api + - name: api-1 groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 + groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples groups: "ui & !slow & (sample | camel)" - name: slow groups: "slow" - # The slowest green shard ("api") takes ~57 min on h2 and ~62 on PostgreSQL; a heap-exhausted - # JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs + # The slowest green shard ("api", before its split) took ~57 min on h2 and ~62 on PostgreSQL; a + # heap-exhausted JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs # fail fast instead of burning runners - but far enough above it that a routine run is never # cancelled: at 60 the PostgreSQL "api" shard was killed two classes short of finishing, which # reads as infrastructure rather than as a test result (#7283). A single method can no longer @@ -240,7 +251,7 @@ jobs: run: ttyd --version - name: Integration tests - run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. @@ -292,8 +303,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 + groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -377,7 +392,7 @@ jobs: echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)" exit 1 fi - mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. diff --git a/.github/workflows/nightly.yml b/.github/workflows/nightly.yml index e0008c76854..d22bc7e237f 100644 --- a/.github/workflows/nightly.yml +++ b/.github/workflows/nightly.yml @@ -5,9 +5,9 @@ name: Nightly Integration Tests # PRs run the fast "smoke-tests" job (see pull-request.yml). Runs on a nightly schedule and on # demand; the same full suite also runs on every push to master (see build.yml). Like build.yml, the # suite is sharded into four tag-selected slices (api / ui / samples / slow) that run in parallel, -# so the wall-clock time is the slowest shard ("api", ~57 min on h2 and ~62 on PostgreSQL), not the -# whole suite. build.yml carries the description of what each shard selects and the measured shard -# times behind the job timeout; keep the two matrices identical. +# "api" itself as two hash partitions (api-1 / api-2), so the wall-clock time is the slowest shard, +# not the whole suite. build.yml carries the description of what each shard selects and the +# measured shard times behind the job timeout; keep the two matrices identical. on: schedule: - cron: '0 2 * * *' @@ -24,8 +24,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 + groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -69,7 +73,7 @@ jobs: run: ttyd --version - name: Integration tests - run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. @@ -183,8 +187,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 + groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -268,7 +276,7 @@ jobs: echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)" exit 1 fi - mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. diff --git a/pom.xml b/pom.xml index 798cb735fef..0d8362f73e0 100644 --- a/pom.xml +++ b/pom.xml @@ -290,6 +290,10 @@ plus the few explicitly-smoke UI journeys. --> + + diff --git a/tests/tests-integrations/pom.xml b/tests/tests-integrations/pom.xml index 6168de9a102..d7e372e37be 100644 --- a/tests/tests-integrations/pom.xml +++ b/tests/tests-integrations/pom.xml @@ -30,6 +30,12 @@ junit-platform-suite compile + + + org.junit.platform + junit-platform-launcher + compile + org.eclipse.dirigible dirigible-application @@ -138,6 +144,11 @@ ${it.groups} ${it.excludedGroups} + + + ${it.partition} +