diff --git a/.claude/docs/ci.md b/.claude/docs/ci.md index e22a60c3eba..6e67dc20173 100644 --- a/.claude/docs/ci.md +++ b/.claude/docs/ci.md @@ -4,7 +4,7 @@ - `code-style`: `mvn -T 1C formatter:validate` - `tests` (ubuntu + windows matrix): `mvn clean install -P unit-tests` -- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow`), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below). +- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow & !upgrade`, itself run as two hash partitions `api-1` / `api-2`, see the budget note below), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below). - `build-deploy`: `mvn clean install -P quick-build` then Docker buildx multi-arch image push to `dirigiblelabs/dirigible` ### PR gate vs full suite (smoke / nightly split) @@ -17,7 +17,7 @@ The full Selenide UI suite takes ~1.5h per DB, so it does **not** run on every P **The job budget, and how a shard that runs out of it reads (#7283).** `timeout-minutes` is **90** on every IT job (150 on the PR `smoke-tests` job) - comfortably above the measured times above, because a job cancelled at its cap reads as infrastructure rather than as a test result. At 60 the PostgreSQL `api` shard of the nightly was killed **two classes short of finishing** and that looked exactly like a hang. Two things to know before concluding "wedged IT" from a cancelled shard: - **Check the last `[INFO] Running org.eclipse...` line against the cancellation timestamp.** In the run that raised the issue the JVM was healthy: it kept starting one IT class per ~40 s until 10 minutes before the cancel, and the class suspected of hanging (`EdmModelRoundTripIT`) had passed in 50.51 s. Searching the downloaded log for a class name lands on its *last* mention, which is its own completion line - everything after it belongs to the next classes. -- **`api` cannot be rebalanced with `@Tag("slow")`.** Its 83 classes cost a near-uniform ~40 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows ~40 s per IT class added and has no long pole to move out. Splitting it needs a partition mechanism, not a tag - and a name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. +- **`api` cannot be rebalanced with `@Tag("slow")`.** Its classes cost a near-uniform ~40-60 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows by that much per IT class added and has no long pole to move out. A name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. So it is split by a **partition**: by October 2026 it had grown to ~133 classes / ~83 min and was cancelled at the 90-minute cap on every master and nightly run, and it now runs as `api-1` / `api-2`, each `-Dit.groups="!ui & !slow & !upgrade" -Dit.partition=k/2`. `ClassPartitionFilter` (tests-integrations `support/`, a JUnit `PostDiscoveryFilter` registered through `META-INF/services`, applied after the tag filter) keeps the classes whose **top-level** class name hashes into partition `k` of `n` - disjoint and complete by construction, nested classes stay with their enclosing class, nothing to tag. Without `-Dit.partition` it keeps everything. When a partition nears the cap again, go to `/3` rather than reshuffling tags; `-Djunit.platform.execution.dryRun.enabled=true` counts what each partition selects in about a minute. **A hang costs one class, not the shard.** `IntegrationTest` carries a class-level **`@Timeout(15, MINUTES)`** (inherited; per test *and* per lifecycle method - the Spring context start is an extension callback and is not counted), so a method that never returns fails its own class and the run continues and reports. The slowest measured method in the suite is 456 s (`CreateNewFileIT`, one `@Test`), so 15 minutes is roughly twice the worst legitimate case; a class needing a different bound declares its own `@Timeout`, which overrides the inherited one (`JavaLspIT`, `JavaDebugIT`, `SchemaTemplateForeignKeyIT` each declare a tighter one). There is deliberately **no `forkedProcessTimeoutInSeconds`**: it bounds the whole fork, and the jobs sharing this configuration range from 13 to 91 minutes, so one value would either be useless or kill the longest job. diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index 23619a66441..4fb3c972c7e 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -162,7 +162,8 @@ jobs: # The full IT suite is split ("sharded") into four disjoint slices selected by JUnit 5 tags, so # the wall-clock time of a run is the slowest shard instead of the whole suite (~2h30 serial): - # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles + # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles, + # run as two jobs (api-1, api-2) - see the partition note below # ui - the browser-driven ITs except the sample-project, camel and long-pole families # samples - the sample-project clone/publish ITs ("sample") + the camel journeys still driven # through the browser ("camel"), which is now only the two template starters - the @@ -181,24 +182,34 @@ jobs: # samples 13 / 13. "api" is the critical path and the "slow" tag cannot rebalance it: its 83 # classes cost a near-uniform ~40 s each, almost all of it the Spring context each class boots # (@DirtiesContext AFTER_EACH_TEST_METHOD), so the shard grows by ~40 s per IT class added and - # has no long pole to move out. #7283 raised the budget rather than split it; splitting "api" - # needs a partition mechanism, not a tag. + # has no long pole to move out. #7283 raised the budget rather than split it, and by October 2026 + # the shard (~133 classes, ~83 min of test time) was cancelled at the 90-minute cap on every run. + # So "api" is split by a partition, not by a tag: -Dit.partition=k/n keeps the classes whose + # top-level class name hashes into partition k of n (ClassPartitionFilter in tests-integrations, a + # JUnit PostDiscoveryFilter applied after the tag filter). The partitions are disjoint and complete + # by construction - a new IT lands in exactly one of them without being tagged - and, the class + # costs being near-uniform, about even (on the 2026-10-05 h2 shard log: 64 classes / 38.5 min + # against 69 / 44.7). Add a third partition (1/3, 2/3, 3/3) when one of them nears the cap again. integration-tests-h2: runs-on: ubuntu-latest strategy: fail-fast: false matrix: shard: - - name: api + - name: api-1 groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 + groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples groups: "ui & !slow & (sample | camel)" - name: slow groups: "slow" - # The slowest green shard ("api") takes ~57 min on h2 and ~62 on PostgreSQL; a heap-exhausted - # JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs + # The slowest green shard ("api", before its split) took ~57 min on h2 and ~62 on PostgreSQL; a + # heap-exhausted JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs # fail fast instead of burning runners - but far enough above it that a routine run is never # cancelled: at 60 the PostgreSQL "api" shard was killed two classes short of finishing, which # reads as infrastructure rather than as a test result (#7283). A single method can no longer @@ -240,7 +251,7 @@ jobs: run: ttyd --version - name: Integration tests - run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. @@ -292,8 +303,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 + groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -377,7 +392,7 @@ jobs: echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)" exit 1 fi - mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. diff --git a/.github/workflows/nightly.yml b/.github/workflows/nightly.yml index e0008c76854..d22bc7e237f 100644 --- a/.github/workflows/nightly.yml +++ b/.github/workflows/nightly.yml @@ -5,9 +5,9 @@ name: Nightly Integration Tests # PRs run the fast "smoke-tests" job (see pull-request.yml). Runs on a nightly schedule and on # demand; the same full suite also runs on every push to master (see build.yml). Like build.yml, the # suite is sharded into four tag-selected slices (api / ui / samples / slow) that run in parallel, -# so the wall-clock time is the slowest shard ("api", ~57 min on h2 and ~62 on PostgreSQL), not the -# whole suite. build.yml carries the description of what each shard selects and the measured shard -# times behind the job timeout; keep the two matrices identical. +# "api" itself as two hash partitions (api-1 / api-2), so the wall-clock time is the slowest shard, +# not the whole suite. build.yml carries the description of what each shard selects and the +# measured shard times behind the job timeout; keep the two matrices identical. on: schedule: - cron: '0 2 * * *' @@ -24,8 +24,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 + groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -69,7 +73,7 @@ jobs: run: ttyd --version - name: Integration tests - run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. @@ -183,8 +187,12 @@ jobs: fail-fast: false matrix: shard: - - name: api + - name: api-1 + groups: "!ui & !slow & !upgrade" + partition: "1/2" + - name: api-2 groups: "!ui & !slow & !upgrade" + partition: "2/2" - name: ui groups: "ui & !slow & !sample & !camel" - name: samples @@ -268,7 +276,7 @@ jobs: echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)" exit 1 fi - mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' + mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}' # Guards #7215: a tag expression that discovers nothing leaves the reports directory absent # and the build green, so what ran is asserted by count instead of by the job's colour. diff --git a/pom.xml b/pom.xml index 798cb735fef..0d8362f73e0 100644 --- a/pom.xml +++ b/pom.xml @@ -290,6 +290,10 @@ plus the few explicitly-smoke UI journeys. --> + + diff --git a/tests/tests-integrations/pom.xml b/tests/tests-integrations/pom.xml index 6168de9a102..d7e372e37be 100644 --- a/tests/tests-integrations/pom.xml +++ b/tests/tests-integrations/pom.xml @@ -30,6 +30,12 @@ junit-platform-suite compile + + + org.junit.platform + junit-platform-launcher + compile + org.eclipse.dirigible dirigible-application @@ -138,6 +144,11 @@ ${it.groups} ${it.excludedGroups} + + + ${it.partition} +