diff --git a/.claude/docs/ci.md b/.claude/docs/ci.md
index e22a60c3eba..6e67dc20173 100644
--- a/.claude/docs/ci.md
+++ b/.claude/docs/ci.md
@@ -4,7 +4,7 @@
- `code-style`: `mvn -T 1C formatter:validate`
- `tests` (ubuntu + windows matrix): `mvn clean install -P unit-tests`
-- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow`), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below).
+- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow & !upgrade`, itself run as two hash partitions `api-1` / `api-2`, see the budget note below), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below).
- `build-deploy`: `mvn clean install -P quick-build` then Docker buildx multi-arch image push to `dirigiblelabs/dirigible`
### PR gate vs full suite (smoke / nightly split)
@@ -17,7 +17,7 @@ The full Selenide UI suite takes ~1.5h per DB, so it does **not** run on every P
**The job budget, and how a shard that runs out of it reads (#7283).** `timeout-minutes` is **90** on every IT job (150 on the PR `smoke-tests` job) - comfortably above the measured times above, because a job cancelled at its cap reads as infrastructure rather than as a test result. At 60 the PostgreSQL `api` shard of the nightly was killed **two classes short of finishing** and that looked exactly like a hang. Two things to know before concluding "wedged IT" from a cancelled shard:
- **Check the last `[INFO] Running org.eclipse...` line against the cancellation timestamp.** In the run that raised the issue the JVM was healthy: it kept starting one IT class per ~40 s until 10 minutes before the cancel, and the class suspected of hanging (`EdmModelRoundTripIT`) had passed in 50.51 s. Searching the downloaded log for a class name lands on its *last* mention, which is its own completion line - everything after it belongs to the next classes.
-- **`api` cannot be rebalanced with `@Tag("slow")`.** Its 83 classes cost a near-uniform ~40 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows ~40 s per IT class added and has no long pole to move out. Splitting it needs a partition mechanism, not a tag - and a name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode.
+- **`api` cannot be rebalanced with `@Tag("slow")`.** Its classes cost a near-uniform ~40-60 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows by that much per IT class added and has no long pole to move out. A name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. So it is split by a **partition**: by October 2026 it had grown to ~133 classes / ~83 min and was cancelled at the 90-minute cap on every master and nightly run, and it now runs as `api-1` / `api-2`, each `-Dit.groups="!ui & !slow & !upgrade" -Dit.partition=k/2`. `ClassPartitionFilter` (tests-integrations `support/`, a JUnit `PostDiscoveryFilter` registered through `META-INF/services`, applied after the tag filter) keeps the classes whose **top-level** class name hashes into partition `k` of `n` - disjoint and complete by construction, nested classes stay with their enclosing class, nothing to tag. Without `-Dit.partition` it keeps everything. When a partition nears the cap again, go to `/3` rather than reshuffling tags; `-Djunit.platform.execution.dryRun.enabled=true` counts what each partition selects in about a minute.
**A hang costs one class, not the shard.** `IntegrationTest` carries a class-level **`@Timeout(15, MINUTES)`** (inherited; per test *and* per lifecycle method - the Spring context start is an extension callback and is not counted), so a method that never returns fails its own class and the run continues and reports. The slowest measured method in the suite is 456 s (`CreateNewFileIT`, one `@Test`), so 15 minutes is roughly twice the worst legitimate case; a class needing a different bound declares its own `@Timeout`, which overrides the inherited one (`JavaLspIT`, `JavaDebugIT`, `SchemaTemplateForeignKeyIT` each declare a tighter one). There is deliberately **no `forkedProcessTimeoutInSeconds`**: it bounds the whole fork, and the jobs sharing this configuration range from 13 to 91 minutes, so one value would either be useless or kill the longest job.
diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml
index 23619a66441..4fb3c972c7e 100644
--- a/.github/workflows/build.yml
+++ b/.github/workflows/build.yml
@@ -162,7 +162,8 @@ jobs:
# The full IT suite is split ("sharded") into four disjoint slices selected by JUnit 5 tags, so
# the wall-clock time of a run is the slowest shard instead of the whole suite (~2h30 serial):
- # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles
+ # api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles,
+ # run as two jobs (api-1, api-2) - see the partition note below
# ui - the browser-driven ITs except the sample-project, camel and long-pole families
# samples - the sample-project clone/publish ITs ("sample") + the camel journeys still driven
# through the browser ("camel"), which is now only the two template starters - the
@@ -181,24 +182,34 @@ jobs:
# samples 13 / 13. "api" is the critical path and the "slow" tag cannot rebalance it: its 83
# classes cost a near-uniform ~40 s each, almost all of it the Spring context each class boots
# (@DirtiesContext AFTER_EACH_TEST_METHOD), so the shard grows by ~40 s per IT class added and
- # has no long pole to move out. #7283 raised the budget rather than split it; splitting "api"
- # needs a partition mechanism, not a tag.
+ # has no long pole to move out. #7283 raised the budget rather than split it, and by October 2026
+ # the shard (~133 classes, ~83 min of test time) was cancelled at the 90-minute cap on every run.
+ # So "api" is split by a partition, not by a tag: -Dit.partition=k/n keeps the classes whose
+ # top-level class name hashes into partition k of n (ClassPartitionFilter in tests-integrations, a
+ # JUnit PostDiscoveryFilter applied after the tag filter). The partitions are disjoint and complete
+ # by construction - a new IT lands in exactly one of them without being tagged - and, the class
+ # costs being near-uniform, about even (on the 2026-10-05 h2 shard log: 64 classes / 38.5 min
+ # against 69 / 44.7). Add a third partition (1/3, 2/3, 3/3) when one of them nears the cap again.
integration-tests-h2:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard:
- - name: api
+ - name: api-1
groups: "!ui & !slow & !upgrade"
+ partition: "1/2"
+ - name: api-2
+ groups: "!ui & !slow & !upgrade"
+ partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
groups: "ui & !slow & (sample | camel)"
- name: slow
groups: "slow"
- # The slowest green shard ("api") takes ~57 min on h2 and ~62 on PostgreSQL; a heap-exhausted
- # JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs
+ # The slowest green shard ("api", before its split) took ~57 min on h2 and ~62 on PostgreSQL; a
+ # heap-exhausted JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs
# fail fast instead of burning runners - but far enough above it that a routine run is never
# cancelled: at 60 the PostgreSQL "api" shard was killed two classes short of finishing, which
# reads as infrastructure rather than as a test result (#7283). A single method can no longer
@@ -240,7 +251,7 @@ jobs:
run: ttyd --version
- name: Integration tests
- run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
+ run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'
# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
@@ -292,8 +303,12 @@ jobs:
fail-fast: false
matrix:
shard:
- - name: api
+ - name: api-1
+ groups: "!ui & !slow & !upgrade"
+ partition: "1/2"
+ - name: api-2
groups: "!ui & !slow & !upgrade"
+ partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
@@ -377,7 +392,7 @@ jobs:
echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)"
exit 1
fi
- mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
+ mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'
# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
diff --git a/.github/workflows/nightly.yml b/.github/workflows/nightly.yml
index e0008c76854..d22bc7e237f 100644
--- a/.github/workflows/nightly.yml
+++ b/.github/workflows/nightly.yml
@@ -5,9 +5,9 @@ name: Nightly Integration Tests
# PRs run the fast "smoke-tests" job (see pull-request.yml). Runs on a nightly schedule and on
# demand; the same full suite also runs on every push to master (see build.yml). Like build.yml, the
# suite is sharded into four tag-selected slices (api / ui / samples / slow) that run in parallel,
-# so the wall-clock time is the slowest shard ("api", ~57 min on h2 and ~62 on PostgreSQL), not the
-# whole suite. build.yml carries the description of what each shard selects and the measured shard
-# times behind the job timeout; keep the two matrices identical.
+# "api" itself as two hash partitions (api-1 / api-2), so the wall-clock time is the slowest shard,
+# not the whole suite. build.yml carries the description of what each shard selects and the
+# measured shard times behind the job timeout; keep the two matrices identical.
on:
schedule:
- cron: '0 2 * * *'
@@ -24,8 +24,12 @@ jobs:
fail-fast: false
matrix:
shard:
- - name: api
+ - name: api-1
groups: "!ui & !slow & !upgrade"
+ partition: "1/2"
+ - name: api-2
+ groups: "!ui & !slow & !upgrade"
+ partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
@@ -69,7 +73,7 @@ jobs:
run: ttyd --version
- name: Integration tests
- run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
+ run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'
# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
@@ -183,8 +187,12 @@ jobs:
fail-fast: false
matrix:
shard:
- - name: api
+ - name: api-1
+ groups: "!ui & !slow & !upgrade"
+ partition: "1/2"
+ - name: api-2
groups: "!ui & !slow & !upgrade"
+ partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
@@ -268,7 +276,7 @@ jobs:
echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)"
exit 1
fi
- mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
+ mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'
# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
diff --git a/pom.xml b/pom.xml
index 798cb735fef..0d8362f73e0 100644
--- a/pom.xml
+++ b/pom.xml
@@ -290,6 +290,10 @@
plus the few explicitly-smoke UI journeys. -->
+
+
diff --git a/tests/tests-integrations/pom.xml b/tests/tests-integrations/pom.xml
index 6168de9a102..d7e372e37be 100644
--- a/tests/tests-integrations/pom.xml
+++ b/tests/tests-integrations/pom.xml
@@ -30,6 +30,12 @@
junit-platform-suite
compile
+
+
+ org.junit.platform
+ junit-platform-launcher
+ compile
+
org.eclipse.dirigible
dirigible-application
@@ -138,6 +144,11 @@
${it.groups}
${it.excludedGroups}
+
+
+ ${it.partition}
+