Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude/docs/ci.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

- `code-style`: `mvn -T 1C formatter:validate`
- `tests` (ubuntu + windows matrix): `mvn clean install -P unit-tests`
- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow`), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below).
- `integration-tests-h2` / `-postgresql`: the **full** Selenide IT suite, `mvn clean install -P integration-tests` with the matching `DIRIGIBLE_DATASOURCE_DEFAULT_*` env vars (MSSQL is no longer a CI leg — removed in #6150). Each DB leg is **sharded into four parallel matrix jobs** selected by tag expression — `api` (`!ui & !slow & !upgrade`, itself run as two hash partitions `api-1` / `api-2`, see the budget note below), `ui` (`ui & !slow & !sample & !camel`), `samples` (`ui & !slow & (sample | camel)`), `slow` (`slow`) — so the run's wall clock is the slowest shard, not the whole ~2h30 suite. The shards partition the suite; keep them disjoint and complete when adding tags. **Measured on nightly run 34428729298 (h2 / postgresql): `api` 57 / 62 min, `slow` 42, `ui` 36 / 36, `samples` 13 / 13**, against a `timeout-minutes: 90` per job (see the budget note below).
- `build-deploy`: `mvn clean install -P quick-build` then Docker buildx multi-arch image push to `dirigiblelabs/dirigible`

### PR gate vs full suite (smoke / nightly split)
Expand All @@ -17,7 +17,7 @@ The full Selenide UI suite takes ~1.5h per DB, so it does **not** run on every P
**The job budget, and how a shard that runs out of it reads (#7283).** `timeout-minutes` is **90** on every IT job (150 on the PR `smoke-tests` job) - comfortably above the measured times above, because a job cancelled at its cap reads as infrastructure rather than as a test result. At 60 the PostgreSQL `api` shard of the nightly was killed **two classes short of finishing** and that looked exactly like a hang. Two things to know before concluding "wedged IT" from a cancelled shard:

- **Check the last `[INFO] Running org.eclipse...` line against the cancellation timestamp.** In the run that raised the issue the JVM was healthy: it kept starting one IT class per ~40 s until 10 minutes before the cancel, and the class suspected of hanging (`EdmModelRoundTripIT`) had passed in 50.51 s. Searching the downloaded log for a class name lands on its *last* mention, which is its own completion line - everything after it belongs to the next classes.
- **`api` cannot be rebalanced with `@Tag("slow")`.** Its 83 classes cost a near-uniform ~40 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows ~40 s per IT class added and has no long pole to move out. Splitting it needs a partition mechanism, not a tag - and a name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode.
- **`api` cannot be rebalanced with `@Tag("slow")`.** Its classes cost a near-uniform ~40-60 s each, nearly all of it the Spring context each class boots (`@DirtiesContext` `AFTER_EACH_TEST_METHOD`), so the shard grows by that much per IT class added and has no long pole to move out. A name-pattern split (`-Dit.test`) would risk silently dropping classes that match neither half, which is the #7215 failure mode. So it is split by a **partition**: by October 2026 it had grown to ~133 classes / ~83 min and was cancelled at the 90-minute cap on every master and nightly run, and it now runs as `api-1` / `api-2`, each `-Dit.groups="!ui & !slow & !upgrade" -Dit.partition=k/2`. `ClassPartitionFilter` (tests-integrations `support/`, a JUnit `PostDiscoveryFilter` registered through `META-INF/services`, applied after the tag filter) keeps the classes whose **top-level** class name hashes into partition `k` of `n` - disjoint and complete by construction, nested classes stay with their enclosing class, nothing to tag. Without `-Dit.partition` it keeps everything. When a partition nears the cap again, go to `/3` rather than reshuffling tags; `-Djunit.platform.execution.dryRun.enabled=true` counts what each partition selects in about a minute.

**A hang costs one class, not the shard.** `IntegrationTest` carries a class-level **`@Timeout(15, MINUTES)`** (inherited; per test *and* per lifecycle method - the Spring context start is an extension callback and is not counted), so a method that never returns fails its own class and the run continues and reports. The slowest measured method in the suite is 456 s (`CreateNewFileIT`, one `@Test`), so 15 minutes is roughly twice the worst legitimate case; a class needing a different bound declares its own `@Timeout`, which overrides the inherited one (`JavaLspIT`, `JavaDebugIT`, `SchemaTemplateForeignKeyIT` each declare a tighter one). There is deliberately **no `forkedProcessTimeoutInSeconds`**: it bounds the whole fork, and the jobs sharing this configuration range from 13 to 91 minutes, so one value would either be useless or kill the longest job.

Expand Down
33 changes: 24 additions & 9 deletions .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,8 @@ jobs:

# The full IT suite is split ("sharded") into four disjoint slices selected by JUnit 5 tags, so
# the wall-clock time of a run is the slowest shard instead of the whole suite (~2h30 serial):
# api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles
# api - every HTTP-level IT (untagged, i.e. everything not tagged "ui") bar the long poles,
# run as two jobs (api-1, api-2) - see the partition note below
# ui - the browser-driven ITs except the sample-project, camel and long-pole families
# samples - the sample-project clone/publish ITs ("sample") + the camel journeys still driven
# through the browser ("camel"), which is now only the two template starters - the
Expand All @@ -181,24 +182,34 @@ jobs:
# samples 13 / 13. "api" is the critical path and the "slow" tag cannot rebalance it: its 83
# classes cost a near-uniform ~40 s each, almost all of it the Spring context each class boots
# (@DirtiesContext AFTER_EACH_TEST_METHOD), so the shard grows by ~40 s per IT class added and
# has no long pole to move out. #7283 raised the budget rather than split it; splitting "api"
# needs a partition mechanism, not a tag.
# has no long pole to move out. #7283 raised the budget rather than split it, and by October 2026
# the shard (~133 classes, ~83 min of test time) was cancelled at the 90-minute cap on every run.
# So "api" is split by a partition, not by a tag: -Dit.partition=k/n keeps the classes whose
# top-level class name hashes into partition k of n (ClassPartitionFilter in tests-integrations, a
# JUnit PostDiscoveryFilter applied after the tag filter). The partitions are disjoint and complete
# by construction - a new IT lands in exactly one of them without being tagged - and, the class
# costs being near-uniform, about even (on the 2026-10-05 h2 shard log: 64 classes / 38.5 min
# against 69 / 44.7). Add a third partition (1/3, 2/3, 3/3) when one of them nears the cap again.
integration-tests-h2:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard:
- name: api
- name: api-1
groups: "!ui & !slow & !upgrade"
partition: "1/2"
- name: api-2
groups: "!ui & !slow & !upgrade"
partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
groups: "ui & !slow & (sample | camel)"
- name: slow
groups: "slow"
# The slowest green shard ("api") takes ~57 min on h2 and ~62 on PostgreSQL; a heap-exhausted
# JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs
# The slowest green shard ("api", before its split) took ~57 min on h2 and ~62 on PostgreSQL; a
# heap-exhausted JVM used to thrash for 4h+ before anyone noticed. Cap the job above the healthy time so hangs
# fail fast instead of burning runners - but far enough above it that a routine run is never
# cancelled: at 60 the PostgreSQL "api" shard was killed two classes short of finishing, which
# reads as infrastructure rather than as a test result (#7283). A single method can no longer
Expand Down Expand Up @@ -240,7 +251,7 @@ jobs:
run: ttyd --version

- name: Integration tests
run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'

# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
Expand Down Expand Up @@ -292,8 +303,12 @@ jobs:
fail-fast: false
matrix:
shard:
- name: api
- name: api-1
groups: "!ui & !slow & !upgrade"
partition: "1/2"
- name: api-2
groups: "!ui & !slow & !upgrade"
partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
Expand Down Expand Up @@ -377,7 +392,7 @@ jobs:
echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)"
exit 1
fi
mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'

# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
Expand Down
22 changes: 15 additions & 7 deletions .github/workflows/nightly.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,9 +5,9 @@ name: Nightly Integration Tests
# PRs run the fast "smoke-tests" job (see pull-request.yml). Runs on a nightly schedule and on
# demand; the same full suite also runs on every push to master (see build.yml). Like build.yml, the
# suite is sharded into four tag-selected slices (api / ui / samples / slow) that run in parallel,
# so the wall-clock time is the slowest shard ("api", ~57 min on h2 and ~62 on PostgreSQL), not the
# whole suite. build.yml carries the description of what each shard selects and the measured shard
# times behind the job timeout; keep the two matrices identical.
# "api" itself as two hash partitions (api-1 / api-2), so the wall-clock time is the slowest shard,
# not the whole suite. build.yml carries the description of what each shard selects and the
# measured shard times behind the job timeout; keep the two matrices identical.
on:
schedule:
- cron: '0 2 * * *'
Expand All @@ -24,8 +24,12 @@ jobs:
fail-fast: false
matrix:
shard:
- name: api
- name: api-1
groups: "!ui & !slow & !upgrade"
partition: "1/2"
- name: api-2
groups: "!ui & !slow & !upgrade"
partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
Expand Down Expand Up @@ -69,7 +73,7 @@ jobs:
run: ttyd --version

- name: Integration tests
run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
run: mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'

# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
Expand Down Expand Up @@ -183,8 +187,12 @@ jobs:
fail-fast: false
matrix:
shard:
- name: api
- name: api-1
groups: "!ui & !slow & !upgrade"
partition: "1/2"
- name: api-2
groups: "!ui & !slow & !upgrade"
partition: "2/2"
- name: ui
groups: "ui & !slow & !sample & !camel"
- name: samples
Expand Down Expand Up @@ -268,7 +276,7 @@ jobs:
echo "::error::The datasource configuration did not reach this step - the suite would run on H2 (see #7249)"
exit 1
fi
mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}'
mvn clean install -P integration-tests -Dit.groups='${{ matrix.shard.groups }}' -Dit.partition='${{ matrix.shard.partition }}'

# Guards #7215: a tag expression that discovers nothing leaves the reports directory absent
# and the build green, so what ran is asserted by count instead of by the job's colour.
Expand Down
4 changes: 4 additions & 0 deletions pom.xml
Original file line number Diff line number Diff line change
Expand Up @@ -290,6 +290,10 @@
plus the few explicitly-smoke UI journeys. -->
<it.groups></it.groups>
<it.excludedGroups></it.excludedGroups>
<!-- Integration-test partition, "k/n" (e.g. 1/2): keeps one of n hash partitions of the
classes the tag expression selected (ClassPartitionFilter in tests-integrations). Empty by
default - every selected class runs. CI splits the "api" shard with it. -->
<it.partition></it.partition>
</properties>

<build>
Expand Down
11 changes: 11 additions & 0 deletions tests/tests-integrations/pom.xml
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,12 @@
<artifactId>junit-platform-suite</artifactId>
<scope>compile</scope>
</dependency>
<!-- PostDiscoveryFilter API for ClassPartitionFilter (-Dit.partition) -->
<dependency>
<groupId>org.junit.platform</groupId>
<artifactId>junit-platform-launcher</artifactId>
<scope>compile</scope>
</dependency>
<dependency>
<groupId>org.eclipse.dirigible</groupId>
<artifactId>dirigible-application</artifactId>
Expand Down Expand Up @@ -138,6 +144,11 @@
<configuration>
<groups>${it.groups}</groups>
<excludedGroups>${it.excludedGroups}</excludedGroups>
<!-- -Dit.partition=k/n keeps one hash partition of the selected classes
(ClassPartitionFilter, registered as a JUnit PostDiscoveryFilter). -->
<systemPropertyVariables>
<dirigible.it.partition>${it.partition}</dirigible.it.partition>
</systemPropertyVariables>
<!-- Fresh JVM per test class. Every IT boots a full Dirigible context per test
method (@DirtiesContext AFTER_EACH_TEST_METHOD) and closed contexts do not
release all their heap (schedulers, engine pools, statics), so a single
Expand Down
Loading
Loading