migrations: a versioned .migration artefact applies once per tenant schema after table sync, recorded in a DIRIGIBLE_MIGRATIONS ledger (#7636) - #7672
Conversation
…chema after table sync, recorded in a DIRIGIBLE_MIGRATIONS ledger (eclipse-dirigible#7636) Structure evolves on publish and CSVIM seeds reference data, but there was no artefact for a data migration - a backfill, a split field, a recomputed roll-up - so such changes were hand-run SQL per instance and per tenant schema, or a job inventing its own "already ran" marker. A new components/data/data-migrations module adds the .migration artefact: <project>/.../<version>__<description>.migration, SQL, optionally headed by "-- tenant: each|system" (default each) and "-- idempotent: true|false" (default false). MigrationsSynchronizer runs at SynchronizersOrder.MIGRATION (260): after schema/table/view/entity, before BPMN and CSVIM. - Each migration runs in one transaction together with the row recording it in DIRIGIBLE_MIGRATIONS (project, version, location, checksum, tenant, applied at, duration). The ledger lives in the database the migration changes - every tenant schema, or the system database for tenant: system - and is keyed by project/version, so a concurrent second node's insert fails and rolls its run back. - The ledger decides, not the artefact lifecycle, so every re-run is safe: the post-provisioning re-trigger migrates the new tenant and finds the others done; a fresh system database re-parses without re-applying. - An applied file whose checksum changed is a FAILED artefact (not a re-run) unless it declares idempotent: true. Checksums ignore line-ending conversion. - A project's migrations apply in version order (numeric, segment-wise): a later version waits for an earlier one still applying in the pass, and fails if the earlier one failed. - The synchronizer iterates the tenants itself and registers one state for all of them; the per-tenant completion of BaseSynchronizer keeps the state the LAST tenant registered, so one tenant's success would hide another's failure. - A failed migration is a failed artefact, so the census already counts it and DIRIGIBLE_READINESS_REQUIRE_CLEAN_BOOT withholds readiness. The new "migrations" health component reports total / pending / failed, and GET /services/core/migrations lists the caller's tenant ledger (plus the system ledger for the default tenant). Verified: data-migrations unit tests (22, H2-backed executor tests), the core-base and core-liquibase suites, DataMigrationIT on H2 and on PostgreSQL 16, formatter:validate with the cache wiped, and the release javadoc build on the touched modules. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| .column(quoted(COLUMN_APPLIED_AT), DataType.TIMESTAMP, false, false, false) | ||
| .column(quoted(COLUMN_DURATION), DataType.BIGINT, false, false, false) | ||
| .build(); | ||
| try (PreparedStatement statement = connection.prepareStatement(sql)) { |
| } | ||
|
|
||
| @GetMapping | ||
| List<Entry> list() { |
…dirigible#7636) The CREATE TABLE is built from constants, but CodeQL reads it as derived from user input (every datasource leaves DataSourcesManager through one static map) and flagged the log line as log injection. Logging the table name says what happened and carries nothing tainted. Verified: data-migrations unit tests (22) green; formatter:validate on the module with the cache wiped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
CodeQL: the log-injection alert ( The two high alerts are false positives. I'd ask a maintainer with Security-tab access to dismiss them as such:
#6384 (number series) raised the same three alerts in the same places when it added a GET endpoint over a per-tenant store that creates its table, and it was merged with them. #6854 added the same Separately, the earlier |
Cause
Structure evolves on publish (
TableAlterProcessor, FK creation, unique reconciliation) and CSVIM seeds reference data, but there was no artefact for the third kind of change: a data migration (backfill a new column, split a field, recompute a roll-up). Such changes were hand-run SQL against every instance and tenant schema, or a.jobinventing its own "did I already run" marker, with no record of whether they ever ran.Change
A new module,
components/data/data-migrations(registered ingroup-database), adds the.migrationartefact:MigrationsSynchronizerruns at the newSynchronizersOrder.MIGRATION = 260. That is after SCHEMA/TABLE/VIEW/ENTITY (210 to 240), so the columns a backfill needs exist, and before BPMN (300) and CSVIM (400).DIRIGIBLE_MIGRATIONS(project, version, location, checksum, tenant, applied_at, duration). The ledger lives in the database the migration changes: each tenant schema, or the system DB fortenant: system. It is created on first use, like the event outbox. The primary key isproject/version, so a second node applying the same migration at the same time fails its insert and rolls its own run back.idempotent: true. The checksum ignores line-ending conversion. Restoring the file heals it.2 < 10. A later version waits (returns undepleted, with no state registered) while an earlier one is still applying in the same pass, and fails if the earlier one failed.BaseSynchronizer's per-tenant completion keeps whatever the last tenant registered, so one tenant's success would hide another's failure. It still reportsmultitenantExecution(), so the post-provisioning re-trigger includes it.artefactscensus therefore counts it, andDIRIGIBLE_READINESS_REQUIRE_CLEAN_BOOTwithholds readiness throughisCleanBoot(), soPlatformReadinessneeded no change. Pending migrations hold readiness through the boot latch. A newmigrationshealth component reportstotal/pending/failed.GET /services/core/migrations(ADMINISTRATOR, OPERATOR) returns the caller's tenant ledger. A default-tenant caller also gets the system ledger; other tenants' administrators don't.DIRIGIBLE_MIGRATION_FILES, with Liquibase changesets for the table and itsARTEFACT_KEYunique constraint.Verified
data-migrationsunit tests, 22:core-base(132) andcore-liquibase(2) suites green, the latter running the changelog with the new changesets.DataMigrationIT(HTTP/JDBC, no browser) is green on H2 (28.8 s) and on PostgreSQL 16 (31.1 s). The PostgreSQL run used the CI legs' env-var shape, andpg_stat_databaseshowed writes to both databases. What it checks:migrations.failed = 1andartefacts.failedByType.migration = 1, andPlatformReadiness.isCleanBoot()is false;mvn -T 1C formatter:validatewith the formatter cache wiped: BUILD SUCCESS.core-baseanddata-migrations: BUILD SUCCESS, javadoc jar produced.Not verified / known limits
PlatformReadiness.isCleanBoot()and the census, which are what the bridge reads.JavaSynchronizer.finishing(), at the end of the pass and after every phase. A migration targeting such a table fails its first pass and heals on the FAILED retry (DIRIGIBLE_SYNCHRONIZER_FAILED_RETRY_INTERVAL_SECONDS, 30 s) once the generation is installed. Tables from.table/.schemaare in place before migrations run.ScriptUtils(;, with comments and quoted literals handled). PostgreSQL dollar-quoted bodies are not supported.migrationshealth component reads the artefact table on each probe.DataMigrationITwas run locally, not the full IT suite. The help portal (separate repository) is not updated yet.Fixes #7636
🤖 Generated with Claude Code