Skip to content

migrations: a versioned .migration artefact applies once per tenant schema after table sync, recorded in a DIRIGIBLE_MIGRATIONS ledger (#7636) - #7672

Open
NicoleNG18 wants to merge 2 commits into
eclipse-dirigible:masterfrom
NicoleNG18:issue-7636-data-migrations
Open

NicoleNG18 wants to merge 2 commits into
eclipse-dirigible:masterfrom
NicoleNG18:issue-7636-data-migrations

Conversation

@NicoleNG18

Copy link
Copy Markdown
Contributor

Cause

Structure evolves on publish (TableAlterProcessor, FK creation, unique reconciliation) and CSVIM seeds reference data, but there was no artefact for the third kind of change: a data migration (backfill a new column, split a field, recompute a roll-up). Such changes were hand-run SQL against every instance and tenant schema, or a .job inventing its own "did I already run" marker, with no record of whether they ever ran.

Change

A new module, components/data/data-migrations (registered in group-database), adds the .migration artefact:

<project>/.../<version>__<description>.migration      e.g. migrations/V2__backfill_status.migration

-- tenant: each          (default: once per tenant schema) | system (once, on the system database)
-- idempotent: false     (default: an edit after it applied FAILS it) | true (an edit re-applies it)
UPDATE "ORDERS" SET "STATUS" = 'OPEN' WHERE "STATUS" IS NULL;
  • Ordering: MigrationsSynchronizer runs at the new SynchronizersOrder.MIGRATION = 260. That is after SCHEMA/TABLE/VIEW/ENTITY (210 to 240), so the columns a backfill needs exist, and before BPMN (300) and CSVIM (400).
  • Exactly once, per database: each migration runs in one transaction together with its ledger row in DIRIGIBLE_MIGRATIONS (project, version, location, checksum, tenant, applied_at, duration). The ledger lives in the database the migration changes: each tenant schema, or the system DB for tenant: system. It is created on first use, like the event outbox. The primary key is project/version, so a second node applying the same migration at the same time fails its insert and rolls its own run back.
  • Re-runs are safe because the ledger decides, not the artefact lifecycle. When a tenant is provisioned, the platform re-synchronizes every multitenant artefact. That run migrates the new tenant and finds the existing tenants done. A boot against a fresh system database re-parses without re-applying anything.
  • Edited after apply: a changed checksum is a FAILED artefact with a message naming the tenant, the applied date and both checksums. It is not re-run, unless the file declares idempotent: true. The checksum ignores line-ending conversion. Restoring the file heals it.
  • Version order within a project: versions compare numerically by segment, so 2 < 10. A later version waits (returns undepleted, with no state registered) while an earlier one is still applying in the same pass, and fails if the earlier one failed.
  • One state across tenants: the synchronizer iterates the tenants itself. BaseSynchronizer's per-tenant completion keeps whatever the last tenant registered, so one tenant's success would hide another's failure. It still reports multitenantExecution(), so the post-provisioning re-trigger includes it.
  • Readiness and health: a failed migration is a failed artefact. The artefacts census therefore counts it, and DIRIGIBLE_READINESS_REQUIRE_CLEAN_BOOT withholds readiness through isCleanBoot(), so PlatformReadiness needed no change. Pending migrations hold readiness through the boot latch. A new migrations health component reports total / pending / failed.
  • Ledger endpoint: GET /services/core/migrations (ADMINISTRATOR, OPERATOR) returns the caller's tenant ledger. A default-tenant caller also gets the system ledger; other tenants' administrators don't.
  • Platform tables: the artefact table is DIRIGIBLE_MIGRATION_FILES, with Liquibase changesets for the table and its ARTEFACT_KEY unique constraint.
  • DELETE and cleanup remove the artefact only. Applied data changes and ledger rows are history.

Verified

  • data-migrations unit tests, 22:
    • Parser: file names, headers, checksum normalisation, version ordering.
    • Executor, on real in-memory H2: apply and record once; already applied; edit fails vs. idempotent re-apply; a failing second statement rolls back the first and records nothing; waiting for, and failing on, an earlier version.
    • Synchronizer aggregation: one tenant's failure is not overwritten by another's success; a waiting migration keeps its lifecycle; system scope runs once on the system DB; predecessors are filtered by scope.
  • core-base (132) and core-liquibase (2) suites green, the latter running the changelog with the new changesets.
  • DataMigrationIT (HTTP/JDBC, no browser) is green on H2 (28.8 s) and on PostgreSQL 16 (31.1 s). The PostgreSQL run used the CI legs' env-var shape, and pg_stat_database showed writes to both databases. What it checks:
    • a table plus two migrations in the default tenant and a provisioned tenant, the backfill applying after the seed;
    • a reset row is not rewritten after a second tenant is provisioned (the full multitenant re-sync), and the new tenant is migrated;
    • one ledger row per migration per tenant;
    • the endpoint lists the rows;
    • an edited applied file shows migrations.failed = 1 and artefacts.failedByType.migration = 1, and PlatformReadiness.isCleanBoot() is false;
    • restoring the file heals it.
  • mvn -T 1C formatter:validate with the formatter cache wiped: BUILD SUCCESS.
  • Release-profile javadoc build on core-base and data-migrations: BUILD SUCCESS, javadoc jar produced.

Not verified / known limits

  • Two real boots were not run. The "boot twice" proof is the re-synchronization that provisioning a tenant triggers, inside one Spring context. The ledger logic is the same either way.
  • The readiness bridge was not booted with the flag on. The IT asserts PlatformReadiness.isCleanBoot() and the census, which are what the bridge reads.
  • Tables from brand-new client-Java entities: these are created in JavaSynchronizer.finishing(), at the end of the pass and after every phase. A migration targeting such a table fails its first pass and heals on the FAILED retry (DIRIGIBLE_SYNCHRONIZER_FAILED_RETRY_INTERVAL_SECONDS, 30 s) once the generation is installed. Tables from .table / .schema are in place before migrations run.
  • SQL only. The JS/Java variants mentioned in the issue are not included. Statements are split by Spring's ScriptUtils (;, with comments and quoted literals handled). PostgreSQL dollar-quoted bodies are not supported.
  • The migrations health component reads the artefact table on each probe.
  • Only DataMigrationIT was run locally, not the full IT suite. The help portal (separate repository) is not updated yet.

Fixes #7636

🤖 Generated with Claude Code

…chema after table sync, recorded in a DIRIGIBLE_MIGRATIONS ledger (eclipse-dirigible#7636)

Structure evolves on publish and CSVIM seeds reference data, but there was
no artefact for a data migration - a backfill, a split field, a recomputed
roll-up - so such changes were hand-run SQL per instance and per tenant
schema, or a job inventing its own "already ran" marker.

A new components/data/data-migrations module adds the .migration artefact:
<project>/.../<version>__<description>.migration, SQL, optionally headed by
"-- tenant: each|system" (default each) and "-- idempotent: true|false"
(default false). MigrationsSynchronizer runs at SynchronizersOrder.MIGRATION
(260): after schema/table/view/entity, before BPMN and CSVIM.

- Each migration runs in one transaction together with the row recording it
  in DIRIGIBLE_MIGRATIONS (project, version, location, checksum, tenant,
  applied at, duration). The ledger lives in the database the migration
  changes - every tenant schema, or the system database for tenant: system -
  and is keyed by project/version, so a concurrent second node's insert
  fails and rolls its run back.
- The ledger decides, not the artefact lifecycle, so every re-run is safe:
  the post-provisioning re-trigger migrates the new tenant and finds the
  others done; a fresh system database re-parses without re-applying.
- An applied file whose checksum changed is a FAILED artefact (not a
  re-run) unless it declares idempotent: true. Checksums ignore line-ending
  conversion.
- A project's migrations apply in version order (numeric, segment-wise): a
  later version waits for an earlier one still applying in the pass, and
  fails if the earlier one failed.
- The synchronizer iterates the tenants itself and registers one state for
  all of them; the per-tenant completion of BaseSynchronizer keeps the state
  the LAST tenant registered, so one tenant's success would hide another's
  failure.
- A failed migration is a failed artefact, so the census already counts it
  and DIRIGIBLE_READINESS_REQUIRE_CLEAN_BOOT withholds readiness. The new
  "migrations" health component reports total / pending / failed, and
  GET /services/core/migrations lists the caller's tenant ledger (plus the
  system ledger for the default tenant).

Verified: data-migrations unit tests (22, H2-backed executor tests), the
core-base and core-liquibase suites, DataMigrationIT on H2 and on
PostgreSQL 16, formatter:validate with the cache wiped, and the release
javadoc build on the touched modules.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
.column(quoted(COLUMN_APPLIED_AT), DataType.TIMESTAMP, false, false, false)
.column(quoted(COLUMN_DURATION), DataType.BIGINT, false, false, false)
.build();
try (PreparedStatement statement = connection.prepareStatement(sql)) {
}

@GetMapping
List<Entry> list() {
…dirigible#7636)

The CREATE TABLE is built from constants, but CodeQL reads it as derived
from user input (every datasource leaves DataSourcesManager through one
static map) and flagged the log line as log injection. Logging the table
name says what happened and carries nothing tainted.

Verified: data-migrations unit tests (22) green; formatter:validate on
the module with the cache wiped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@NicoleNG18

Copy link
Copy Markdown
Contributor Author

CodeQL: the log-injection alert (MigrationLedger.java, medium) is fixed in f966d2f. The ledger creation now logs the table name instead of the generated SQL.

The two high alerts are false positives. I'd ask a maintainer with Security-tab access to dismiss them as such:

  • Query built from user-controlled sources (MigrationLedger.java, the CREATE TABLE). The statement is built only from constants: the table name, the column names and sizes. The one non-constant input is the dialect, which SqlFactory.getNative(connection) reads from the connection. CodeQL treats that connection as user-controlled because every datasource leaves DataSourcesManager through one static map (DataSourceInitializer.DATASOURCES), and some datasource reaches that map from a request.
  • HTTP request type unprotected from CSRF (MigrationsEndpoint.list()). The handler only reads the ledger. The "state-changing action" CodeQL found is the lazy datasource initialization inside getDefaultDataSource(), which every GET touching the default datasource reaches.

#6384 (number series) raised the same three alerts in the same places when it added a GET endpoint over a per-tenant store that creates its table, and it was merged with them. #6854 added the same CREATE TABLE code without an HTTP endpoint and got none. That points at DataSourcesManager as the taint source rather than at this change. Silencing the alerts in the code would mean building the DDL without the dialect, which breaks it on SQL Server. The lasting fix is a CodeQL model for DataSourcesManager, which is repository-wide work outside this PR.

Separately, the earlier smoke-tests cancellation wasn't a test failure. The test step passed (07:43 to 09:59). The job's last step, the checkout post-cleanup, then hung until the 150-minute job timeout. The push above starts a fresh run.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

migrations: a versioned data-migration artefact applied once per tenant after table sync, with a ledger table, readiness hold and health counts

2 participants