tenants: shut the delayed tenant provisioning down with its context - #7693
Merged
Merged
Conversation
TenantsInitializer scheduled the provisioning 30 s after ApplicationReady on a ScheduledThreadPoolExecutor it never shut down. A context closed within those 30 s still ran it later, through the closed context's beans, which Spring re-creates on demand - a new SystemDB pool, an EntityManagerFactory and the SystemDB Liquibase update - against a database that by then belonged to the next context. On the PostgreSQL CI leg the one LocalNativeAppLifecycleIT left behind fired 0.3 s after DirigibleCleaner had wiped the system schema and left the Liquibase changelog lock held: every later IT class in the slow shard then waited five minutes for the lock and failed to load its context. The executor is now a field of the bean and shut down in destroy(), so a provisioning that has not started when its context closes never runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cause
The PostgreSQL
slowshard of master and the nightly has been failing with the same cascade (runs 37275439965, 37088834792, 36903920667). Every class afterLocalNativeAppLifecycleIT(IntentEngineIT,ActAsSessionIT,ExternalFrontendTenantSelectionIT, ... 170 errors) fails after ~308 s withTenantsInitializerscheduled the tenant provisioning 30 s afterApplicationReadyEventon aScheduledThreadPoolExecutorit never shut down.LocalNativeAppLifecycleITboots one context per test method, each living about 13 s, so every one of its contexts left a provisioning timer behind. Those timers fired after their context had closed, five of them during the class, each on apool-N-thread-1thread. A late run goes through the closed context's beans, and Spring re-creates them on demand. The log shows a new SystemDB Hikari pool and the SystemDB Liquibase update running on that pool thread (07:28:21 [pool-13-thread-1] liquibase.changelog - Reading from public.databasechangelog). The last one fired at07:29:25.206, 0.1 s afterDirigibleCleanerhad dropped and recreated the PostgreSQL system schema. The next context's Liquibase then found the changelog lock held and waited for it, and so did every class after it. On H2 the schema lives in a file under the fork's owntarget/, which is why only the PostgreSQL leg shows it.The same leak exists outside tests: any context that closes within 30 s of becoming ready keeps a timer that later reaches into its closed beans.
Change
The executor is now a field of the bean and is shut down in
destroy(). A provisioning that has not started when its context closes never runs. The delay is still 30 s. A package-private constructor takes the delay so the unit test does not have to wait 30 s.Verification
mvn -pl components/core/core-tenants test: 33/33 green, including the newTenantsInitializerTest. It checks that the provisioning runs after the delay, and that a provisioning scheduled beforedestroy()never runs. Mutation check: withdestroy()turned into a no-op, the second test fails.-Dit.test=LocalNativeAppLifecycleIT,ActAsSessionIT -Dfailsafe.runOrder=reversealphabetical, i.e.ActAsSessionITbooting right afterLocalNativeAppLifecycleIT, which is the order that fails on CI. Result: 10/10 green andBUILD SUCCESS. The log has noTenantsInitializerrun after a context closed, no Liquibase on a non-mainthread and noWaiting for changelog lock. For comparison, the CI log of the same class on master has five late provisioning runs.formatter:validateon the module, with the cache wiped.slowshard. The timing that poisons the lock depends on the runner, so the shard's next run on master is the real proof.The other master failure, the
apishards cancelled at their 90-minute cap, is unrelated and is addressed separately.🤖 Generated with Claude Code