Repository navigation
Clarify cache behavior with parallel test processes #80
Description
Activity
Hi there! Very good question, but unfortunately, we don't have parallel testing in our suite so I can't test it out. I think the best to do is to set
config.cache_path = "tmp/cache/fixture_kit/#{ENV["TEST_ENV_NUMBER"]}"so that each test process gets its own cache directory so they don't clobber each other. There might be some edge cases? I'm not sure. If anyone doing this has some other workaround, I'm happy to incorporate doc updates or fixes.We run FixtureKit on a large suite in parallel, so here is what we found. The answer to "can one process reuse another's cache files" turns out to be different for each runner, which is probably why this is confusing. I've opened #82 to put it in the reference.
Rails
parallelize— files are shared and reused across processes, by design.ActiveSupport::Testing::Parallelizationforks its workers inMinitest.run, before any suite runs, and each worker runs individual test methods (Minitest.run_one_method), neverrun_suite. FixtureKit generates fromFixtureKit::Minitest::ClassMethods#run_suite, soRunner#startand everygeneratehappen in the parent and the workers onlymountwhat the parent wrote. One sharedcache_pathis correct here, and it is what the default gives you. There is also nothing to key the path on: Rails names the per-worker databases itself and never setsTEST_ENV_NUMBER(no reference to it in activerecord, activesupport or railties). Keying on a worker number would point the workers at directories the parent never writes to.parallel_tests— your suggestion is right. Each worker is a full process that boots the app and callsRunner#start, so each one clears the directory the others are generating into or mounting from.config.cache_path = "tmp/cache/fixture_kit/#{ENV["TEST_ENV_NUMBER"]}"fixes it; the cost is that every worker generates its own copy of the fixtures its share of the suite needs.One directory shared between processes — a warm-up run that fills the cache before the suite, or a CI cache restored between jobs. This needs
FIXTURE_KIT_PRESERVE_CACHE, otherwise the next process to start deletes it, and preserving it hands invalidation to the caller:Cache#exists?isFile.exist?, so nothing compares that file against the definitions, the factories they call, or the schema. This is where it bit us. A cache written before a PaperTrail guard landed in our test helper kept being mounted, and surfaced asPG::UniqueViolation: versions_pkeyin tests with no visible relationship to fixtures. We bound it by folding a digest of the definitions, factories and schema intocache_path, so a cache written under any other state of the code is ignored and regenerated rather than read. #82 documents the pattern and its two costs (obsolete digest directories are not reclaimed, and one changed file regenerates everything).I also opened #83 for one thing that can only be fixed in the gem:
FileCache#writeis a plainFile.write, which truncates the destination and fills it back in, so a shared directory can hand a concurrent reader a prefix of the JSON and aJSON::ParserErrorthat looks like a corrupt cache. Writing a sibling file and renaming it over the destination makes that window disappear.Happy to adjust the wording in #82 if any of it reads wrong to you.
@ngan
Thanks for the reply!As @navidemad suggested, our team enabled
FIXTURE_KIT_PRESERVE_CACHE, generated the cache files before running the parallel tests, and then shared the samecache_pathacross multiple processes. This worked without issues and significantly improved our test performance.Using a process-specific path such as
config.cache_path = "tmp/cache/fixture_kit/#{ENV["TEST_ENV_NUMBER"]}"is certainly one possible workaround. However, each process would need to generate its own cache files, which may reduce the performance benefits of FixtureKit.I think a brief mention of the shared-cache approach in the documentation would be sufficient. What do you think?
Hi! Thanks for FixtureKit.
I'd like to clarify how the cache behaves with process-based parallel tests such as
parallel_tests.The docs mention that cache data is mounted per
connection_pool. If multiple test processes use the samecache_path, can cache files generated by one process be reused by the other processes, or is the cache effectively scoped to each process/connection pool?If the cache is shared across processes, would you consider mentioning this explicitly in the docs?
I think this would be helpful since teams using FixtureKit for large test suites are also likely to run tests in parallel.