Run hourly usageStatistics sequentially with lean queries - #871
Open
stevenweaver wants to merge 2 commits into
Open
stevenweaver wants to merge 2 commits into
stevenweaver wants to merge 2 commits into
Conversation
The hourly usageStatistics looper fired all 12 per-method queries at once, each hydrating a year of completed jobs into full Mongoose documents. The resulting memory burst lined up with host-level OOMs. - Run the per-method calls one at a time and skip a tick if the previous run has not finished. - Use .lean() for both usageStatistics queries; results are only serialized to Redis, so hydration was pure overhead.
… run can't block stats forever
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
lib/usageStatistics.jsrunssetInterval(usageStatisticsLooper, 3600000), and every hour the looper starts all 12 per-methodusageStatistics()calls at the same moment. Each call loads the last year of completed jobs for its collection (created,msa.sites,msa.sequences) and builds every row into a full Mongoose document, then turns the whole result set into JSON for Redis. With all 12 running at once, those result sets and their JSON strings sit in memory together. The timer is anchored to process start, so the burst lands at the same minute past the hour every hour.Evidence summary
mongod.JSON.stringifyand a RedisSET. No document methods, virtuals or saves are used, so building full documents is wasted work.Change
lib/usageStatistics.js: the per-methodusageStatistics()calls now run one after another through a small callback chain (the codebase uses callbacks, and the pinnedasync@^0.1is very old). A module-level start time skips a tick (with a warning) if the previous run is still going; the guard expires after 3 hours so a hung run can't stop stats refreshing permanently. Errors from each method are now logged instead of silently dropped. The hourly cadence is unchanged.app/models/analysis.js: both queries inAnalysisSchema.statics.usageStatisticsnow use.lean(). The projection is unchanged, and the JSON stored in Redis has the same fields (created,msa[].sites,msa[].sequences).How to test
node --check lib/usageStatistics.js app/models/analysis.jsusageStatisticsLooper()once after boot. Check that the calls run one at a time (only one method's query in flight indb.currentOp()), and that each<db>_<collection>_job_statsRedis key is filled in.FEL.cachePath()) and confirm the charts render as before.Deploy notes
lib/usageStatistics.js,app/models/analysis.js). No schema, config or dependency changes.