Skip to content

Run hourly usageStatistics sequentially with lean queries - #871

Open
stevenweaver wants to merge 2 commits into
masterfrom
fix/usage-statistics-memory-burst
Open

stevenweaver wants to merge 2 commits into
masterfrom
fix/usage-statistics-memory-burst

Conversation

@stevenweaver

@stevenweaver stevenweaver commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Problem

lib/usageStatistics.js runs setInterval(usageStatisticsLooper, 3600000), and every hour the looper starts all 12 per-method usageStatistics() calls at the same moment. Each call loads the last year of completed jobs for its collection (created, msa.sites, msa.sequences) and builds every row into a full Mongoose document, then turns the whole result set into JSON for Redis. With all 12 running at once, those result sets and their JSON strings sit in memory together. The timer is anchored to process start, so the burst lands at the same minute past the hour every hour.

Evidence summary

  • On the production host, the hourly run causes a memory burst of about 0.2–1.1 GiB.
  • Two global OOM events happened in the same minute as this hourly run (2026-09-07 and 2026-09-23). In the second one the kernel killed mongod.
  • The query results only go to JSON.stringify and a Redis SET. No document methods, virtuals or saves are used, so building full documents is wasted work.

Change

  • lib/usageStatistics.js: the per-method usageStatistics() calls now run one after another through a small callback chain (the codebase uses callbacks, and the pinned async@^0.1 is very old). A module-level start time skips a tick (with a warning) if the previous run is still going; the guard expires after 3 hours so a hung run can't stop stats refreshing permanently. Errors from each method are now logged instead of silently dropped. The hourly cadence is unchanged.
  • app/models/analysis.js: both queries in AnalysisSchema.statics.usageStatistics now use .lean(). The projection is unchanged, and the JSON stored in Redis has the same fields (created, msa[].sites, msa[].sequences).

How to test

  1. node --check lib/usageStatistics.js app/models/analysis.js
  2. Locally, temporarily switch to the commented-out short interval, or call the exported usageStatisticsLooper() once after boot. Check that the calls run one at a time (only one method's query in flight in db.currentOp()), and that each <db>_<collection>_job_stats Redis key is filled in.
  3. Load a usage/stats page (for example the FEL usage endpoint that reads FEL.cachePath()) and confirm the charts render as before.
  4. Optional: watch the Node process RSS across the hourly run before and after the change. The peak should drop a lot.

Deploy notes

  • The change only touches the frontend (lib/usageStatistics.js, app/models/analysis.js). No schema, config or dependency changes.
  • Apply with a normal pull and a restart of the datamonkey-js process.
  • The first run happens one hour after the process restarts, as before.

The hourly usageStatistics looper fired all 12 per-method queries at
once, each hydrating a year of completed jobs into full Mongoose
documents. The resulting memory burst lined up with host-level OOMs.

- Run the per-method calls one at a time and skip a tick if the
  previous run has not finished.
- Use .lean() for both usageStatistics queries; results are only
  serialized to Redis, so hydration was pure overhead.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant