Skip to content

Bound the workerd collector's memory, ingest queue and stored spans - #5

Merged
RhysSullivan merged 2 commits into
mainfrom
fix/bounded-collector-memory
Sep 28, 2026
Merged

RhysSullivan merged 2 commits into
mainfrom
fix/bounded-collector-memory

Conversation

@RhysSullivan

@RhysSullivan RhysSullivan commented Sep 27, 2026 •

Copy link
Copy Markdown

The workerd collector runs inside Executor's self-host process and shares its memory limit. Its memory and store are now bounded by configuration.

  • One HTTP router per durable object instead of one per request. Building the typed API per request left ~6 MB of buffers behind each time; at 400 spans/s the isolate held 150–400 MiB of garbage buffers.
  • OTLP exports are parsed from the body and written straight to SQLite, without keeping the body text alive while spans are stored. With the router fix alone (2f5bd17), the isolate's heap grew about 22 MiB/min in an Executor soak until a full collection; with this change it stays level (see below).
  • At most MOTEL_OTEL_MAX_PENDING_INGEST exports (default 16) are read at once, each at most MOTEL_OTEL_MAX_INGEST_BYTES (default 16 MiB). Beyond that the collector answers 429 with Retry-After: 1 or 413.
  • Every export the collector does not store is counted by reason (queueFull, tooLarge, invalid, storeFailed) with its declared bytes. The counts are kept in the durable object's storage, so they survive eviction and restarts. GET /api/ingest reports them. Storage failures are logged each time with their cause; other refusals at most once a minute.
  • Malformed exports get 400. Store failures get 500, and a store that cannot be opened gets 503, so OTLP exporters retry them. Before this, a storage failure returned 400, which exporters treat as final.
  • Retention evicts until the store is back within age, MOTEL_OTEL_MAX_SPANS (new, default 1,000,000) and MOTEL_OTEL_MAX_DB_SIZE_MB, within a writer-time budget (MOTEL_OTEL_RETENTION_PASS_BUDGET_MS, default 500). The size bound is also enforced after each export, so the database file exceeds it by at most one export rather than by what arrives between passes. Before, a pass evicted at most one batch, and an Executor soak store reached 1.7 GB against a 1 GB bound.

Limits: workerd owns the SQLite write-ahead log, and PRAGMA journal_size_limit / wal_checkpoint are not authorized from a worker, so the WAL beside the file is not bounded by this configuration. An exporter that exhausts its retries on 429 drops what it buffers next without sending it, and the collector cannot count that.

In an Executor soak (30 minutes, ~480 spans/s, inspector sampling), the collector isolate's V8 heap on 89fbfd3 levelled at 87–92 MiB total from minute 10 to minute 25, and a forced collection left 17.5 MiB live. On 2f5bd17 it grew from 71 to 380 MiB over 14 minutes.

Tests are in the PR stacked above this one.

…nd per export

Malformed exports get 400; store failures get 500 or 503, which exporters retry.
Refusal counts persist in the collector's storage and include declared bytes.
@RhysSullivan
RhysSullivan changed the base branch from fix/ingest-cancellation to main September 28, 2026 13:12
@RhysSullivan
RhysSullivan marked this pull request as ready for review September 28, 2026 13:12
@RhysSullivan
RhysSullivan merged commit e59404f into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant