Stop page loads from spending the whole Cloud Armor rate-limit budget - #152
Merged
Merged
Conversation
An operator hit a bare "429 Too Many Requests" page twice on 2026-09-17 by
refreshing the dashboard. Measured from the load-balancer logs, the denied
minute was exactly 100 requests allowed then 6 denied -- the rule's threshold
is 100/60s per IP, and it counts the document and assets as well as API calls.
Half the budget went to two unbatched per-pool loops:
/api/fleet/pool-sources 35
/api/fleet/android-pool-sources 14
other API (28 endpoints) 39
document + assets + favicon 12
Both per-pool computations were already cached server-side, so the cost was
purely request COUNT -- exactly what a per-IP limiter charges for. Hence:
* Batch endpoints (`/fleet/pool-sources-batch`, `/fleet/android-pool-sources-batch`)
built on the same cache keys as the single-pool routes, so per-pool payloads
are byte-identical and no card changes behaviour. Cold pools are computed with
bounded parallelism rather than serially. The four call sites in Pools.tsx (3)
and Overview.tsx (1) now issue one request each: 49 -> 3.
* `android_pool_sources` is now cached like its hardware sibling. It was the one
uncached path -- a live Taskcluster workers call (limit=1000) plus a task fetch
per running task, per pool, 14 times per Pools load.
* In-flight GET dedupe in api.ts. Components that mount together ask for the same
endpoint (Layout + Overview both want /fleet/summary; Workers + CommandPalette
both want /fleet/pools), and a duplicate is never free under a shared budget.
* Retry with jittered backoff on 429/502/503/504, honouring Retry-After, so a
throttled API call degrades one card instead of erroring the page.
On the Terraform side the limit is split in two. Denying an /api call degrades a
card in a loaded page; denying the document replaces the entire dashboard with
Cloud Armor's error page, which is what was actually reported. API calls get
1200/60s (~24 heavy page loads/min after batching), the app shell 3000/60s. IAP
already restricts this origin to @mozilla.com, so these limits are DoS hygiene,
not access control -- and `enforce_on_key = "IP"` means corp VPN/NAT users share
one budget, so the headroom matters more as the audience grows, not less.
Recorded in a comment but deliberately NOT changed here: Cloud Armor enforces the
first matching rule and stops, and these rules match every request, so the three
OWASP rules (2000 XSS, 2001 SQLi, 2002 RFI) are unreachable. Verified in the LB
logs -- every request, including a scanner walking /cgi-bin and /docSQL, reports
enforcedSecurityPolicy priority 1000. The policy's own default rule at
2147483647 is "Default allow", so the swap in this commit has no window in which
traffic is denied -- only a brief one in which it is not rate limited. Activating them is a separate change that
wants preview = true first to measure false positives.
Verified: 28 backend tests pass (12 new); reverting just the android cache fails
test_android_single_route_is_now_cached, so it pins the fix rather than
documenting it. Frontend build, terraform fmt and terraform validate all clean.
eslint and ruff at exact parity with main on every touched file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rcurranmoz
force-pushed
the
fix/rate-limit-fanout
branch
from
September 17, 2026 14:41
6c9bc38 to
3033de3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An operator hit a bare
429 Too Many Requestspage twice on 2026-09-17 simply by refreshing the dashboard. This fixes the cause on both sides: the page asks for far less, and the limit stops being able to blank the app.What the load-balancer logs actually show
The denied minute was exactly 100 allowed, then 6 denied —
hangar-armor's threshold is 100/60s per IP, and it counts the HTML document and hashed assets alongside API calls. Breakdown of those 100:/api/fleet/pool-sources/api/fleet/android-pool-sourcesA repeat 12 minutes later had the identical signature. In both cases the document was denied last — the API burst spent the budget, then the navigation was refused. That ordering is why the symptom is a blank error page rather than a degraded card.
So two unbatched per-pool loops were 49% of the entire budget. Both per-pool computations were already cached server-side, so the cost was purely request count — which is exactly what a per-IP limiter charges for.
Changes
Batch the per-pool loops. New
/fleet/pool-sources-batchand/fleet/android-pool-sources-batchbuild on the same cache keys as the single-pool routes, so per-pool payloads are byte-identical and no card changes behaviour. Cold pools are computed with bounded parallelism rather than serially. The four call sites —Pools.tsx(pinned, linux/windows, android) andOverview.tsx(monitored) — now issue one request each: 49 → 3.Cache
android_pool_sources. It was the one uncached path: a live Taskcluster workers call (limit=1000) plus a task fetch per running task, per pool, 14 times per Pools load. Now cached like its hardware sibling.In-flight GET dedupe in
api.ts. Components that mount together request the same endpoint — Layout + Overview both want/fleet/summary; Workers + CommandPalette both want/fleet/pools. Identical concurrent GETs share one response; nothing is cached past settlement, so no semantics change.Retry with jittered backoff on 429/502/503/504, honouring
Retry-After. A throttled API call now degrades one card instead of erroring the page.Split the Terraform limit in two. Denying an
/apicall degrades a card inside a loaded page; denying the document replaces the whole dashboard with Cloud Armor's error page. Those deserve different ceilings: API1200/60s(~24 heavy page loads/min after batching), app shell3000/60s.IAP already restricts this origin to
@mozilla.com, so these limits are DoS hygiene, not access control. Andenforce_on_key = "IP"means corp VPN/NAT users share one budget — which is why the headroom matters more as the audience grows, not less.Net effect: a heavy load goes ~100 → ~54 requests while the
/apiceiling goes 100 → 1200.Recorded but deliberately NOT changed here
The three OWASP rules are unreachable —
2000XSS,2001SQLi,2002RFI. Cloud Armor enforces the first matching rule and stops; these rate-limit rules match every request. Verified in the LB logs — every request reportsenforcedSecurityPolicy priority 1000, including a scanner walking/cgi-binand/docSQL, which gets a 429 rather than a WAF 403. Today the rate limiter is the only Cloud Armor control actually running.Activating them means reordering, which carries real false-positive risk and wants
preview = truefirst to measure. That's its own change; there's a comment inlb.tfso the next reader isn't surprised.Verification
tests/test_pool_sources_batch.py).test_android_single_route_is_now_cached.npm run build,terraform fmt -check,terraform validateall clean.mainon every touched file — diffed per-file rather than trusting a total.The request-count reduction is derived from the changed call sites plus the measured before-state; the after-number comes from LB logs post-deploy:
No window where traffic is denied
The plan removes one rule and adds two, which raises the fair question of what happens in between. The policy's own default rule is
2147483647 / allow("Default allow"), so during the swap a request that matches no rate-limit rule falls through to allow. The failure direction is fail-open — briefly un-rate-limited, never denied — and IAP authenticates throughout, since it lives on the backend service and is independent of the security policy.Worth one post-apply check: the new rules plan as
preview = (known after apply). They must come backfalse. If either landed astruethe rule would log instead of enforce, which is harmless but would mean no rate limiting at all.Deploy note
The code half auto-deploys on merge. The Terraform half is a manual
applyand is the part that makes the blank page impossible — merging alone roughly halves consumption but leaves the 100/min ceiling in place. A good plan touches onlygoogle_compute_security_policy.hangar; if it wants to changeiap{}or the Cloud Runimage, stop.🤖 Generated with Claude Code