Skip to content

Stop page loads from spending the whole Cloud Armor rate-limit budget - #152

Merged
rcurranmoz merged 1 commit into
mainfrom
fix/rate-limit-fanout
Sep 17, 2026
Merged

rcurranmoz merged 1 commit into
mainfrom
fix/rate-limit-fanout

Conversation

@rcurranmoz

@rcurranmoz rcurranmoz commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

An operator hit a bare 429 Too Many Requests page twice on 2026-09-17 simply by refreshing the dashboard. This fixes the cause on both sides: the page asks for far less, and the limit stops being able to blank the app.

What the load-balancer logs actually show

The denied minute was exactly 100 allowed, then 6 denied — hangar-armor's threshold is 100/60s per IP, and it counts the HTML document and hashed assets alongside API calls. Breakdown of those 100:

Path Count
/api/fleet/pool-sources 35
/api/fleet/android-pool-sources 14
other API (28 distinct endpoints) 39
document + assets + favicon 12

A repeat 12 minutes later had the identical signature. In both cases the document was denied last — the API burst spent the budget, then the navigation was refused. That ordering is why the symptom is a blank error page rather than a degraded card.

So two unbatched per-pool loops were 49% of the entire budget. Both per-pool computations were already cached server-side, so the cost was purely request count — which is exactly what a per-IP limiter charges for.

Changes

Batch the per-pool loops. New /fleet/pool-sources-batch and /fleet/android-pool-sources-batch build on the same cache keys as the single-pool routes, so per-pool payloads are byte-identical and no card changes behaviour. Cold pools are computed with bounded parallelism rather than serially. The four call sites — Pools.tsx (pinned, linux/windows, android) and Overview.tsx (monitored) — now issue one request each: 49 → 3.

Cache android_pool_sources. It was the one uncached path: a live Taskcluster workers call (limit=1000) plus a task fetch per running task, per pool, 14 times per Pools load. Now cached like its hardware sibling.

In-flight GET dedupe in api.ts. Components that mount together request the same endpoint — Layout + Overview both want /fleet/summary; Workers + CommandPalette both want /fleet/pools. Identical concurrent GETs share one response; nothing is cached past settlement, so no semantics change.

Retry with jittered backoff on 429/502/503/504, honouring Retry-After. A throttled API call now degrades one card instead of erroring the page.

Split the Terraform limit in two. Denying an /api call degrades a card inside a loaded page; denying the document replaces the whole dashboard with Cloud Armor's error page. Those deserve different ceilings: API 1200/60s (~24 heavy page loads/min after batching), app shell 3000/60s.

IAP already restricts this origin to @mozilla.com, so these limits are DoS hygiene, not access control. And enforce_on_key = "IP" means corp VPN/NAT users share one budget — which is why the headroom matters more as the audience grows, not less.

Net effect: a heavy load goes ~100 → ~54 requests while the /api ceiling goes 100 → 1200.

Recorded but deliberately NOT changed here

The three OWASP rules are unreachable — 2000 XSS, 2001 SQLi, 2002 RFI. Cloud Armor enforces the first matching rule and stops; these rate-limit rules match every request. Verified in the LB logs — every request reports enforcedSecurityPolicy priority 1000, including a scanner walking /cgi-bin and /docSQL, which gets a 429 rather than a WAF 403. Today the rate limiter is the only Cloud Armor control actually running.

Activating them means reordering, which carries real false-positive risk and wants preview = true first to measure. That's its own change; there's a comment in lb.tf so the next reader isn't surprised.

Verification

  • 28 backend tests pass (12 new in tests/test_pool_sources_batch.py).
  • The new tests pin the fix rather than document it: reverting just the android cache wrapper fails test_android_single_route_is_now_cached.
  • npm run build, terraform fmt -check, terraform validate all clean.
  • eslint (6 problems) and ruff (7) at exact parity with main on every touched file — diffed per-file rather than trusting a total.

The request-count reduction is derived from the changed call sites plus the measured before-state; the after-number comes from LB logs post-deploy:

gcloud logging read 'resource.type="http_load_balancer" AND httpRequest.status=429' \
  --project=relops-dashboard --freshness=1d --limit=100 \
  --format="value(httpRequest.remoteIp,httpRequest.requestUrl)"

No window where traffic is denied

The plan removes one rule and adds two, which raises the fair question of what happens in between. The policy's own default rule is 2147483647 / allow ("Default allow"), so during the swap a request that matches no rate-limit rule falls through to allow. The failure direction is fail-open — briefly un-rate-limited, never denied — and IAP authenticates throughout, since it lives on the backend service and is independent of the security policy.

Worth one post-apply check: the new rules plan as preview = (known after apply). They must come back false. If either landed as true the rule would log instead of enforce, which is harmless but would mean no rate limiting at all.

Deploy note

The code half auto-deploys on merge. The Terraform half is a manual apply and is the part that makes the blank page impossible — merging alone roughly halves consumption but leaves the 100/min ceiling in place. A good plan touches only google_compute_security_policy.hangar; if it wants to change iap{} or the Cloud Run image, stop.

🤖 Generated with Claude Code

An operator hit a bare "429 Too Many Requests" page twice on 2026-09-17 by
refreshing the dashboard. Measured from the load-balancer logs, the denied
minute was exactly 100 requests allowed then 6 denied -- the rule's threshold
is 100/60s per IP, and it counts the document and assets as well as API calls.

Half the budget went to two unbatched per-pool loops:

    /api/fleet/pool-sources          35
    /api/fleet/android-pool-sources  14
    other API (28 endpoints)         39
    document + assets + favicon      12

Both per-pool computations were already cached server-side, so the cost was
purely request COUNT -- exactly what a per-IP limiter charges for. Hence:

* Batch endpoints (`/fleet/pool-sources-batch`, `/fleet/android-pool-sources-batch`)
  built on the same cache keys as the single-pool routes, so per-pool payloads
  are byte-identical and no card changes behaviour. Cold pools are computed with
  bounded parallelism rather than serially. The four call sites in Pools.tsx (3)
  and Overview.tsx (1) now issue one request each: 49 -> 3.
* `android_pool_sources` is now cached like its hardware sibling. It was the one
  uncached path -- a live Taskcluster workers call (limit=1000) plus a task fetch
  per running task, per pool, 14 times per Pools load.
* In-flight GET dedupe in api.ts. Components that mount together ask for the same
  endpoint (Layout + Overview both want /fleet/summary; Workers + CommandPalette
  both want /fleet/pools), and a duplicate is never free under a shared budget.
* Retry with jittered backoff on 429/502/503/504, honouring Retry-After, so a
  throttled API call degrades one card instead of erroring the page.

On the Terraform side the limit is split in two. Denying an /api call degrades a
card in a loaded page; denying the document replaces the entire dashboard with
Cloud Armor's error page, which is what was actually reported. API calls get
1200/60s (~24 heavy page loads/min after batching), the app shell 3000/60s. IAP
already restricts this origin to @mozilla.com, so these limits are DoS hygiene,
not access control -- and `enforce_on_key = "IP"` means corp VPN/NAT users share
one budget, so the headroom matters more as the audience grows, not less.

Recorded in a comment but deliberately NOT changed here: Cloud Armor enforces the
first matching rule and stops, and these rules match every request, so the three
OWASP rules (2000 XSS, 2001 SQLi, 2002 RFI) are unreachable. Verified in the LB
logs -- every request, including a scanner walking /cgi-bin and /docSQL, reports
enforcedSecurityPolicy priority 1000. The policy's own default rule at
2147483647 is "Default allow", so the swap in this commit has no window in which
traffic is denied -- only a brief one in which it is not rate limited. Activating them is a separate change that
wants preview = true first to measure false positives.

Verified: 28 backend tests pass (12 new); reverting just the android cache fails
test_android_single_route_is_now_cached, so it pins the fix rather than
documenting it. Frontend build, terraform fmt and terraform validate all clean.
eslint and ruff at exact parity with main on every touched file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@rcurranmoz
rcurranmoz force-pushed the fix/rate-limit-fanout branch from 6c9bc38 to 3033de3 Compare September 17, 2026 14:41
@rcurranmoz
rcurranmoz merged commit cd271ae into main Sep 17, 2026
5 checks passed
@rcurranmoz
rcurranmoz deleted the fix/rate-limit-fanout branch September 17, 2026 14:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant