From 4a417164618c951ff4c9972eb15af08457189d7e Mon Sep 17 00:00:00 2001 From: Kostya Farber Date: Wed, 2 Sep 2026 16:20:23 +0000 Subject: [PATCH] fix(internal): stop the Anthropic token alert paging on cache reads and fix the cost rules at month rollover (#10654) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit This PR fixes a bug where `AnthropicTokenRateAnomaly` kept paging on ordinary agent and eval traffic after two rounds of threshold tuning (#10579, #10582). It also fixes the two Anthropic cost rules going NoData at the start of each month, and documents how the three rules work. ### Before The token rule compared every `gen_ai_usage_tokens_total` series (key × model × token type, about 70 of them) against 5x its own 7-day average. Two things made that unfixable by tuning: - Usage is bursty and mostly idle (median hour under 10 tokens/sec), so any baseline is near zero and a normal working hour clears 5x. - Cache reads counted the same as uncached input despite costing a tenth as much. On Sep 1 and 2 the rule fired repeatedly on Opus 5 cache reads from the internal evals key, a known workload that spent about $50 that day. The cost rules subtracted `gen_ai_cost offset 1d` from the current value. That metric is month-to-date, not lifetime, and updates once a day at about 03:00 UTC. At the rollover the subtraction has nothing to diff against, so both rules have been in "Normal (NoData)" since Sep 1. Their descriptions also claimed to measure "the last 24h", which the metric cannot do. ### After `AnthropicTokenBurst` (renamed from `AnthropicTokenRateAnomaly`, same uid) sums the token rate across all series, excludes `input_cache_read`, and fires above a flat 1000 tokens/sec with no baseline. Checked against the last 14 days of data in Grafana: it is true only during the Aug 21 to 22 Sonnet 4.5 incident and today's uncached Sonnet burst, and stays quiet during the evals traffic (which peaks around 500 non-cache-read tokens/sec). The cost rules use `increase(gen_ai_cost[1d])`, which treats the monthly reset as a counter reset and counts from zero. Checked at Aug 21, Aug 22, Aug 24 and Sep 2: the spike rule fires on the two incident days and nowhere else. The descriptions now say they describe the last completed day. The scrape gap on the 1st is still a blind spot and is noted in the README. The README gains an "Anthropic rules" section covering the two metric semantics and the reasoning behind each rule, so the next tuning attempt starts from that instead of rediscovering it. ### Implementation notes - The alertname label changes with the rename. Notification policies are hand-managed and the Anthropic rules carry no receiver label, so they route through the default policy either way. Any silence keyed on the old name stops applying, which is the intent. - The $1000/day critical threshold and the $50 floor are unchanged. The Aug 22 day was $325, so $1000 has never come close; lowering it is a separate decision. - Nothing was pushed to Grafana from a local machine. The deploy workflow pushes on merge. ### Change type - [x] `bugfix` ### Test plan 1. `gcx resources validate -p internal/grafana/resources` passes locally (alert rules report as "skipped: client-side check only", expected per the README). 2. Each new PromQL expression was run against the live datasource at the timestamps above with `gcx metrics query`. 3. After merge, confirm in Grafana that `AnthropicTokenBurst` shows a single instance and the two cost rules leave NoData at the next 03:00 UTC update. ### Code changes | Section | LOC change | | -------------- | ---------- | | Config/tooling | +36 / -37 | --- internal/grafana/README.md | 19 ++++++++++++++- .../8451fbba-5835-56e2-b2f5-7f5e44c520b8.yaml | 19 +++++++-------- .../e75d2f6c-44cf-54bf-8a74-ced2e045fcb6.yaml | 24 +++++-------------- .../eec2ad9e-30b9-5c8f-802f-53e26ab8fdc9.yaml | 11 ++++----- 4 files changed, 36 insertions(+), 37 deletions(-) diff --git a/internal/grafana/README.md b/internal/grafana/README.md index 92867aabece9..ff1d5d1c70a0 100644 --- a/internal/grafana/README.md +++ b/internal/grafana/README.md @@ -6,7 +6,7 @@ This directory currently provisions: - Eleven dashboards: `Zero on Fly.io — service health` (`zero-fly-health`), `Zero slow queries` (`slow-queries`), and `Zero-cache HTTP edge (Fly.io)` (`ni7hhgd`) in the `Zero` folder; `Anthropic scrape overview` in the `Anthropic` folder; `Dotcom events` (`ni5k8zc`), `Dotcom events V1` (`adgumkhou3chsd`), `File effect outbox`, `MCP server (shared boards)` (`mcp-server`), `MCP app sessions` (`kotx6wj`), and `Bemo analytics` in the `Dotcom` folder; and `Cloudflare platform metrics` (`nivs2rl`), which sits at the root rather than in a folder. The events pipeline currently reports from **staging** — set the environment variable accordingly when a dashboard looks empty. - All 13 Grafana-managed alert rules (Zero replication/errors/backups, Fly.io container health, Anthropic cost/usage, file room DO exceptions) -- The folders `Zero` (dashboards + 9 rules), `Dotcom` (6 dashboards + 1 rule), and `Anthropic` (1 dashboard + 3 cost rules) +- The folders `Zero` (dashboards + 9 rules), `Dotcom` (6 dashboards + 1 rule), and `Anthropic` (1 dashboard + 3 rules, see below) - The `discord` and `email` notification templates (`templates/*.gotmpl`, defining `tldraw.discord.title` / `tldraw.discord.text` and `tldraw.email.subject` / `tldraw.email.message`), which compress alert notifications to severity + summary + links (no Silence link on Discord: silence URLs carry one matcher per label and blow Discord’s 2000-char message cap on multi-instance alerts, truncating mid-link) and explain `DatasourceNoData` as a likely metrics-source outage. Templates can't go through `gcx resources push` — gcx ignores the `notifications.alerting.grafana.app` API group — so they live in `templates/` instead of `resources/` and the deploy workflow PUTs them through the classic provisioning API. Each contact point's optional Title/Subject and Message fields reference these templates; those references are set by hand (contact points stay hand-managed), so the templates only take effect while a contact point points at them. There are two MCP dashboards, for two different servers. `MCP server (shared boards)` covers the public server on the sync worker at `POST /app/mcp` (`apps/dotcom/sync-worker`), which reads its protocol metrics from the `MEASURE` dataset. `MCP app sessions` covers the separate tldraw MCP app worker (`apps/mcp-app`) and its own `MCP_ANALYTICS` dataset. Neither one's panels belong on `Dotcom events` — that dashboard covers the sync worker's own events. @@ -19,6 +19,23 @@ Contact points, notification policies, and data sources are hand-managed in Graf Every provisioned dashboard carries the tags `provisioned` and `repo:tldraw/tldraw`, so anyone looking at it in Grafana knows where it's managed from. Add both tags when adopting a new dashboard into this tree. +## Anthropic rules + +The three rules in the `Anthropic` folder watch the Grafana Anthropic integration's scrape of the Anthropic Admin API. The two metrics behave very differently, and the rules are shaped around that: + +- `gen_ai_usage_tokens_total` is near real time and is labelled by API key, model, and token type (`input`, `output`, `cache_creation_5m`, `cache_creation_1h`, `input_cache_read`). About 70 series at the time of writing. This is the only intraday signal. +- `gen_ai_cost` is **month-to-date** spend in cents, with no API key label, updated **once a day at about 03:00 UTC**. It resets on the 1st, and has had no data at all on the 1st. Anything built on it describes the last completed day, sees a runaway up to 27 hours late, and is blind for a day or two at each month rollover. + +`AnthropicTokenBurst` (warning) fires when the summed rate across all series, excluding cache reads, exceeds 1000 tokens/sec over an hour (about 3.6M tokens/hour). Three deliberate choices, each learned the hard way: + +- **Summed, not per series.** A per-series rule makes one alert instance per key × model × token type, and every lightly used series has a tiny baseline that any real session clears. +- **No baseline.** Usage is bursty and mostly idle (the median hour is under 10 tokens/sec), so a 7-day average is close to zero and "5x the average" is a normal working hour. #10579 and #10582 tuned the multiplier and floor twice and it still paged on ordinary sessions. +- **Cache reads excluded.** They cost a tenth of uncached input and dominate agent and eval traffic. On 2026-09-01 and 09-02 the rule paged repeatedly on Opus 5 cache reads from the internal evals key, which spent about $50 that day. The floor is chosen so that workload stays quiet (it peaks around 500 non-cache-read tokens/sec) while a big-numbers bulk cold-outreach research run (`tldraw-internal`, `crm-dashboard-api`), roughly 2000 tokens/sec of uncached Sonnet 4.5 input at $20 to $25 an hour and about $500 over 2026-08-21 to 08-22, clears it within the hour. Those runs are expected work, so a page from this rule usually means "check whether someone started a bulk run" rather than "something is broken". + +`AnthropicDailyCostSpike` (warning) fires when the last completed day cost more than 1.5x the day before, with a $50 floor. `AnthropicHighCostThreshold` (critical) fires when the last completed day cost more than $1000. Both use `increase(gen_ai_cost[1d])` rather than `offset` subtraction so the monthly reset reads as a counter reset and counts from zero instead of going negative. The scrape gap on the 1st is still a blind spot: the first day of the month reads as roughly $0 because there is nothing to diff against. Because the metric only steps once a day, a firing rule stays firing until the next 03:00 update. + +All three are filter queries (`expr > threshold`), so the alert series only exists while the condition holds. `missingSeriesEvalsToResolve: 10` is what stops a single quiet minute from resolving and re-firing them. + ## Making a change 1. Edit the files under `resources/`. diff --git a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/8451fbba-5835-56e2-b2f5-7f5e44c520b8.yaml b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/8451fbba-5835-56e2-b2f5-7f5e44c520b8.yaml index 7aece59deb96..b678dbbbd44b 100644 --- a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/8451fbba-5835-56e2-b2f5-7f5e44c520b8.yaml +++ b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/8451fbba-5835-56e2-b2f5-7f5e44c520b8.yaml @@ -14,9 +14,10 @@ metadata: namespace: stacks-890472 spec: annotations: - description: Token processing rate of {{ $value | humanize }} tokens/sec is 5x - higher than the 7-day average for job {{ $labels.job }}. - summary: Anthropic API token rate anomaly detected. + description: 'Anthropic API usage over the last hour averaged {{ $value | humanize }} + tokens/sec excluding cache reads, above the 1000 tokens/sec floor (about 3.6M + tokens/hour). Check the scrape overview dashboard for which key and model.' + summary: Anthropic API token burst detected. execErrState: Ok expressions: prometheus_math: @@ -54,13 +55,9 @@ spec: type: prometheus uid: grafanacloud-prom expr: | - ( - rate(gen_ai_usage_tokens_total{job=~"integrations/anthropic/.*"}[1h]) - > - 5 * avg_over_time(rate(gen_ai_usage_tokens_total{job=~"integrations/anthropic/.*"}[1h])[7d:1h]) - ) - and - rate(gen_ai_usage_tokens_total{job=~"integrations/anthropic/.*"}[1h]) > 1000 + sum( + rate(gen_ai_usage_tokens_total{job=~"integrations/anthropic/.*", gen_ai_token_type!="input_cache_read"}[1h]) + ) > 1000 instant: true intervalMs: 1000 maxDataPoints: 43200 @@ -110,6 +107,6 @@ spec: severity: warning missingSeriesEvalsToResolve: 10 noDataState: Ok - title: AnthropicTokenRateAnomaly + title: AnthropicTokenBurst trigger: interval: 1m diff --git a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/e75d2f6c-44cf-54bf-8a74-ced2e045fcb6.yaml b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/e75d2f6c-44cf-54bf-8a74-ced2e045fcb6.yaml index dcd1a15f7a20..f617a720b682 100644 --- a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/e75d2f6c-44cf-54bf-8a74-ced2e045fcb6.yaml +++ b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/e75d2f6c-44cf-54bf-8a74-ced2e045fcb6.yaml @@ -14,8 +14,9 @@ metadata: namespace: stacks-890472 spec: annotations: - description: 'Cost over the last 24h of ${{ $value | humanize }} is more than - 1.5x the previous 24h.' + description: 'The last completed day cost ${{ $value | humanize }}, more than + 1.5x the day before. The cost metric updates once a day at about 03:00 UTC, + so this describes yesterday.' summary: Anthropic API daily cost spike detected. execErrState: Ok expressions: @@ -55,25 +56,12 @@ spec: uid: grafanacloud-prom expr: | ( - ( - sum(gen_ai_cost{job=~"integrations/anthropic/.*"}) - - - sum(gen_ai_cost{job=~"integrations/anthropic/.*"} offset 1d) - ) / 100 + sum(increase(gen_ai_cost{job=~"integrations/anthropic/.*"}[1d])) / 100 > - 1.5 * - ( - sum(gen_ai_cost{job=~"integrations/anthropic/.*"} offset 1d) - - - sum(gen_ai_cost{job=~"integrations/anthropic/.*"} offset 2d) - ) / 100 + 1.5 * sum(increase(gen_ai_cost{job=~"integrations/anthropic/.*"}[1d] offset 1d)) / 100 ) and - ( - sum(gen_ai_cost{job=~"integrations/anthropic/.*"}) - - - sum(gen_ai_cost{job=~"integrations/anthropic/.*"} offset 1d) - ) / 100 > 50 + sum(increase(gen_ai_cost{job=~"integrations/anthropic/.*"}[1d])) / 100 > 50 instant: true intervalMs: 1000 maxDataPoints: 43200 diff --git a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/eec2ad9e-30b9-5c8f-802f-53e26ab8fdc9.yaml b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/eec2ad9e-30b9-5c8f-802f-53e26ab8fdc9.yaml index 38d0e9eac256..1ebed9ec8a23 100644 --- a/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/eec2ad9e-30b9-5c8f-802f-53e26ab8fdc9.yaml +++ b/internal/grafana/resources/alertrules.v0alpha1.rules.alerting.grafana.app/eec2ad9e-30b9-5c8f-802f-53e26ab8fdc9.yaml @@ -14,8 +14,9 @@ metadata: namespace: stacks-890472 spec: annotations: - description: 'Cost over the last 24h of ${{ $value | humanize }} has exceeded - the $1000 threshold. Current cost: ${{ $value }}.' + description: 'The last completed day cost ${{ $value | humanize }}, above the + $1000 threshold. The cost metric updates once a day at about 03:00 UTC, so + this describes yesterday.' summary: Anthropic API cost threshold exceeded. execErrState: Ok expressions: @@ -54,11 +55,7 @@ spec: type: prometheus uid: grafanacloud-prom expr: | - ( - sum(gen_ai_cost{job=~"integrations/anthropic/.*"}) - - - sum(gen_ai_cost{job=~"integrations/anthropic/.*"} offset 1d) - ) / 100 > 1000 + sum(increase(gen_ai_cost{job=~"integrations/anthropic/.*"}[1d])) / 100 > 1000 instant: true intervalMs: 1000 maxDataPoints: 43200