Part of the StatusBus documentation set. Overview & quick start:
../README.md
StatusBus is a globally-distributed website uptime monitoring service. It orchestrates consumers running in VMs across different world regions that ping a set of websites at frequent intervals, measure response times, and record the results (ticks) into a central database. Users register, add websites through a web dashboard, and see the latest status/response time for each site.
The multi-region deployment (live) runs the edge (API + dashboard) on Cloud Run and the monitoring mesh on disposable VMs per region — India consumers on a self-managed kubeadm cluster, the US consumer on a standalone spot VM:
- The API and database are centralized (Cloud Run + Neon Postgres; Redis on Upstash).
- A single shared Redis Stream fans monitoring jobs out to every region.
- Each region runs its own consumers (a consumer group named after the region) so that every region independently checks every website.
- All check results are aggregated into Postgres for the dashboard and future analytics.
The worker infrastructure (bootstrap scripts, chaos-proven self-healing) is documented in
gcp-infra/k8s/README.md and its rationale in
ARCHITECTURE-DECISIONS.md.
┌───────────────────────────────────────────────┐
│ User / Browser │
└──────────────────────┬────────────────────────┘
│ HTTPS (browser)
▼
┌──────────────────────┴────────────────────────┐
│ apps/fe (Next.js) │
│ landing · signup/signin · dashboard │
└──────────────────────┬────────────────────────┘
│ CORS + JWT (GET /websites,
│ POST /website, ...)
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ apps/api (Express 5) │
│ REST API · JWT auth · zod validation · Prisma (packages/store) │
│ Public endpoints: user/signup, user/signin, website, websites, status/:id │
│ Internal endpoints: monitoring/websites, monitoring/tick, health │
└─────────────┬────────────────────────────────────┬───────────────────────────┘
│ │ POST /monitoring/tick
│ GET /monitoring/websites │ {website_id, region_id, rt_ms, status}
│ [{id, url, user_id}] ▼
│ ┌───────────────────────────┐
▼ │ Postgres (Prisma/store) │
┌──────────────────────┐ │ User → Website → Tick │
│ apps/producer │ │ Region │
│ polls API every 60s │ └───────────────────────────┘
│ XADD → stream │
└──────────┬───────────┘
│ XADD statusbus:web {url, id} (per website, 1/min)
▼
┌─────────────────────────────────┐
│ Redis Stream: statusbus:web │
│ (shared, fan-out to regions) │
└──────────────┬──────────────────┘
│ XREADGROUP GROUP=<region_id> CONSUMER=<consumer_id> COUNT 5
▼
┌──────────────────────────────────────────────────────────────┐
│ apps/consumer — one consumer group per region │
│ e.g. group "1" (India): consumer india-consumer-1 │
│ group "2" (US): consumer us-consumer-1 │
│ for each job: axios.get(url) → rt_ms → status Up/Down │
│ → POST /monitoring/tick → XACK │
└──────────────────────────────────────────────────────────────┘
The monitoring substrate is a single Redis Stream named statusbus:web. The design is
fan-out, not partitioning:
- The
producerenqueues every website once per cycle, regardless of region. - Redis consumer groups are keyed by the region id (
REGION_ID). Each region forms its own group over the same stream. - Because every group reads with
XREADGROUP ... >(new messages only), every region receives a full copy of every job and therefore checks every website locally. - Individual consumers within a group (identified by
CONSUMER_ID) share the group's load, so scaling a region means adding more consumer replicas.
This gives:
- Independence: a check from India and a check from the US both validate a site’s availability; regional network issues don’t affect other regions.
- Simplicity: no job routing, no mapping websites to regions, no per-region queues.
| Component | Role in the fan-out |
|---|---|
| Producer | Global scheduler. Polls the API for the full website list, enqueues every site. |
| Redis Stream | Job buffer / queue + delivery ledger (via consumer-group pending list). |
| Consumer (per region) | Reads its region's copy of the queue, performs the actual probe, reports a tick. |
Express 5 (TypeScript) single-file app (apps/api/index.ts), run on the Bun runtime.
- Auth: JWT (HS256,
jsonwebtoken). Signed payload{ sub: user.id }. No expiry. Tokens are sent in the rawAuthorizationheader (noBearerprefix). - Validation:
zod(AuthInput→{ username, password }). - Database access: exclusively through
packages/store(import { prisma } from "store/client"). - CORS: allow list of the frontend origins (
https://statusbus.byaniket.site,http://localhost:3000). - Endpoints (8): see
docs/WORKFLOW.mdfor the full reference table.
- App Router + Tailwind v4 + shadcn/ui tokens + lucide icons.
- Marketing landing page (
/),/signup,/signin,/dashboard, and a placeholder/website/[websiteId]route. - Talks to the API cross-origin via
NEXT_PUBLIC_BACKEND_URL(inlined at build time, defaulthttps://api-statusbus.byaniket.site). - Stores the JWT in
localStorage["token"]and sends it in theAuthorizationheader. - No state-management library, no realtime updates (dashboards fetch once per mount), no per-website detail view or charts yet.
- Loop: every
PRODUCER_INTERVAL_SEC(default 300 s),GET {API_URL}/monitoring/websiteswith thex-internal-keyheader → for each site,XADD statusbus:web * url <url> id <id>(viapackages/redisq), thencapStream()trims the stream to ~100 K entries. - Runs one cycle immediately on start, then on the interval.
- Retries API connectivity 10× with 2 s backoff before giving up.
- Region-agnostic; only enqueues. Note: winnows the API list to
{url, id}— the stream carries only those two fields.
- Requires
REGION_ID(string) andCONSUMER_ID. - Consumer group name =
REGION_ID; consumer name =CONSUMER_ID. - Creates the group idempotently (
XGROUP CREATE ... MKSTREAM, swallowsBUSYGROUP). - Loop:
XREADGROUPwithCOUNT 5,id: '>'; idle-sleepCONSUMER_POLL_SEC(default 60 s) when no jobs. EveryCONSUMER_RECLAIM_INTERVAL_SEC(default 300 s) it runsXAUTOCLAIM(min idle 5 min, count 10) to reclaim in-flight messages whose consumer vanished (spot eviction). - For each job:
axios.get(url, { timeout: 10_000 }); on resolve →status: "Up", on reject →status: "Down"— wherert_ms = Date.now() - startTime; thenPOST /monitoring/tick(withx-internal-key) andXACKthe event, always (.finally). - Up to 5 probes run concurrently per batch (
Promise.all).
- Prisma 7 +
@prisma/adapter-pgPostgreSQL driver adapter. - Singleton Prisma client cached on
globalThisin dev. - Schema below; migrations in
packages/store/prisma/migrations/. - Generated client under
packages/store/generated/prisma/(git-ignored, regenerated in Docker builds withbunx prisma generate --config=prisma.config.ts).
Hard-coded to the single stream statusbus:web with message shape { url, id }.
Exports: xAdd, xAddBulk, xGroupCreate, xReadGroup, xAck, xAckBulk.
Connection: one module-level createClient({ url: process.env.REDIS_URL }).
Shared types used by the queue: MessageType ({url, id}), StreamEntry<T>,
RawRedisMessage.
Bun test runner integration tests (user.test.ts, website.test.ts) that hit a live API
at http://127.0.0.1:3001.
Authoritative schema: packages/store/prisma/schema.prisma.
User
id String @id @default(uuid())
username String @unique
password String ← currently stored in plaintext
websites Website[]
Website
id String @id @default(uuid())
url String
createdAt DateTime @default(now())
user_id String → User.id
ticks WebsiteTick[]
Region
id String @id @default(uuid()) ← seeded as '1' = India, '2' = US
name String
ticks WebsiteTick[]
WebsiteTick
id String @id @default(uuid())
rt_ms Int
status WebsiteStatus (enum: Up | Down | Unknown)
region_id String → Region.id (ON DELETE RESTRICT)
website_id String → Website.id (ON DELETE CASCADE)
createdAt DateTime @default(now())
enum WebsiteStatus { Up Down Unknown }
Migration history:
| Migration | Purpose |
|---|---|
20250722191418_init |
Initial Website, Region, WebsiteTick, WebsiteStatus enum |
20250722204455_web_model_update |
Drop Website.timeAdded, add createdAt |
20250724014954_add_user |
Add User, Website.user_id, WebsiteTick.createdAt |
20250725183020_make_username_unique |
Unique index on User.username |
20250821180444_website_tick_on_delete_cascading |
Ticks cascade-delete with their website |
Storage note: Postgres is currently the single source of truth for both entities and
monitoring results (WebsiteTick). A dedicated time-series database is planned but
not yet implemented — the current docs and code treat WebsiteTick in Postgres as the
store. The Region id is used as the Redis consumer-group name, so seeding the Region
rows (India=1, US=2) is a prerequisite for consumers.
| Layer | Technology |
|---|---|
| Monorepo | Turborepo + Bun workspaces (bun@1.2.18) |
| API | Express 5, TypeScript, JWT, zod |
| Frontend | Next.js 15 (App Router, Turbopack), Tailwind v4, shadcn/ui |
| Workers | Plain Bun/Node processes (producer, consumer) |
| Queue | Redis 7 Streams on Upstash (packages/redisq) |
| Database | PostgreSQL via Neon (Prisma 7, @prisma/adapter-pg) |
| Tests | Bun test runner + axios integration tests |
| Deployment | Docker / docker-compose, kind (local); live: Cloud Run (api/fe) + Neon + Upstash + kubeadm worker mesh (see below) |
| Worker infra | GCP: kubeadm cluster (asia-south1-a, on-demand control plane + 2 spot workers) + standalone spot consumer VM (us-central1-a); bootstrap via VM startup scripts, join token + CA hash in Secret Manager |
- Redis Streams over a plain queue: consumer groups give per-region fan-out, at-least-once delivery semantics and a pending-entry list for recovery, all built in.
- Fan-out to all regions (vs partitioning): every region maintains a complete, independent view of every website, which is the point of global monitoring. Current util is capped by the same number of probe jobs every region performs, but with one or a few consumers per region that scales linearly by region count.
- Central API + Postgres: all writes funnel through the API; workers never touch the DB directly (recently refactored from direct Prisma calls). This keeps the DB ACL simple (only the API connects, via a Neon connection string held in Secret Manager) and gives a natural choke point for auth.
- Region id doubles as the consumer-group name: keeps worker config to two env vars
(
REGION_ID,CONSUMER_ID) at the cost of coupling region identity to the queue machinery. - Probe = plain HTTP GET:
Upiff axios resolves (HTTP 2xx) within 10 s, otherwiseDown. Simple, but it conflates DNS/TLS/status-code failures (see Known Issues indocs/DEPLOYMENT.md).