A production-style backend and data engineering platform for ingesting, validating, processing, securing, and analyzing event data.
Built with FastAPI, PostgreSQL, Redis, Celery, SQLAlchemy, Alembic, Docker Compose, Prometheus, JWT/RBAC, pytest, and GitHub Actions.
Designed to demonstrate production-oriented backend and data-platform engineering beyond a basic CRUD API.
- Versioned REST API with single and batch event ingestion
- PostgreSQL persistence with SQLAlchemy 2.x + Alembic migrations
- Asynchronous event processing with Celery + Redis
- Idempotent writes using
Idempotency-Key - Duplicate-event protection and safe request replay
- Redis-backed analytics caching and cache invalidation
- JWT authentication + role-based access control
- Redis-backed authentication rate limiting
- Event quality assessment and value normalization
- Prometheus metrics, structured JSON logging, and request tracing
- Health and dependency-aware readiness endpoints
- Docker Compose multi-service environment
- 33 passing automated tests
- Ruff linting and GitHub Actions CI
┌─────────────────────┐
│ Client / Producer │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ FastAPI │
│ REST + Validation │
│ JWT + RBAC │
│ Rate Limiting │
│ Idempotency │
└──────┬───────┬──────┘
│ │
┌───────────┘ └────────────┐
▼ ▼
┌──────────────┐ ┌──────────────┐
│ PostgreSQL │ │ Redis │
│ Events/Users │ │ Broker/Cache │
└──────┬───────┘ └──────┬───────┘
│ │
│ ▼
│ ┌──────────────┐
│ │ Celery │
│ │ Worker │
│ └──────┬───────┘
│ │
└───────────────┬────────────────┘
│
▼
┌─────────────────────┐
│ Analytics / Metrics │
│ Redis TTL Cache │
│ Prometheus │
│ Structured Logs │
└─────────────────────┘
Client
│
▼
FastAPI validation
│
▼
Idempotency / duplicate checks
│
▼
PostgreSQL persistence
│
▼
processing_status = pending
│
▼
Redis → Celery worker
│
▼
Quality assessment + normalization
│
▼
processing_status = processed
│
▼
Analytics cache invalidation
Events can also be processed synchronously by setting PROCESS_ASYNC=false, which is useful for local development and deterministic testing.
The FastAPI service exposes versioned endpoints for event ingestion, processing, analytics, authentication, user management, health checks, and observability.
Single-event ingestion supports request validation, duplicate protection, and Redis-backed idempotency.
A successful request is persisted with an initial pending processing state before asynchronous processing.
Events are queued through Redis and processed by a Celery worker.
The processing pipeline performs quality assessment and value normalization before persisting the completed state.
The API also supports batch ingestion and reports accepted and duplicate events independently.
Processed event data is exposed through aggregate analytics endpoints.
The demo below contains 13 ingested and 13 successfully processed events.
Prometheus-compatible metrics expose ingestion, duplicate, processing, and latency telemetry.
| Area | Technologies |
|---|---|
| Backend | Python 3.11, FastAPI, Pydantic, Uvicorn |
| Persistence | PostgreSQL, SQLAlchemy 2.x, Alembic |
| Async Processing | Celery, Redis |
| Caching | Redis TTL cache |
| Security | JWT, password hashing, RBAC, rate limiting |
| Reliability | Idempotency keys, duplicate protection, health/readiness checks |
| Observability | Prometheus, structured JSON logging, request IDs |
| Testing | pytest |
| Infrastructure | Docker, Docker Compose |
| Quality / CI | Ruff, GitHub Actions |
GET /health
GET /ready
GET /metrics
POST /api/v1/auth/register
POST /api/v1/auth/login
GET /api/v1/users/me
GET /api/v1/admin/users
POST /api/v1/events
POST /api/v1/events/batch
GET /api/v1/events
GET /api/v1/events/{event_id}
POST /api/v1/events/{event_id}/process
GET /api/v1/analytics/summary
GET /api/v1/analytics/by-source
GET /api/v1/analytics/by-type
Interactive OpenAPI documentation is available through Swagger UI at /docs.
Single-event ingestion supports an Idempotency-Key header.
POST /api/v1/events
Idempotency-Key: event-request-123On the first request, the event is created and the serialized response is stored in Redis.
If the client retries the request with the same key, the stored response can be replayed without performing another event write.
This models a common production requirement for clients retrying requests after network failures or timeouts.
Each event has a unique event_id, preventing duplicate event persistence independently of request-level idempotency.
Authentication endpoints use Redis-backed fixed-window rate limiting.
When the configured threshold is exceeded, the API returns:
429 Too Many Requests
Retry-After: ...
The API provides:
- user registration and login
- JWT bearer authentication
- secure password hashing
- authenticated current-user access
- typed
userandadminroles - admin-only user-management endpoints
Authorization is implemented through reusable FastAPI dependencies rather than duplicated route-level checks.
The analytics layer exposes summary and grouped event statistics.
GET /api/v1/analytics/summary
GET /api/v1/analytics/by-source
GET /api/v1/analytics/by-type
Summary analytics are cached in Redis with a configurable TTL.
Request
│
▼
Redis lookup
┌─┴──────────┐
│ │
hit miss
│ │
▼ ▼
response PostgreSQL
│
▼
aggregate
│
▼
Redis SETEX
│
▼
response
Relevant event writes invalidate the analytics cache. Duplicate-only batches avoid unnecessary invalidation.
Metrics are exposed through:
GET /metrics
Custom application metrics include:
data_platform_events_ingested_total
data_platform_events_duplicate_total
data_platform_events_processed_total
data_platform_event_ingest_duration_seconds
These provide visibility into ingestion volume, duplicate rejection, processing outcomes, and request latency.
Every request receives a request ID.
Clients may provide X-Request-ID; otherwise the service generates one automatically and returns it in the response.
This allows API responses to be correlated with server-side logs.
Application and request lifecycle events are emitted as structured JSON rather than unstructured console messages.
GET /health
GET /ready
/health reports application liveness.
/ready additionally checks dependencies such as PostgreSQL and Redis, separating process health from actual service readiness.
Run the automated test suite:
pytest -qVerified result:
33 passed
The suite covers:
- event ingestion and retrieval
- batch ingestion
- duplicate rejection
- synchronous and asynchronous processing behavior
- analytics aggregation and Redis caching
- cache hits, misses, and invalidation
- user registration and authentication
- JWT-protected endpoints
- RBAC and admin functionality
- authentication rate limiting
Retry-Afterbehavior- request ID generation and preservation
- idempotency storage and replay
- health and readiness behavior
External Redis/Celery behavior is mocked where appropriate to keep automated tests deterministic.
production-data-platform/
│
├── src/
│ ├── api/
│ │ ├── dependencies.py
│ │ ├── errors.py
│ │ ├── main.py
│ │ ├── middleware.py
│ │ ├── routes_admin.py
│ │ ├── routes_analytics.py
│ │ ├── routes_auth.py
│ │ ├── routes_events.py
│ │ └── routes_users.py
│ │
│ ├── core/
│ │ ├── config.py
│ │ ├── idempotency.py
│ │ ├── logging.py
│ │ ├── metrics.py
│ │ ├── rate_limit.py
│ │ ├── redis.py
│ │ ├── request_context.py
│ │ └── security.py
│ │
│ ├── db/
│ ├── models/
│ ├── repositories/
│ ├── schemas/
│ ├── services/
│ └── workers/
│
├── alembic/
├── tests/
├── docs/
│ └── screenshots/
├── .github/
│ └── workflows/
├── docker-compose.yml
├── Dockerfile
├── pyproject.toml
└── requirements.txt
The repository separates API transport, persistence, domain services, infrastructure concerns, and background processing rather than concentrating application logic in route handlers.
The easiest way to run the complete stack is Docker Compose.
docker compose up --buildThis starts:
api FastAPI application
db PostgreSQL
redis Redis broker/cache/state store
worker Celery worker
Once running:
API http://localhost:8000
Swagger http://localhost:8000/docs
Metrics http://localhost:8000/metrics
Stop the stack with:
docker compose downCreate and activate a virtual environment:
python -m venv .venv
source .venv/bin/activateInstall dependencies:
pip install -r requirements.txtCopy the environment template:
cp .env.example .envApply database migrations:
alembic upgrade headRun the API:
uvicorn src.api.main:app --reloadRun linting:
ruff check .Run tests:
pytest -qGitHub Actions runs the quality gate on pushes and pull requests to main.
Install dependencies
│
▼
Ruff
│
▼
pytest
This prevents linting and test regressions from being merged unnoticed.
PostgreSQL provides durable relational persistence, constraints, transactions, and analytical queries.
Redis acts as the Celery broker, analytics cache, rate-limit state store, idempotency store, and readiness dependency.
Celery moves event processing outside the HTTP request path when asynchronous processing is enabled.
Repository and service layers separate database access and domain processing from HTTP transport concerns.
Alembic provides explicit, version-controlled database schema migrations.
Idempotency keys protect write operations against accidental re-execution during client retries.
Explicit cache invalidation keeps analytics consistent after relevant writes instead of relying only on TTL expiration.
Prometheus metrics + structured logs + request IDs provide operational visibility and request correlation.
Docker Compose provides a reproducible local multi-service environment.
This repository is intentionally broader than a CRUD API.
It demonstrates practical backend and data-platform concerns including:
API design · relational persistence · schema migrations · asynchronous processing · caching · idempotency · authentication · authorization · rate limiting · observability · testing · CI/CD · multi-service containerization
The project is designed as a portfolio demonstration for Software Engineer, Backend Engineer, Data Engineer, and platform-oriented engineering roles.
This project is licensed under the MIT License.





