Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ See respective README files in sub-directories for details.
- [LoR](lor/README.md): Using Dataflow, Cloud Run and Spanner to explore Lord of the Rings characters.
- [Network Digital Twin](telco-and-csp/README.md): Advanced usage of graph to Spanner Graph to model, visualize, and query a complex telecommunications network.
- [Transit Fraud Detector](TransitFraud/README.md): Advanced usage of graph capabilities to detect fraud.

- [Orders Streaming Analytics](orders-streaming-analytics/README.md): Dual-path streaming with Spanner: Precomputed Lakehouse aggregates and Ad-Hoc analytics via Data Boost.

## Notebooks
Some of these notebooks are hosted in external Google Cloud repositories.
Expand Down
4 changes: 4 additions & 0 deletions orders-streaming-analytics/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
src/java-tpch-stream-generator/target
output
checkpoints
.env
27 changes: 27 additions & 0 deletions orders-streaming-analytics/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Spanner DataBoost Streaming Demo

This demo implements a dual-path streaming architecture ingesting orders (TCP-H dataset) where each branch addresses a distinct analytical goal.

The left branch (Lakehouse Aggregates path) handles predictable, day-to-day reporting by precomputing cumulative aggregates, such as daily order totals, and saving them directly into Spanner aggregate tables only when relevant data changes.

The right branch (Direct Spanner path) continuously maintains the raw orders table to support ad-hoc queries not covered by precomputations. By leveraging Spanner Data Boost alongside its columnar engine, these ad-hoc queries run efficiently without impacting the primary Spanner instance, thereby eliminating the need to over-provision it.

![DataBoost Streaming Demo](SpannerStreamingLakehouse-Dataflow.drawio.png)

### System requirements

- gcloud cli
- uv https://docs.astral.sh/uv/#installation (for dependencies management)
- jq https://jqlang.org/ (for parsing gcloud json response reliably, used by the bash scripts during the notebook)


## Setup

The entire demo architecture notes, GCP infrastructure provisioning, Spark job sources, producer execution, and dashboard/ad-hoc query code is in [`demo.ipynb`](demo.ipynb).

```bash
uv sync
uv run jupyter lab
```

Open `demo.ipynb` and run the cells top to bottom.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading