From 7cfd7b7ac9975fddbbce7e14174ead318f173ccb Mon Sep 17 00:00:00 2001 From: Jimmy Angelakos Date: Wed, 30 Sep 2026 20:15:57 +0100 Subject: [PATCH 1/3] docs: title-case headings and apply the pgEdge docs style pass --- DUCKDB_1.5_PATCHED.md | 14 +- README.md | 21 +-- docs/architecture.md | 318 ++++++++++++++++----------------- docs/architecture_decoupled.md | 85 +++++---- docs/architecture_tiered.md | 204 ++++++++++----------- docs/architecture_vectors.md | 34 ++-- docs/compaction.md | 24 +-- docs/formal/README.md | 106 +++++------ docs/index.md | 12 +- docs/installation.md | 32 ++-- docs/object_store.md | 61 +++---- docs/usage.md | 135 +++++++------- docs/usage_vectors.md | 62 ++++--- docs/walkthrough.md | 8 +- docs/walkthrough_demos.md | 145 +++++++++------ 15 files changed, 645 insertions(+), 616 deletions(-) diff --git a/DUCKDB_1.5_PATCHED.md b/DUCKDB_1.5_PATCHED.md index 520bbee..bb61c3c 100644 --- a/DUCKDB_1.5_PATCHED.md +++ b/DUCKDB_1.5_PATCHED.md @@ -121,12 +121,14 @@ inside `DoTableUpdates` (PG `PRE_COMMIT`, while the ticket is held): 3 then land the write on the live head. An expired start snapshot keeps the error. -**Formally verified** before the code (the project rule): -`docs/formal/Bakery.tla` models the async ordering; `Bakery_async.cfg` -(patched) holds `NoLakekeeperConflict`, `Bakery_race.cfg` (async **without** -the patch) violates it — the standing proof the patch is mandatory for async. -**Validated** over Azure ADLS: journey 6b (4 concurrent mixed-tier writers → -8/8, 0 loss) and 9b (8 concurrent cold writers → 8/8). +### Formally verified + +Before the code (the project rule): `docs/formal/Bakery.tla` models the async +ordering; `Bakery_async.cfg` (patched) holds `NoLakekeeperConflict`, +`Bakery_race.cfg` (async **without** the patch) violates it — the standing +proof the patch is mandatory for async. **Validated** over Azure ADLS: journey +6b (4 concurrent mixed-tier writers → 8/8, 0 loss) and 9b (8 concurrent cold +writers → 8/8). ## 3. Strict-reader interop (two patches + one upstream fix) diff --git a/README.md b/README.md index 0f640ad..3c8d8e1 100644 --- a/README.md +++ b/README.md @@ -70,13 +70,14 @@ modes in action with copy-pasteable commands. ColdFront is open source under the PostgreSQL License and runs on stock PostgreSQL 16, 17, and 18. The full build workflow lives in the -**[Installation guide](docs/installation.md)**: build the thin coldfront layer +**[Installation guide](docs/installation.md)**: build the thin ColdFront layer on top of the published DuckDB 1.5.x base image (or build the base yourself), or install bare-metal. Then continue with the Quickstart below. -**Setting up on cloud S3?** Once the image is built, the -**[S3 setup guide](docs/object_store.md)** takes you from an empty bucket to a -working cold tier end-to-end. +### Setting Up on Cloud S3? + +Once the image is built, the **[S3 setup guide](docs/object_store.md)** takes +you from an empty bucket to a working cold tier end-to-end. ## Quickstart @@ -106,18 +107,18 @@ SELECT count(*) FROM events; A table that already exists in the Iceberg catalog is adopted rather than created: `coldfront.adopt_iceberg_table()` reads its schema from the catalog and gives it the same wrapper view and registry row, read-only unless writes -are asked for. `coldfront.release_iceberg_table()` hands it back with the -Iceberg table untouched. See -[Adopting a table that already exists in the catalog](docs/usage.md#adopting-a-table-that-already-exists-in-the-catalog). +are asked for. `coldfront.release_iceberg_table()` hands the table back with +the Iceberg table untouched. See +[Adopting a Table That Already Exists in the Catalog](docs/usage.md#adopting-a-table-that-already-exists-in-the-catalog). To remove a table again, `coldfront.drop_iceberg_table()` unregisters it and drops the Iceberg table, deleting the stored objects only when asked to. See -[Dropping an Iceberg table](docs/usage.md#dropping-an-iceberg-table-both-modes). +[Dropping an Iceberg Table](docs/usage.md#dropping-an-iceberg-table-both-modes). For compliance environments that cannot store an object-store credential, `coldfront.set_storage_secret_vended()` runs with no credential in the database: Lakekeeper issues short-lived per-table credentials at access time. -See [Vended credentials](docs/usage.md#vended-credentials). +See [Vended Credentials](docs/usage.md#vended-credentials). ## Documentation @@ -135,7 +136,7 @@ The following table lists the ColdFront guides and what each one covers: | **[Architecture: decoupled](docs/architecture_decoupled.md)** | Decoupled (iceberg-only) deep dive | | **[Architecture: vectors](docs/architecture_vectors.md)** | Vector storage internals - type mapping, routing state, cluster assignment, layout | -## Least-privilege application roles +## Least-Privilege Application Roles Application roles need no superuser and no server-file access, yet they read and write the cold tier through the same transparent view. Onboarding an diff --git a/docs/architecture.md b/docs/architecture.md index 0a7f065..d94e6c6 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -7,32 +7,16 @@ rewrites reads and writes to hit the right tier, and **all Iceberg I/O goes through `pg_duckdb` running in-process inside PostgreSQL** - no external query engine, no Go Iceberg libraries. -## Contents - -This document is organized into the following sections: - -- [Operating modes and topologies](#operating-modes-and-topologies) - the three - axes + read target (primary / standby) -- [System Overview](#system-overview) - the moving parts -- [Core Mechanics: pg_duckdb](#core-mechanics-pg_duckdb) - how Iceberg I/O - happens -- [Application Interface](#application-interface) - the shared rewrite hook -- [Concurrency and pgEdge Spock Deployments](#concurrency-and-pgedge-spock-deployments) - - bakery, cold-write strategy, DDL, OID-vs-name -- [Known Limitations](#known-limitations) - cross-cutting -- [Infrastructure (Docker)](#infrastructure-docker) -- [Upstream Requests](#upstream-requests) - open asks to pg_duckdb / - duckdb-iceberg - For mode-specific design, see [architecture_tiered.md](architecture_tiered.md) (hot PG + cold Iceberg) · [architecture_decoupled.md](architecture_decoupled.md) (all-Iceberg) · [architecture_vectors.md](architecture_vectors.md) (vector storage). -## Operating modes and topologies +## Operating Modes and Topologies -Three independent axes describe any ColdFront deployment, selected as shown -below. They compose freely - e.g. tiered + mesh + permissive writes: +Three independent axes describe any ColdFront deployment; they compose freely - +e.g. tiered + mesh + permissive writes. The following table shows how each axis +is selected: | Axis | Values | Selected by | |---|---|---| @@ -44,14 +28,14 @@ Both storage modes coexist in one database and share **one** code path: the transparent view and read rewriter, the INSERT/UPDATE/DELETE hook (`emit_cold` / `emit_hot` / `emit_dual` in [`extension/coldfront/src/coldfront.c`](https://github.com/pgEdge/ColdFront/blob/main/extension/coldfront/src/coldfront.c)), -and the `_exec_iceberg_with_claim` write chokepoint. Decoupled mode simply -always classifies as `TIER_COLD` and never reaches `emit_hot`; vanilla and mesh -differ only in how that chokepoint serializes cold writes. This document covers -the shared mechanics and the tiered path; see +and the `_exec_iceberg_with_claim` write chokepoint. Decoupled mode always +classifies as `TIER_COLD` and never reaches `emit_hot`; vanilla and mesh differ +only in how that chokepoint serializes cold writes. This document covers the +shared mechanics and the tiered path; see [architecture_decoupled.md](architecture_decoupled.md) for the decoupled mode's ACID model and distributed scaling story. -### Read target: primary or physical standby +### Read Target: Primary or Physical Standby Orthogonal to the three axes above, any ColdFront node - vanilla or a mesh member - can have one or more **physical (streaming) standbys that serve @@ -95,7 +79,7 @@ below: └──────────────┬───────────────────────────────────────────┘ │ ┌──────────────▼───────────────────────────────────────────┐ -│ Lakekeeper — Iceberg REST catalog (own dedicated Postgres)│ +│ Lakekeeper - Iceberg REST catalog (own dedicated Postgres)│ │ Manages Iceberg metadata, snapshots, commit concurrency │ └──────────────┬───────────────────────────────────────────┘ │ @@ -110,7 +94,7 @@ The following table describes each component, its role, and its license: | Component | Role | License | |-----------|------|---------| | PostgreSQL 16+ | Heap storage; range partitioning for the tiered hot tier. Works uniformly on PG 16, 17, and 18 - the cold-tier secret is a DuckDB persistent secret loaded at instance init, with no version-gated mechanism. | PostgreSQL | -| pg_duckdb | DuckDB in-process. Iceberg read + write. Analytics. pg_duckdb 1.5.4 (PR #1025). The `duckdb-iceberg` carries the bakery-aware commit-refresh patch (async parquet overlap, no 409); see [Cold-write strategy](#cold-write-strategy-stock-vs-patched-duckdb-iceberg). | MIT | +| pg_duckdb | DuckDB in-process. Iceberg read + write. Analytics. pg_duckdb 1.5.4 (PR #1025). The `duckdb-iceberg` carries the bakery-aware commit-refresh patch (async parquet overlap, no 409); see [Cold-Write Strategy](#cold-write-strategy-stock-vs-patched-duckdb-iceberg). | MIT | | coldfront | PGXS C extension. `post_parse_analyze_hook` rewrites INSERT/UPDATE/DELETE on registered views to the correct tier and, on a SELECT DuckDB will run, the spellings DuckDB lacks (`date_bin`, `::jsonb`, the JSON builders); `planner_hook` folds bound parameters into such a read; `ProcessUtility_hook` handles DDL; the hook lazily ATTACHes the Iceberg catalog on the first query touching a tiered view. | PostgreSQL | | Lakekeeper | Iceberg REST catalog. Single Rust binary. | Apache 2.0 | | S3-compatible store | Any: SeaweedFS, MinIO, GCS, Azure Blob, etc. | Varies | @@ -118,7 +102,7 @@ The following table describes each component, its role, and its license: How rows move through this depends on the storage mode: the tiered hot heap + archiver + `UNION ALL` data-flow is in -[architecture_tiered.md → Data flow](architecture_tiered.md#data-flow); the +[architecture_tiered.md → Data Flow](architecture_tiered.md#data-flow); the all-Iceberg flow is in [architecture_decoupled.md](architecture_decoupled.md). ## Core Mechanics: pg_duckdb @@ -127,7 +111,7 @@ All Iceberg I/O goes through SQL executed against PostgreSQL. There are no Go DuckDB/Iceberg/Arrow libraries. DuckDB Iceberg writes require a REST catalog - Lakekeeper fills this role. -### Session setup +### Session Setup The cold-tier S3 secret is set once per cluster. A single call records the credentials and materializes a DuckDB persistent secret: @@ -173,14 +157,14 @@ The Iceberg catalog ATTACH is **lazy**: the coldfront C extension hook issues `coldfront.warehouse` and `coldfront.lakekeeper_endpoint` GUCs - on the **first query that touches a tiered view** (read or write), per DuckDB cached connection. There is no arming step and no per-session boilerplate: both reads -(`iceberg_scan`) and writes (`duckdb.raw_query`) just work on a fresh psql -session. Until a tiered view is touched no ATTACH is attempted, so a -pre-bootstrap connection is never blocked by a missing warehouse. +(`iceberg_scan`) and writes (`duckdb.raw_query`) work on a fresh psql session. +Until a tiered view is touched no ATTACH is attempted, so a pre-bootstrap +connection is never blocked by a missing warehouse. -### Non-superuser app roles (least privilege) +### Non-Superuser App Roles (Least Privilege) pg_duckdb force-disables DuckDB's `LocalFileSystem` for non-superusers (see -[Upstream Requests](#pg_duckdb-non-superuser-localfilesystem-blocks-side-loaded-extensions)), +[Upstream requests](#pg_duckdb-non-superuser-localfilesystem-blocks-side-loaded-extensions)), which would block the side-loaded iceberg/postgres DuckDB extensions from loading on `ATTACH`. So `coldfront.ensure_attached()` / `ensure_pg_attached()` are `SECURITY DEFINER` with a pinned `search_path`: the extension load + @@ -217,7 +201,7 @@ still violates `NoLakekeeperConflict`). See README "Security"; asserted by the journey's `story_app_privilege`, `ci/ops.sh` check 3, and the `privilege_model` pg_regress test. -### Temp table bridge: PG → Iceberg +### Temp Table Bridge: PG → Iceberg `duckdb.raw_query()` cannot see PG tables directly. The bridge is a DuckDB temp table: @@ -230,7 +214,7 @@ SELECT duckdb.raw_query($$INSERT INTO ice.public.events DROP TABLE duck_stage; ``` -### Cold-side column references +### Cold-Side Column References `iceberg_scan()` requires `r['col']::type` syntax: @@ -256,7 +240,7 @@ the statement the view is named (a CTE, a sub-select, a set-operation branch): `date_bin`, the `::jsonb` cast, `jsonb_array_length` and the JSON builders (`jsonb_build_object`, `jsonb_agg` and their `json_` twins) are rewritten into spellings both engines accept (see -[usage.md → Supported column types](usage.md#supported-column-types)). A +[usage.md → Supported Column Types](usage.md#supported-column-types)). A `planner_hook` folds bound parameters into such a read before pg_duckdb plans it when a parameter sits where DuckDB cannot type a placeholder (a direct argument of a pg_duckdb function, any argument of a table function); the plan @@ -265,7 +249,7 @@ priced above any custom plan, so under the plan cache's cost-based selection the read is planned from its values on every execution. `plan_cache_mode = force_generic_plan` bypasses that selection and picks the value-less generic plan, which fails with `only works with DuckDB execution` -(see [usage.md → Supported column types](usage.md#supported-column-types)). +(see [usage.md → Supported Column Types](usage.md#supported-column-types)). The following table maps each operation to its interface and routing path: @@ -289,23 +273,22 @@ How the hook splits a write is mode-specific: - **Decoupled** always classifies `TIER_COLD`: every write is a single-tier Iceberg write. See [architecture_decoupled.md](architecture_decoupled.md). -### Cold-tier DML from inside plpgsql (functions, DO blocks, triggers) +### Cold-Tier DML from Inside PL/pgSQL (Functions, DO Blocks, Triggers) Cold-tier `INSERT`/`UPDATE`/`DELETE` work as top-level statements *and* from inside a plpgsql function / `DO` block / trigger, via two mechanisms: -1. **Parameters are emitted as a runtime `format(