diff --git a/.claude/agents/golang-expert.md b/.claude/agents/golang-expert.md index 33cacae..71f029d 100644 --- a/.claude/agents/golang-expert.md +++ b/.claude/agents/golang-expert.md @@ -26,7 +26,8 @@ You are a Go development specialist for ColdFront. ## Testing Approach - Table-driven tests preferred -- Integration tests with a real database (the `ci/journey.sh` harness), not mocks +- Integration tests with a real database (the `ci/journey.sh` harness), not + mocks - Hand-written mocks defined locally in test files when a unit needs one - Test files next to the code they test - Use `t.Helper()` in test utilities diff --git a/.claude/agents/postgres-expert.md b/.claude/agents/postgres-expert.md index e2c338b..2368dc8 100644 --- a/.claude/agents/postgres-expert.md +++ b/.claude/agents/postgres-expert.md @@ -23,11 +23,12 @@ You are a PostgreSQL specialist for ColdFront. - Index naming: `idx_{table}_{column}` - Constraint naming: `chk_`, `fk_`, `{table}_{cols}_unique` - `COMMENT ON` for schema objects -- Parameterized queries only; interpolate only sanitized identifiers, never values +- Parameterized queries only; interpolate only sanitized identifiers, never + values - pgerrcode for error classification - Idempotent migrations (`IF NOT EXISTS`) -- NEVER write plpgsql `EXCEPTION`/`SAVEPOINT` in the cold path — pg_duckdb rejects - subtransactions; use precondition checks instead +- NEVER write plpgsql `EXCEPTION`/`SAVEPOINT` in the cold path — pg_duckdb + rejects subtransactions; use precondition checks instead ## Replication Safety diff --git a/.claude/agents/security-auditor.md b/.claude/agents/security-auditor.md index bddede5..48e01b2 100644 --- a/.claude/agents/security-auditor.md +++ b/.claude/agents/security-auditor.md @@ -17,11 +17,14 @@ You are a security specialist for ColdFront. ## Standards -- No hardcoded secrets (use config / environment variables / DuckDB persistent secrets) -- No real hostnames, buckets, accounts, keys, or paths in committed code, tests, docs, or commit messages +- No hardcoded secrets (use config / environment variables / DuckDB persistent + secrets) +- No real hostnames, buckets, accounts, keys, or paths in committed code, + tests, docs, or commit messages - Parameterized queries; never concatenate values into SQL - Input validation at all boundaries - gitleaks must pass (no committed secrets) - gosec findings must be addressed - Principle of least privilege for database connections and app roles - (`grant_app_access`, the SECURITY DEFINER attach helpers, PGC_SUSET config GUCs) + (`grant_app_access`, the SECURITY DEFINER attach helpers, PGC_SUSET config + GUCs) diff --git a/.github/PULL_REQUEST_TEMPLATE.md b/.github/PULL_REQUEST_TEMPLATE.md index 56479da..3562339 100644 --- a/.github/PULL_REQUEST_TEMPLATE.md +++ b/.github/PULL_REQUEST_TEMPLATE.md @@ -4,10 +4,12 @@ ## Checklist -- [ ] `./run-ci-local.sh` passes (gofmt, golangci-lint, build, pg_regress, journey) +- [ ] `./run-ci-local.sh` passes (gofmt, golangci-lint, build, pg_regress, + journey) - [ ] Tests added/updated (test-first) - [ ] Docs updated where applicable (README / USAGE / INSTALL / ARCHITECTURE) -- [ ] Bakery / mesh / distributed changes: TLA+ model updated and TLC re-checked (`docs/formal/`) +- [ ] Bakery / mesh / distributed changes: TLA+ model updated and TLC + re-checked (`docs/formal/`) - [ ] Commit messages are short and imperative ## Related issues diff --git a/DUCKDB_1.5_PATCHED.md b/DUCKDB_1.5_PATCHED.md index 520bbee..bb61c3c 100644 --- a/DUCKDB_1.5_PATCHED.md +++ b/DUCKDB_1.5_PATCHED.md @@ -121,12 +121,14 @@ inside `DoTableUpdates` (PG `PRE_COMMIT`, while the ticket is held): 3 then land the write on the live head. An expired start snapshot keeps the error. -**Formally verified** before the code (the project rule): -`docs/formal/Bakery.tla` models the async ordering; `Bakery_async.cfg` -(patched) holds `NoLakekeeperConflict`, `Bakery_race.cfg` (async **without** -the patch) violates it — the standing proof the patch is mandatory for async. -**Validated** over Azure ADLS: journey 6b (4 concurrent mixed-tier writers → -8/8, 0 loss) and 9b (8 concurrent cold writers → 8/8). +### Formally verified + +Before the code (the project rule): `docs/formal/Bakery.tla` models the async +ordering; `Bakery_async.cfg` (patched) holds `NoLakekeeperConflict`, +`Bakery_race.cfg` (async **without** the patch) violates it — the standing +proof the patch is mandatory for async. **Validated** over Azure ADLS: journey +6b (4 concurrent mixed-tier writers → 8/8, 0 loss) and 9b (8 concurrent cold +writers → 8/8). ## 3. Strict-reader interop (two patches + one upstream fix) diff --git a/LICENSE.md b/LICENSE.md index bd01495..dc796f4 100644 --- a/LICENSE.md +++ b/LICENSE.md @@ -2,12 +2,24 @@ The PostgreSQL License Portions copyright (c) 2026, pgEdge, Inc. -Permission to use, copy, modify, and distribute this software and its documentation for any purpose, without fee, and without a written agreement is hereby granted, provided that the above copyright notice and this paragraph and the following two paragraphs appear in all copies. +Permission to use, copy, modify, and distribute this software and its +documentation for any purpose, without fee, and without a written agreement is +hereby granted, provided that the above copyright notice and this paragraph and +the following two paragraphs appear in all copies. -IN NO EVENT SHALL pgEdge, Inc. BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF pgEdge, Inc. HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. +IN NO EVENT SHALL pgEdge, Inc. BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, +SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING +OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF pgEdge, Inc. HAS +BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. -pgEdge, Inc. SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE SOFTWARE PROVIDED HEREUNDER IS ON AN "AS IS" BASIS, AND pgEdge, Inc. HAS NO OBLIGATIONS TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS. +pgEdge, Inc. SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED +TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR +PURPOSE. THE SOFTWARE PROVIDED HEREUNDER IS ON AN "AS IS" BASIS, AND pgEdge, +Inc. HAS NO OBLIGATIONS TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, +OR MODIFICATIONS. --- -This software redistributes third-party components under their own licenses (DuckDB, pg_duckdb, duckdb-iceberg, and others). Those copyright and license notices are reproduced in [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). +This software redistributes third-party components under their own licenses +(DuckDB, pg_duckdb, duckdb-iceberg, and others). Those copyright and license +notices are reproduced in [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). diff --git a/README.md b/README.md index 0f640ad..4aca500 100644 --- a/README.md +++ b/README.md @@ -1,18 +1,40 @@ # pgEdge ColdFront -> [!WARNING] -> ColdFront is beta software under active development. Do not use it in -> production. Interfaces, on-disk formats, and behavior may change without -> notice, and data loss is possible. - [![CI](https://github.com/pgEdge/ColdFront/actions/workflows/ci.yml/badge.svg)](https://github.com/pgEdge/ColdFront/actions/workflows/ci.yml) +## Table of Contents + +The ColdFront documentation consists of the following guides: + +- [Introduction](docs/index.md) +- Getting Started + - [Running the Walkthrough](docs/walkthrough.md) + - [Exploring the Walkthrough Demos](docs/walkthrough_demos.md) + - [Building ColdFront from Source](docs/installation.md) + - [Setting Up an Object Store](docs/object_store.md) +- Architecture + - [Architecture Overview](docs/architecture.md) + - [Tiered Mode](docs/architecture_tiered.md) + - [Decoupled Mode](docs/architecture_decoupled.md) + - [Vector Storage](docs/architecture_vectors.md) +- [Using ColdFront](docs/usage.md) +- [Storing and Searching Embeddings](docs/usage_vectors.md) +- [Compacting the Cold Tier](docs/compaction.md) +- Developer Resources + - [Verifying the Bakery Protocol](docs/formal/README.md) +- [Release Notes](docs/changelog.md) + ColdFront keeps tables in PostgreSQL and cold data in Apache Iceberg (Parquet on S3-compatible, Azure, or GCS storage), and the cold tier is both readable and writable through the same SQL with no application changes. The application queries every table as an ordinary PostgreSQL relation, and both operating modes present the same standard SQL surface. +> [!WARNING] +> ColdFront is beta software under active development. Do not use it in +> production. Interfaces, on-disk formats, and behavior may change without +> notice, and data loss is possible. + ColdFront provides two operating modes: - Tiered mode keeps recent data in native PostgreSQL partitions and archives @@ -65,18 +87,35 @@ tier, so the application sees one relation: ## Installation -New here? Run the [guided walkthrough](docs/walkthrough.md) to see all three -modes in action with copy-pasteable commands. +If you are new to ColdFront, run the [guided walkthrough](docs/walkthrough.md) +to see all three modes in action with copy-pasteable commands. -ColdFront is open source under the PostgreSQL License and runs on stock -PostgreSQL 16, 17, and 18. The full build workflow lives in the -**[Installation guide](docs/installation.md)**: build the thin coldfront layer +ColdFront is open source under the PostgreSQL License and runs on PostgreSQL +16, 17, and 18: stock PostgreSQL for a single node, and pgEdge's PostgreSQL +build with Spock, which the Docker image is based on, for distributed mode. The +full build workflow lives in the +**[Installation guide](docs/installation.md)**: build the thin ColdFront layer on top of the published DuckDB 1.5.x base image (or build the base yourself), or install bare-metal. Then continue with the Quickstart below. -**Setting up on cloud S3?** Once the image is built, the -**[S3 setup guide](docs/object_store.md)** takes you from an empty bucket to a -working cold tier end-to-end. +### Setting Up on Cloud S3? + +Once the image is built, the **[S3 setup guide](docs/object_store.md)** takes +you from an empty bucket to a working cold tier end-to-end. + +## Configuration + +ColdFront reads its settings from two places. The server settings live in +`postgresql.conf`: `shared_preload_libraries = 'pg_duckdb,coldfront'`, +`coldfront.warehouse` and `coldfront.lakekeeper_endpoint`, plus +`snowflake.node` and `coldfront.dblink_self` on every node of a mesh. The +Docker image writes them on first start. The archiver, partitioner, and +compactor read a deployment YAML, modeled on +[config.example.yaml](config.example.yaml), for the database DSN, the Iceberg +catalog, and the cold-store credentials, and each table's lifecycle lives in +`coldfront.partition_config`. For every setting, see +[Using ColdFront → One-Time Setup](docs/usage.md#one-time-setup) and +[Tuning Knobs](docs/usage.md#tuning-knobs). ## Quickstart @@ -84,7 +123,7 @@ Build the image (see the [Installation](docs/installation.md) guide) and bring up the stack: ```bash -docker compose up -d --build +docker compose --profile local-store up -d --build ``` Bootstrap Lakekeeper and create a warehouse (see the one-time setup in the @@ -106,18 +145,26 @@ SELECT count(*) FROM events; A table that already exists in the Iceberg catalog is adopted rather than created: `coldfront.adopt_iceberg_table()` reads its schema from the catalog and gives it the same wrapper view and registry row, read-only unless writes -are asked for. `coldfront.release_iceberg_table()` hands it back with the -Iceberg table untouched. See -[Adopting a table that already exists in the catalog](docs/usage.md#adopting-a-table-that-already-exists-in-the-catalog). +are asked for. `coldfront.release_iceberg_table()` hands the table back with +the Iceberg table untouched. See +[Adopting a Table That Already Exists in the Catalog](docs/usage.md#adopting-a-table-that-already-exists-in-the-catalog). To remove a table again, `coldfront.drop_iceberg_table()` unregisters it and drops the Iceberg table, deleting the stored objects only when asked to. See -[Dropping an Iceberg table](docs/usage.md#dropping-an-iceberg-table-both-modes). +[Dropping an Iceberg Table](docs/usage.md#dropping-an-iceberg-table-both-modes). For compliance environments that cannot store an object-store credential, `coldfront.set_storage_secret_vended()` runs with no credential in the database: Lakekeeper issues short-lived per-table credentials at access time. -See [Vended credentials](docs/usage.md#vended-credentials). +See [Vended Credentials](docs/usage.md#vended-credentials). + +## Using ColdFront + +The Quickstart covers a decoupled table from start to finish. The +[Using ColdFront](docs/usage.md) guide covers both modes in depth, the +standalone partition manager and its CLI, the storage backends, and the +distributed setup; the [walkthrough demos](docs/walkthrough_demos.md) run each +mode on a sample table. ## Documentation @@ -125,17 +172,19 @@ The following table lists the ColdFront guides and what each one covers: | Doc | Contents | |---|---| -| **[Embeddings](docs/usage_vectors.md)** | Storing and searching embeddings with the pgvector interface | -| **[Usage](docs/usage.md)** | Day-to-day use - both modes plus the standalone partition manager, one-time setup, reading/writing, supported types, the partition CLI, storage backends, distributed (mesh) setup, tuning | -| **[Installation](docs/installation.md)** | Build from source (Docker or bare-metal); Testing & CI | -| **[Object store setup](docs/object_store.md)** | Get ColdFront running on cloud S3 (virtual-hosted), end-to-end | -| **[Compaction](docs/compaction.md)** | Cold-tier table maintenance - compaction, snapshot expiry, orphan-file removal | -| **[Architecture](docs/architecture.md)** | Shared architecture and core mechanics | -| **[Architecture: tiered](docs/architecture_tiered.md)** | Tiered (hot PG + cold Iceberg) deep dive | -| **[Architecture: decoupled](docs/architecture_decoupled.md)** | Decoupled (iceberg-only) deep dive | -| **[Architecture: vectors](docs/architecture_vectors.md)** | Vector storage internals - type mapping, routing state, cluster assignment, layout | - -## Least-privilege application roles +| [Walkthrough](docs/walkthrough.md) | Sets up the demo stack and runs ColdFront hands-on. | +| [Walkthrough demos](docs/walkthrough_demos.md) | Walks through the tiered, decoupled, partitioner, and distributed demos. | +| [Embeddings](docs/usage_vectors.md) | Covers storing and searching embeddings with the pgvector interface. | +| [Usage](docs/usage.md) | Covers day-to-day use: both modes plus the standalone partition manager, one-time setup, reading and writing, supported types, the partition CLI, storage backends, distributed (mesh) setup, and tuning. | +| [Installation](docs/installation.md) | Covers building from source (Docker or bare-metal), and testing and CI. | +| [Object store setup](docs/object_store.md) | Gets ColdFront running on cloud S3 (virtual-hosted), end to end. | +| [Compaction](docs/compaction.md) | Covers cold-tier table maintenance: compaction, snapshot expiry, and orphan-file removal. | +| [Architecture](docs/architecture.md) | Describes the shared architecture and core mechanics. | +| [Architecture: tiered](docs/architecture_tiered.md) | Describes tiered mode (hot PG plus cold Iceberg) in depth. | +| [Architecture: decoupled](docs/architecture_decoupled.md) | Describes decoupled (iceberg-only) mode in depth. | +| [Architecture: vectors](docs/architecture_vectors.md) | Describes vector storage internals: type mapping, routing state, cluster assignment, and layout. | + +## Least-Privilege Application Roles Application roles need no superuser and no server-file access, yet they read and write the cold tier through the same transparent view. Onboarding an @@ -146,8 +195,10 @@ SELECT coldfront.grant_app_access('alice'); ``` grant_app_access grants only the minimum the cold path needs: membership in -duckdb.postgres_role, schema USAGE, SELECT on the registry, and DML on every -registered view and the hot table and sequences behind it. Those objects are +duckdb.postgres_role, SET on the duckdb.unsafe_allow_execution_inside_functions +parameter, schema USAGE, SELECT on the registry and the watermark table, DML on +the dual-write anchor table, DML on every registered view and the hot table +behind it, and USAGE and SELECT on the hot table's sequences. Those objects are derived from the registry, not hardcoded, and the call also grants EXECUTE on a fixed allow-list of runtime cold-path functions. The call is idempotent and is not executable by PUBLIC, so an application role can never self-grant. The role @@ -155,10 +206,16 @@ is never granted pg_read_server_files or pg_write_server_files, so it has no host-file access. CREATE ROLE and GRANT both replicate over Spock, so you onboard a role once on any node and it propagates across the mesh. -For how the non-superuser path works under the hood - the `SECURITY DEFINER` -attach helpers, the `PGC_SUSET` / `GUC_SUPERUSER_ONLY` config hardening, the -turnkey `duckdb.postgres_role` default, and how least privilege holds across a -Spock mesh - see +The Docker image sets `duckdb.postgres_role` to `coldfront_duckdb` and creates +that role when it initializes a new data directory. To name a different role, +set the container's `COLDFRONT_DUCKDB_ROLE` environment variable before that +first start. An empty value keeps pg_duckdb's default, under which only +superusers run DuckDB, so grant_app_access then fails with an error. + +For how the non-superuser path works - the `SECURITY DEFINER` attach helpers, +the `PGC_SUSET` / `GUC_SUPERUSER_ONLY` config hardening, the +`duckdb.postgres_role` default that the image sets up, and how least privilege +holds across a Spock mesh - see [Architecture: non-superuser app roles](docs/architecture.md#non-superuser-app-roles-least-privilege). ## Caveats @@ -197,17 +254,19 @@ pgedge-coldfront/ │ ├── topo/ ← vanilla.sh (1 node) · mesh.sh (3-node Spock) │ └── runbooks/ ← failover-patroni.md (failover delegated to Patroni) ├── docker/ -│ ├── Dockerfile.duckdb15-base ← DuckDB 1.5.x base (pg_duckdb 1.5.4 + patched iceberg) +│ ├── Dockerfile.duckdb15-base ← DuckDB 1.5.x base (pg_duckdb on DuckDB 1.5.4 + patched iceberg) │ ├── Dockerfile.duckdb15 ← thin coldfront app layer (ARG PG_MAJOR=16|17|18) │ ├── iceberg-*.patch ← duckdb-iceberg patches (bakery commit-refresh + strict-reader interop) │ ├── iceberg-azure-extension-config-v15.cmake ← Azure ADLS extension build config │ ├── entrypoint.sh │ └── seaweedfs-s3.json ← SeaweedFS S3 auth config (example) ├── docs/ ← MkDocs site (user docs; mkdocs.yml at repo root) -│ ├── index.md · installation.md · object_store.md · usage.md · compaction.md +│ ├── index.md · walkthrough.md · walkthrough_demos.md · installation.md +│ ├── object_store.md · usage.md · compaction.md │ ├── architecture.md · architecture_tiered.md · architecture_decoupled.md │ ├── architecture_vectors.md · usage_vectors.md · changelog.md │ └── formal/ ← TLA+ model of the bakery protocol (Bakery.tla) +├── examples/walkthrough/ ← interactive walkthrough (guide.sh, its compose stack and configs) ├── docker-compose.yml ← END-USER single-node stack (ports published) ├── docker-compose.matrix.yml ← CI only: single-node vanilla matrix ├── docker-compose.matrix-azure.yml ← CI only: vanilla matrix on Azure ADLS @@ -225,11 +284,12 @@ The following table lists the services and components ColdFront runs against: | Component | Version | Purpose | |-----------|---------|---------| -| PostgreSQL | 16, 17, or 18 | Database with native partitioning (stock upstream; no fork) | -| pg_duckdb | 1.5.4 (PR #1025) | Iceberg reads + writes via DuckDB in-process | -| duckdb-iceberg | `v1.5-variegata` @ `5edc45f0`, patched | Iceberg catalog/IO for DuckDB; carries ColdFront's four patches (see [docker/Dockerfile.duckdb15-base](docker/Dockerfile.duckdb15-base)) | -| Lakekeeper | latest | Iceberg REST catalog (Rust binary) | -| S3-compatible store | any | SeaweedFS, MinIO, GCS, Azure Blob, etc. | +| PostgreSQL | 16, 17, or 18 | Provides the database with native partitioning; stock PostgreSQL serves a single node, and the Docker image and distributed mode use pgEdge's PostgreSQL build with Spock. | +| pg_duckdb | commit c04e6a2 (PR #1025), on DuckDB 1.5.4 | Runs Iceberg reads and writes through in-process DuckDB. | +| duckdb-iceberg | `v1.5-variegata` @ `5edc45f0`, patched | Provides Iceberg catalog and I/O for DuckDB, with ColdFront's four patches applied (see [docker/Dockerfile.duckdb15-base](docker/Dockerfile.duckdb15-base)). | +| Lakekeeper | latest | Provides the Iceberg REST catalog (a Rust binary). | +| S3-compatible store | any | Stores the cold data; SeaweedFS, MinIO, AWS S3, and GCS all work. | +| Azure ADLS Gen2 | any | Stores the cold data instead of an S3 store, through `set_storage_secret_azure` and an `adls` warehouse. | Building from source needs the Go toolchain (the version is pinned in [go.mod](go.mod)). The Go module dependencies are the source of truth in @@ -239,7 +299,7 @@ static, CGO-free binaries on `pgx/v5`, and the compactor is a separate module ## Versioning -ColdFront carries two independent version numbers, each following its own +ColdFront has two independent version numbers, each following its own convention: - Release tags use three-part [Semantic Versioning](https://semver.org) @@ -248,21 +308,35 @@ convention: required because ColdFront is a Go module, and the toolchain recognizes only full `vX.Y.Z` tags as releases. The patch field keeps a bugfix-only release (`v1.0.1`) distinct from a feature release (`v1.1.0`), which matters for a - data-writing extension where "same behavior, one safety fix" is worth stating - plainly. + data-writing extension, where a release that keeps the same behavior and adds + one safety fix must be easy to recognize. - The PostgreSQL extension uses the conventional two-part version in its control file (`default_version = '1.0'`) and upgrade-script filenames (`coldfront--1.0--1.1.sql`), as is standard for PostgreSQL extensions. The two map cleanly: extension `1.0` ships inside release `v1.0.0`, and a patch -release may carry the same extension version or bump it with an upgrade script +release may keep the same extension version or bump it with an upgrade script when the SQL changes. +## Support & Resources + +For more information about pgEdge products, visit +[docs.pgedge.com](https://docs.pgedge.com). + +To report an issue with the software, visit +[GitHub Issues](https://github.com/pgEdge/ColdFront/issues). + +## Contributing + +We welcome your project contributions; for more information, see +[CONTRIBUTING.md](CONTRIBUTING.md). + ## Author -Created by Jimmy Angelakos. +Jimmy Angelakos created ColdFront. ## License -PostgreSQL License. See [LICENSE.md](LICENSE.md). Redistributed third-party -components and their notices: [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). +This project is licensed under the [PostgreSQL License](LICENSE.md). The +redistributed third-party components and their notices are listed in +[THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). diff --git a/THIRD_PARTY_NOTICES.md b/THIRD_PARTY_NOTICES.md index 14226d5..9cf8a0e 100644 --- a/THIRD_PARTY_NOTICES.md +++ b/THIRD_PARTY_NOTICES.md @@ -11,22 +11,23 @@ Copyright © Stichting DuckDB Foundation — **DuckDB** (2018-2025), **pg_duckdb **postgres_scanner** extensions. duckdb-iceberg is modified by ColdFront (`docker/iceberg-bakery-aware-commit-refresh-v15.patch`). -> Permission is hereby granted, free of charge, to any person obtaining a copy of -> this software and associated documentation files (the "Software"), to deal in the -> Software without restriction, including without limitation the rights to use, -> copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the -> Software, and to permit persons to whom the Software is furnished to do so, -> subject to the following conditions: +> Permission is hereby granted, free of charge, to any person obtaining a copy +> of this software and associated documentation files (the "Software"), to deal +> in the Software without restriction, including without limitation the rights +> to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +> copies of the Software, and to permit persons to whom the Software is +> furnished to do so, subject to the following conditions: > -> The above copyright notice and this permission notice shall be included in all -> copies or substantial portions of the Software. +> The above copyright notice and this permission notice shall be included in +> all copies or substantial portions of the Software. > > THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR -> IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS -> FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR -> COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN -> AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION -> WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. +> IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +> AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +> LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +> OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +> SOFTWARE. ## PostgreSQL License @@ -43,17 +44,18 @@ client is httplib). It is distributed under the curl License: > Copyright (C) Daniel Stenberg, daniel@haxx.se, and many contributors. > > Permission to use, copy, modify, and distribute this software for any purpose -> with or without fee is hereby granted, provided that the above copyright notice -> and this permission notice appear in all copies. +> with or without fee is hereby granted, provided that the above copyright +> notice and this permission notice appear in all copies. > > THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR -> IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS -> FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT OF THIRD PARTY RIGHTS. IN NO EVENT -> SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER -> LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT -> OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE -> SOFTWARE. +> IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT OF THIRD PARTY RIGHTS. +> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, +> DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR +> OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE +> OR OTHER DEALINGS IN THE SOFTWARE. > -> Except as contained in this notice, the name of a copyright holder shall not be -> used in advertising or otherwise to promote the sale, use or other dealings in -> this Software without prior written authorization of the copyright holder. +> Except as contained in this notice, the name of a copyright holder shall not +> be used in advertising or otherwise to promote the sale, use or other +> dealings in this Software without prior written authorization of the +> copyright holder. diff --git a/docs/LICENSE.md b/docs/LICENSE.md index 94cd706..944fbc3 100644 --- a/docs/LICENSE.md +++ b/docs/LICENSE.md @@ -2,8 +2,18 @@ The PostgreSQL License Portions copyright (c) 2026, pgEdge, Inc. -Permission to use, copy, modify, and distribute this software and its documentation for any purpose, without fee, and without a written agreement is hereby granted, provided that the above copyright notice and this paragraph and the following two paragraphs appear in all copies. +Permission to use, copy, modify, and distribute this software and its +documentation for any purpose, without fee, and without a written agreement is +hereby granted, provided that the above copyright notice and this paragraph and +the following two paragraphs appear in all copies. -IN NO EVENT SHALL pgEdge, Inc. BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF pgEdge, Inc. HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. +IN NO EVENT SHALL pgEdge, Inc. BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, +SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING +OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF pgEdge, Inc. HAS +BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. -pgEdge, Inc. SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE SOFTWARE PROVIDED HEREUNDER IS ON AN "AS IS" BASIS, AND pgEdge, Inc. HAS NO OBLIGATIONS TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS. +pgEdge, Inc. SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED +TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR +PURPOSE. THE SOFTWARE PROVIDED HEREUNDER IS ON AN "AS IS" BASIS, AND pgEdge, +Inc. HAS NO OBLIGATIONS TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, +OR MODIFICATIONS. diff --git a/docs/architecture.md b/docs/architecture.md index 0a7f065..08be941 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -7,65 +7,51 @@ rewrites reads and writes to hit the right tier, and **all Iceberg I/O goes through `pg_duckdb` running in-process inside PostgreSQL** - no external query engine, no Go Iceberg libraries. -## Contents - -This document is organized into the following sections: - -- [Operating modes and topologies](#operating-modes-and-topologies) - the three - axes + read target (primary / standby) -- [System Overview](#system-overview) - the moving parts -- [Core Mechanics: pg_duckdb](#core-mechanics-pg_duckdb) - how Iceberg I/O - happens -- [Application Interface](#application-interface) - the shared rewrite hook -- [Concurrency and pgEdge Spock Deployments](#concurrency-and-pgedge-spock-deployments) - - bakery, cold-write strategy, DDL, OID-vs-name -- [Known Limitations](#known-limitations) - cross-cutting -- [Infrastructure (Docker)](#infrastructure-docker) -- [Upstream Requests](#upstream-requests) - open asks to pg_duckdb / - duckdb-iceberg - For mode-specific design, see [architecture_tiered.md](architecture_tiered.md) (hot PG + cold Iceberg) · [architecture_decoupled.md](architecture_decoupled.md) (all-Iceberg) · [architecture_vectors.md](architecture_vectors.md) (vector storage). -## Operating modes and topologies +## Operating Modes and Topologies -Three independent axes describe any ColdFront deployment, selected as shown -below. They compose freely - e.g. tiered + mesh + permissive writes: +Three independent axes describe any ColdFront deployment; they compose freely - +e.g. tiered + mesh + permissive writes. The following table shows how each axis +is selected: | Axis | Values | Selected by | |---|---|---| -| **Storage mode** | **Tiered** - hot PG heap + cold Iceberg, unified by a `UNION ALL` view; an archiver moves rows hot→cold on a cron. · **Decoupled** - the table lives entirely in Iceberg; PG holds only a wrapper view + a registry row (no archiver, no PG storage, no watermark). | Per relation at creation, via the `is_iceberg_only` flag on `coldfront.tiered_views` (short-circuited in the hook's `classify_tier()`). | -| **Topology** | **Vanilla** - single node; `spock`/`snowflake` not loaded; cold writes serialize on a local advisory lock. · **Mesh** - 3-node pgEdge Spock active-active; cold writes serialize cluster-wide via the bakery protocol. | Whether `spock`/`snowflake` are in `shared_preload_libraries`. One image and one SQL surface serve both; the `_exec_iceberg_with_claim` chokepoint self-selects via its `v_armed` gate. | -| **Write mode** | **Permissive** (default) - an ambiguous cross-tier `UPDATE`/`DELETE` writes both tiers. · **Strict** - it is rejected with a hint. | `coldfront.allow_mixed_writes` (USERSET). | +| Storage mode | In tiered mode, a `UNION ALL` view unifies the hot PG heap and cold Iceberg, and an archiver moves rows from hot to cold on a cron. In decoupled mode, the table lives entirely in Iceberg, and PG holds only a wrapper view and a registry row (no archiver, no PG storage, no watermark). | The `is_iceberg_only` flag on `coldfront.tiered_views` selects the mode per relation at creation, and the hook's `classify_tier()` short-circuits on that flag. | +| Topology | In a vanilla topology (a single node, with `spock`/`snowflake` not loaded), cold writes serialize on a local advisory lock. In a mesh (3-node pgEdge Spock active-active), cold writes serialize cluster-wide via the bakery protocol. | The topology follows whether `spock`/`snowflake` are in `shared_preload_libraries`. One image and one SQL surface serve both. The serializer follows `coldfront._bakery_armed()`: the bakery when both `snowflake.node` and `coldfront.dblink_self` are set, and a local advisory lock otherwise; the image sets both only when `MESH=on`. | +| Write mode | In permissive mode (the default), an ambiguous cross-tier `UPDATE`/`DELETE` writes both tiers. In strict mode, such a statement is rejected with a hint. | `coldfront.allow_mixed_writes` (USERSET) selects the write mode. | Both storage modes coexist in one database and share **one** code path: the transparent view and read rewriter, the INSERT/UPDATE/DELETE hook (`emit_cold` / `emit_hot` / `emit_dual` in [`extension/coldfront/src/coldfront.c`](https://github.com/pgEdge/ColdFront/blob/main/extension/coldfront/src/coldfront.c)), -and the `_exec_iceberg_with_claim` write chokepoint. Decoupled mode simply -always classifies as `TIER_COLD` and never reaches `emit_hot`; vanilla and mesh -differ only in how that chokepoint serializes cold writes. This document covers +and the per-table claim (`coldfront._take_iceberg_claim`) that every cold write +path in the hook takes, mostly through `_exec_iceberg_with_claim`. Decoupled +mode always classifies as `TIER_COLD` and never reaches `emit_hot`; vanilla and +mesh differ only in how that claim serializes cold writes. This document covers the shared mechanics and the tiered path; see [architecture_decoupled.md](architecture_decoupled.md) for the decoupled mode's ACID model and distributed scaling story. -### Read target: primary or physical standby +### Read Target: Primary or Physical Standby Orthogonal to the three axes above, any ColdFront node - vanilla or a mesh member - can have one or more **physical (streaming) standbys that serve read-only cross-tier reads**. The hot tier arrives by physical replication; the cold tier is read by `iceberg_scan` executing on the read-only backend. A base -backup carries everything a replica needs - the coldfront catalog -(`tiered_views`, `archive_watermark`, `storage_secret`), the DuckDB persistent -S3 secret (loaded at instance init), and the GUCs (in `postgresql.conf`, not -`ALTER SYSTEM`) - so a replica is byte-identical to its primary (same OIDs) -with zero extra setup. Cold **writes** are refused on a standby: every cold -write funnels through `_exec_iceberg_with_claim`, which raises when -`pg_is_in_recovery()`, so a read replica can never become an uncoordinated -writer to the shared Iceberg table (hot writes hit a PG heap and PG rejects -them natively). +backup contains the coldfront catalog (`tiered_views`, `archive_watermark`, +`storage_secret`) and the GUCs (in `postgresql.conf`, not `ALTER SYSTEM`), so a +replica is byte-identical to its primary (same OIDs). The DuckDB persistent +secret file sits under the OS user's home directory, outside PGDATA, so with +static credentials run `SELECT coldfront.materialize_storage_secret()` once on +the replica. Cold **writes** are refused on a standby: every cold write path in +the hook calls `coldfront._reject_on_standby`, which raises when +`pg_is_in_recovery()`, before it takes the claim, so a read replica can never +become an uncoordinated writer to the shared Iceberg table (hot writes hit a PG +heap and PG rejects them natively). Standby reads are gated by [`ci/probe-standby.sh`](https://github.com/pgEdge/ColdFront/blob/main/ci/probe-standby.sh) @@ -95,7 +81,7 @@ below: └──────────────┬───────────────────────────────────────────┘ │ ┌──────────────▼───────────────────────────────────────────┐ -│ Lakekeeper — Iceberg REST catalog (own dedicated Postgres)│ +│ Lakekeeper - Iceberg REST catalog (own dedicated Postgres)│ │ Manages Iceberg metadata, snapshots, commit concurrency │ └──────────────┬───────────────────────────────────────────┘ │ @@ -109,16 +95,16 @@ The following table describes each component, its role, and its license: | Component | Role | License | |-----------|------|---------| -| PostgreSQL 16+ | Heap storage; range partitioning for the tiered hot tier. Works uniformly on PG 16, 17, and 18 - the cold-tier secret is a DuckDB persistent secret loaded at instance init, with no version-gated mechanism. | PostgreSQL | -| pg_duckdb | DuckDB in-process. Iceberg read + write. Analytics. pg_duckdb 1.5.4 (PR #1025). The `duckdb-iceberg` carries the bakery-aware commit-refresh patch (async parquet overlap, no 409); see [Cold-write strategy](#cold-write-strategy-stock-vs-patched-duckdb-iceberg). | MIT | -| coldfront | PGXS C extension. `post_parse_analyze_hook` rewrites INSERT/UPDATE/DELETE on registered views to the correct tier and, on a SELECT DuckDB will run, the spellings DuckDB lacks (`date_bin`, `::jsonb`, the JSON builders); `planner_hook` folds bound parameters into such a read; `ProcessUtility_hook` handles DDL; the hook lazily ATTACHes the Iceberg catalog on the first query touching a tiered view. | PostgreSQL | -| Lakekeeper | Iceberg REST catalog. Single Rust binary. | Apache 2.0 | -| S3-compatible store | Any: SeaweedFS, MinIO, GCS, Azure Blob, etc. | Varies | -| Archiver (tiered mode) | Go binary, invoked by cron. Thin SQL orchestrator that moves rows hot→cold. | PostgreSQL | +| PostgreSQL 16+ | Provides heap storage and range partitioning for the tiered hot tier. ColdFront works uniformly on PG 16, 17, and 18, because the cold-tier secret is a DuckDB persistent secret loaded at instance init, with no version-gated mechanism. | PostgreSQL | +| pg_duckdb | Runs DuckDB in-process for Iceberg reads, writes, and analytics; the build is pg_duckdb at commit c04e6a2 (PR #1025) on DuckDB 1.5.4. The bundled `duckdb-iceberg` includes the bakery-aware commit-refresh patch (async parquet overlap, no 409); see [Cold-Write Strategy](#cold-write-strategy-stock-vs-patched-duckdb-iceberg). | MIT | +| coldfront | This PGXS C extension's `post_parse_analyze_hook` rewrites INSERT/UPDATE/DELETE on registered views to the correct tier and, on a SELECT DuckDB will run, the spellings DuckDB lacks (`date_bin`, `::jsonb`, the JSON builders); `planner_hook` folds bound parameters into such a read; `ProcessUtility_hook` handles DDL; the hook lazily ATTACHes the Iceberg catalog on the first query touching a tiered view. | PostgreSQL | +| Lakekeeper | Provides the Iceberg REST catalog as a single Rust binary. | Apache 2.0 | +| S3-compatible store or Azure ADLS Gen2 | Stores the cold data; any S3-compatible store works, including SeaweedFS, MinIO, AWS S3, and GCS, and Azure ADLS Gen2 works through `set_storage_secret_azure`. | Varies | +| Archiver (tiered mode) | Moves rows from hot to cold; this Go binary is a thin SQL orchestrator that cron invokes. | PostgreSQL | How rows move through this depends on the storage mode: the tiered hot heap + archiver + `UNION ALL` data-flow is in -[architecture_tiered.md → Data flow](architecture_tiered.md#data-flow); the +[architecture_tiered.md → Data Flow](architecture_tiered.md#data-flow); the all-Iceberg flow is in [architecture_decoupled.md](architecture_decoupled.md). ## Core Mechanics: pg_duckdb @@ -127,22 +113,27 @@ All Iceberg I/O goes through SQL executed against PostgreSQL. There are no Go DuckDB/Iceberg/Arrow libraries. DuckDB Iceberg writes require a REST catalog - Lakekeeper fills this role. -### Session setup +### Session Setup The cold-tier S3 secret is set once per cluster. A single call records the credentials and materializes a DuckDB persistent secret: ```sql SELECT coldfront.set_storage_secret('', '', ''); +-- Azure ADLS Gen2 instead of an S3-compatible store: +SELECT coldfront.set_storage_secret_azure(''); ``` -This does two things. (1) It stores the secret in the +`set_storage_secret` does two things. (1) The function stores the secret in the `coldfront.storage_secret` table - an extension-member table, so its data is excluded from `pg_dump` by default, and it is added to the Spock replication set, so the secret replicates **by value** to every mesh node with no per-node -file syncing. (2) It materializes a DuckDB **persistent secret**, which DuckDB -loads automatically at instance init - so every backend, including the first -fresh one, sees the secret at a committed timestamp before any query runs. +file syncing. (2) The function materializes a DuckDB **persistent secret**, +which DuckDB loads automatically at instance init - so every backend, including +the first fresh one, sees the secret at a committed timestamp before any query +runs. `set_storage_secret_azure` takes one connection string holding the +account name, account key and endpoint suffix, writes the same row, and +materializes a `TYPE azure` persistent secret. For deployments that must not store a credential at all, `coldfront.set_storage_secret_vended()` records a vended @@ -159,28 +150,29 @@ per-table secret inside the commit transaction. For backup and restore (`pg_dump`), the durable tiering metadata - `coldfront.tiered_views` (registry), `archive_watermark` (cutoffs) and `partition_config` - is marked with `pg_extension_config_dump`, so a logical -`pg_dump` carries it and a restore re-attaches to the **same** Iceberg cold +`pg_dump` includes it and a restore re-attaches to the **same** Iceberg cold tier with no re-provisioning. Two things are deliberately **not** dumped: the credential (`coldfront.storage_secret` above - re-run `set_storage_secret` (or -`set_storage_secret_vended`) once after restoring into a fresh instance) and -the bakery's transient claim tables (`claims` / `claim_acks` / `deferred_acks`, -per-node mesh state). Until the credential is re-established a restored node -serves hot reads but fails cold I/O cleanly. `ci/ops.sh` Check 4 exercises -exactly this. +`set_storage_secret_azure` or `set_storage_secret_vended`) once after restoring +into a fresh instance) and the bakery's transient claim tables (`claims` / +`claim_acks` / `deferred_acks`, per-node mesh state). Until the credential is +re-established a restored node serves hot reads but fails cold I/O cleanly. +`ci/ops.sh` Check 4 exercises exactly this. The Iceberg catalog ATTACH is **lazy**: the coldfront C extension hook issues `ATTACH IF NOT EXISTS` against Lakekeeper - using the cluster's `coldfront.warehouse` and `coldfront.lakekeeper_endpoint` GUCs - on the **first query that touches a tiered view** (read or write), per DuckDB cached -connection. There is no arming step and no per-session boilerplate: both reads -(`iceberg_scan`) and writes (`duckdb.raw_query`) just work on a fresh psql -session. Until a tiered view is touched no ATTACH is attempted, so a -pre-bootstrap connection is never blocked by a missing warehouse. +connection. There is no setup step and no per-session boilerplate: both reads +(`iceberg_scan`) and writes (`duckdb.raw_query`) work on a fresh psql session. +Until a tiered view is touched no ATTACH is attempted, so a pre-bootstrap +connection is never blocked by a missing warehouse. -### Non-superuser app roles (least privilege) +### Non-Superuser App Roles (Least Privilege) -pg_duckdb force-disables DuckDB's `LocalFileSystem` for non-superusers (see -[Upstream Requests](#pg_duckdb-non-superuser-localfilesystem-blocks-side-loaded-extensions)), +pg_duckdb force-disables DuckDB's `LocalFileSystem` for any role that lacks +both `pg_read_server_files` and `pg_write_server_files` (see +[Upstream requests](#pg_duckdb-non-superuser-localfilesystem-blocks-side-loaded-extensions)), which would block the side-loaded iceberg/postgres DuckDB extensions from loading on `ATTACH`. So `coldfront.ensure_attached()` / `ensure_pg_attached()` are `SECURITY DEFINER` with a pinned `search_path`: the extension load + @@ -188,36 +180,43 @@ are `SECURITY DEFINER` with a pinned `search_path`: the extension load + because the DuckDB instance is per-backend the attach persists for the session - every subsequent `iceberg_scan` / `_exec_iceberg_with_claim` then runs as the **app role** over S3/httpfs, never touching `LocalFileSystem`. The -app role needs only `duckdb.postgres_role` membership + object grants; **no -superuser, no `pg_{read,write}_server_files`**. +app role needs only `duckdb.postgres_role` membership, object grants, and SET +on the superuser-only `duckdb.unsafe_allow_execution_inside_functions` +parameter, which the cross-tier move needs; **no superuser, no +`pg_{read,write}_server_files`**. Because the attach helpers run elevated, the deployment-config GUCs they consume (`coldfront.warehouse`, `coldfront.lakekeeper_endpoint`, `coldfront.local_pg_dsn`) are registered `PGC_SUSET` (the last also `GUC_SUPERUSER_ONLY`) in `_PG_init`, so a non-superuser cannot redirect the -elevated `ATTACH` at an attacker endpoint. Onboarding is one operator call, +elevated `ATTACH` at an attacker endpoint. `coldfront.dblink_self`, the DSN of +the bakery's loopback, is `PGC_SUSET` as well, because the loopback runs claim +statements as the user that DSN names. It stays readable by every role, since +the invoker-rights `coldfront._bakery_armed()` reads it on every cold write. +All four GUCs default to an empty string. Onboarding is one operator call, `coldfront.grant_app_access(role)` - idempotent, registry-derived (schemas, views, the hot heap + its identity sequence, the cold-path function EXECUTE allow-list), not `PUBLIC`-executable. The image defaults `duckdb.postgres_role = coldfront_duckdb` (env `COLDFRONT_DUCKDB_ROLE`) and -creates the role, so the path is turnkey. +creates the role, so the path needs no further setup. In a **Spock mesh** the role and its grants replicate via Spock DDL - onboard -once on any node. Mesh cold *writes* route through the R-A bakery; its +once on any node. Mesh cold *writes* route through the bakery protocol; its coordination function `_claim_iceberg_lock` is itself `SECURITY DEFINER` (search_path-pinned, fully schema-qualified) so a non-superuser drives the cross-node serialization (`pg_stat_replication` liveness + the loopback claim) with the privilege it requires. `_exec_iceberg_with_claim` deliberately stays `SECURITY INVOKER` - it runs the caller's cold DML, which must execute as the -caller. The bakery SD is **protocol-neutral**: it changes the PG execution -privilege, not the claim/ack/lock/ticket protocol, re-verified against -[the TLA+ model](formal/Bakery.tla) (all safe configs pass; the race config -still violates `NoLakekeeperConflict`). +caller. This `SECURITY DEFINER` setting is **protocol-neutral**: it changes the +PG execution privilege, not the claim/ack/lock/ticket protocol, re-verified +against [the TLA+ model](formal/Bakery.tla) (all safe configs pass; the race +config still violates `NoLakekeeperConflict`). -See README "Security"; asserted by the journey's `story_app_privilege`, -`ci/ops.sh` check 3, and the `privilege_model` pg_regress test. +See README "Security"; the setting is asserted by the journey's +`story_app_privilege`, `ci/ops.sh` check 3, and the `privilege_model` +pg_regress test. -### Temp table bridge: PG → Iceberg +### Temp Table Bridge: PG → Iceberg `duckdb.raw_query()` cannot see PG tables directly. The bridge is a DuckDB temp table: @@ -230,7 +229,7 @@ SELECT duckdb.raw_query($$INSERT INTO ice.public.events DROP TABLE duck_stage; ``` -### Cold-side column references +### Cold-Side Column References `iceberg_scan()` requires `r['col']::type` syntax: @@ -246,8 +245,8 @@ Applications use the transparent view exactly like a table. A `post_parse_analyze_hook` in the coldfront extension intercepts INSERT/UPDATE/DELETE whose target is a registered relation - resolved in `coldfront.tiered_views` by name (`schema_name`, `relname`) - and rewrites the -parsed `Query` so it lands in the correct tier; cold-side writes funnel through -the `_exec_iceberg_with_claim` chokepoint (see +parsed `Query` so it lands in the correct tier; cold-side writes go through +`_exec_iceberg_with_claim` (see [Concurrency](#concurrency-and-pgedge-spock-deployments)). A `ProcessUtility_hook` handles DDL on the same relations. @@ -256,56 +255,56 @@ the statement the view is named (a CTE, a sub-select, a set-operation branch): `date_bin`, the `::jsonb` cast, `jsonb_array_length` and the JSON builders (`jsonb_build_object`, `jsonb_agg` and their `json_` twins) are rewritten into spellings both engines accept (see -[usage.md → Supported column types](usage.md#supported-column-types)). A +[usage.md → Supported Column Types](usage.md#supported-column-types)). A `planner_hook` folds bound parameters into such a read before pg_duckdb plans it when a parameter sits where DuckDB cannot type a placeholder (a direct argument of a pg_duckdb function, any argument of a table function); the plan -cache's generic-plan build, which carries no values, gets a PostgreSQL plan -priced above any custom plan, so under the plan cache's cost-based selection -the read is planned from its values on every execution. +cache's generic-plan build, which has no values, gets a PostgreSQL plan priced +above any custom plan, so under the plan cache's cost-based selection the read +is planned from its values on every execution. `plan_cache_mode = force_generic_plan` bypasses that selection and picks the value-less generic plan, which fails with `only works with DuckDB execution` -(see [usage.md → Supported column types](usage.md#supported-column-types)). +(see [usage.md → Supported Column Types](usage.md#supported-column-types)). The following table maps each operation to its interface and routing path: | Operation | Interface | Routed via | |---|---|---| -| SELECT | `SELECT FROM events` | pg_duckdb (UNION ALL in tiered, wrapper view in decoupled) | +| SELECT | `SELECT FROM events` | pg_duckdb (UNION ALL in tiered, wrapper view in decoupled); a single-table tiered read whose WHERE proves it hot is rerouted to the hot heap in plain PostgreSQL | | INSERT | `INSERT INTO events ...` | coldfront `post_parse_analyze_hook` | | UPDATE / DELETE | `... events WHERE ...` | coldfront `post_parse_analyze_hook` | | DDL (ALTER / RENAME) | `ALTER TABLE _events ...` | coldfront `ProcessUtility_hook` (see [Transparent DDL](#transparent-ddl-via-coldfront)) | -| DROP / TRUNCATE | `DROP TABLE _events` | blocked by the hook (see [Transparent DDL](#transparent-ddl-via-coldfront)) | +| DROP / TRUNCATE | `DROP TABLE _events` | The hook blocks the statement (see [Transparent DDL](#transparent-ddl-via-coldfront)). | With `duckdb.force_execution = true`, hot-side queries are also accelerated by DuckDB's vectorized columnar engine. How the hook splits a write is mode-specific: -- **Tiered** routes by the partition-column watermark - hot heap vs cold - Iceberg, with dual-tier writes for ambiguous predicates. See +- In tiered mode, the hook routes by the partition-column watermark - hot heap + vs cold Iceberg, with dual-tier writes for ambiguous predicates. See [architecture_tiered.md → Transparent INSERT](architecture_tiered.md#transparent-insert) and [→ UPDATE/DELETE](architecture_tiered.md#transparent-updatedelete). -- **Decoupled** always classifies `TIER_COLD`: every write is a single-tier - Iceberg write. See [architecture_decoupled.md](architecture_decoupled.md). +- In decoupled mode, the hook always classifies `TIER_COLD`, so every write is + a single-tier Iceberg write. See + [architecture_decoupled.md](architecture_decoupled.md). -### Cold-tier DML from inside plpgsql (functions, DO blocks, triggers) +### Cold-Tier DML from Inside PL/pgSQL (Functions, DO Blocks, Triggers) Cold-tier `INSERT`/`UPDATE`/`DELETE` work as top-level statements *and* from inside a plpgsql function / `DO` block / trigger, via two mechanisms: -1. **Parameters are emitted as a runtime `format(