diff --git a/API_CONFIG_CONTRACT.md b/API_CONFIG_CONTRACT.md index 7b748bcc..2bb3b986 100644 --- a/API_CONFIG_CONTRACT.md +++ b/API_CONFIG_CONTRACT.md @@ -1,6 +1,6 @@ -# Version 1 Configuration Contracts +# Versioned Configuration Contracts -## Generic `cwl-pingora-gateway` +## Generic `cwl-pingora-gateway` version 1 ```yaml version: 1 @@ -25,7 +25,7 @@ upstreams: Unknown fields are rejected. `version` must be `1`. `listener`, `metrics_listener`, and `address` are socket addresses with non-zero ports. Port zero is rejected because production authority must remain operator-declared rather than OS-selected or unusable. Traffic and metrics listeners must not overlap one effective socket authority: equal sockets, same-port same-family wildcard/concrete aliases, exact native/IPv4-mapped aliases, native or mapped IPv4 wildcard aliases, and the platform-dependent same-port IPv6-wildcard/IPv4 combination fail closed. Distinct concrete non-aliased addresses may share a non-zero port. `max_request_body_bytes`, `max_in_flight_requests`, and `upstream_keepalive_pool_size` must all be positive. Generic v1 requires exactly one upstream and a non-empty stable upstream name. Every timeout must be positive. -The timeout fields map directly to the pinned Pingora peer options rather than defining a second gateway timer model. In particular, `read_ms` is a **per-read inactivity budget**: Pingora waits at most that long for each individual upstream `read()` and resets the timer after a successful read. It is not a total-response deadline. A connected upstream that sends no response bytes is therefore bounded by `read_ms`, while a slow-drip response can remain alive across multiple successful reads. Whole-response lifetime remains an explicit open runtime-isolation requirement and must not be inferred from `read_ms`. +The timeout fields map directly to the pinned Pingora peer options rather than defining a second gateway timer model. In particular, `read_ms` is a **per-read inactivity budget**: Pingora waits at most that long for each individual upstream `read()` and resets the timer after a successful read. It is not a total-response deadline. A connected upstream that sends no response bytes is therefore bounded by `read_ms`, while a slow-drip response can remain alive across multiple successful reads. Generic v1 has no whole-response lifetime and must not infer one from `read_ms`. `max_in_flight_requests` is a process-local backpressure boundary for non-health downstream requests. When the budget is exhausted, the runtime fails fast with HTTP 503 instead of admitting unbounded work. `/livez` and `/readyz` bypass this application admission budget so saturation does not hide process health. The admission lease is released when the request context ends, including failed requests. `upstream_keepalive_pool_size` is wired directly into Pingora's `ServerConf`; the runtime does not inherit Pingora's framework default of 128 reusable upstream connections. @@ -37,14 +37,15 @@ Generic v1 downstream transport is cleartext TCP. Before proxying, the generic a ## Bounded `cwl-pingora-pg-erd-migration` candidate -The dedicated pg-erd migration binary consumes a different, migration-specific Admin Config profile. It deliberately reuses the same top-level deployment value names while admitting exactly two fixed transport authorities: +The dedicated pg-erd migration binary consumes a different, migration-specific Admin Config profile. Version 1 preserves the existing unreleased characterization semantics and rejects `max_upstream_response_body_ms`. Version 2 is the explicit response-lifetime increment and requires a positive `max_upstream_response_body_ms`; the runtime never injects a hidden default. ```yaml -version: 1 +version: 2 listener: 0.0.0.0:6188 metrics_listener: 127.0.0.1:6192 max_request_body_bytes: 1048576 max_in_flight_requests: 128 +max_upstream_response_body_ms: 30000 upstream_keepalive_pool_size: 32 upstreams: - name: backend @@ -67,6 +68,12 @@ upstreams: idle_ms: 10000 ``` +The numeric response-lifetime value above is illustrative configuration, not a production SLO. A deployment owner must choose the version-2 value from its observed long-response contract before canary or cutover. + +`max_upstream_response_body_ms` starts at the first non-informational upstream response header. Runtime Isolation compares elapsed monotonic time only when a non-empty upstream body chunk is observed. Progress at or beyond the configured lifetime raises an upstream-scoped fatal error; empty/end-of-stream bookkeeping callbacks do not manufacture a timeout. If the final response has already been written, `fail_to_proxy` observes Pingora's `Session::response_written()` commitment state and emits no second status. Pre-commit upstream failures retain the existing policy-complete local error response behavior. No route failover is introduced by this contract. + +This callback guard is not an exact timer interrupt. `read_ms` remains a per-read inactivity budget that resets after a successful read. A continuously progressing body is stopped at the first non-empty body callback at or beyond `max_upstream_response_body_ms`; a response that becomes quiescent is bounded by `read_ms`. The current callback surface does not wake a pending read exactly at the body-lifetime instant, and slow delivery of an incomplete response header remains a separate transport gap. + This is not a generic multi-route configuration language. Operator input can bind only concrete transport/TLS values for the compiled `backend` and `frontend` identities. Missing, extra, duplicate, renamed, port-zero, or otherwise invalid listener/metrics/upstream transport authorities fail closed before listener activation. Listener and metrics sockets consume the same effective-authority invariant as generic v1, while the migration profile keeps its specific zero-transport-authority error contract. Routes and edge-owned response fields are not configurable: the characterized profile fixes exact `/healthz -> backend`, raw `PathPrefix(`/api`) -> backend` semantics including `/apiary`, fallback `/ -> frontend`, and the four captured response fields `X-Content-Type-Options: nosniff`, `X-Frame-Options: DENY`, `Referrer-Policy: no-referrer`, and `Permissions-Policy: geolocation=(), microphone=(), camera=()`. Admin parsing validates only deterministic configuration and authority invariants. It does not read custom trust-bundle bytes. If an admitted TLS upstream supplies `trust_bundle_file`, the canonical Pingora peer adapter reads and parses that material exactly once during `build_proxy`, still before listeners are registered. An unreadable or invalid bundle therefore blocks activation without a validate-then-reload trust-file window. diff --git a/CHANGELOG.md b/CHANGELOG.md index 8876952c..ccd1b363 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,50 +4,32 @@ All notable changes are tracked here. No release has been published yet. ## Unreleased -- Bootstrapped an executable Rust Pingora proxy through pull-request governance. -- Added strict v1 configuration, explicit one-upstream network authority, TLS identity verification, and explicit upstream I/O budgets. -- Added a transport-neutral Edge Routing characterization for the live `pg-erd-cloud` Traefik contract: exact `/healthz`, raw-prefix `/api`, fallback `/`, explicit numeric precedence, and fail-closed ambiguous-priority/malformed-route rejection. This does not by itself claim traffic cutover. -- Added a separate transport-neutral HTTP Policy characterization for the live `pg-erd-cloud` Traefik response-security middleware: exact `X-Content-Type-Options`, `X-Frame-Options`, `Referrer-Policy`, and `Permissions-Policy` values, ASCII case-insensitive field identity, duplicate authority rejection, and RFC 9110 field-value admission that rejects invalid controls plus leading/trailing SP or HTAB while preserving valid interior whitespace. -- Added a transport-neutral `EdgeMigrationPlan` application composition over the characterized route and HTTP-policy contracts plus an explicit normalized upstream-authority set. The pg-erd-cloud plan admits only `backend` and `frontend`, rejects undeclared route targets, and does not create product-domain or service-discovery authority. -- Added `MigrationDeliveryPlan` to bind every characterized migration upstream identity to exactly one explicit, prevalidated Pingora `HttpPeer`; missing, duplicate, and undeclared concrete transport authority fails closed. This does not widen generic `GatewayConfig` v1. -- Added `MigrationGatewayProxy` over characterized routing/HTTP policy, explicit peer binding, runtime isolation, transport-derived forwarding trust, failure-response policy, and shared observability. It selects only prevalidated peers, rejects unmatched routes, and enforces body/in-flight budgets. -- Added bounded `PgErdMigrationConfig` and a separate `cwl-pingora-pg-erd-migration` composition root. Operators may configure listener/metrics sockets, non-zero runtime/keepalive budgets, and concrete `backend`/`frontend` transport/TLS values, but routes, response-policy fields, product auth/business logic, service discovery, and new migration authorities remain non-configurable. Generic `GatewayConfig` v1 public shape and one-upstream semantics remain unchanged. -- Tightened the shared edge network-authority invariant used by generic and pg-erd configuration. Generic v1 rejects port-zero traffic, metrics and upstream bindings; both profiles reject same-port wildcard listener aliases, native/IPv4-mapped aliases and IPv6-wildcard/IPv4 dual-stack ambiguity while retaining distinct concrete non-aliased addresses. -- Added a fail-closed HTTP/1 protocol-transition admission boundary for generic v1 and the bounded pg-erd composition root. Any `Upgrade` field or case-insensitive comma-delimited `upgrade` token in `Connection` returns HTTP 501; callback ordering places that rejection before request admission and route/upstream selection. Compiled real-listener acceptance then observes no origin connection during a fixed 500 ms window after the verified 501 response and requires `/readyz` to stay available. `x-upgrade` and other unrelated tokens do not match. This is a deliberate non-support contract, not WebSocket implementation; HTTP/2 Extended CONNECT and HTTP/3/QUIC remain separate versioned work. -- Aligned immutable Pingora peer construction with the same fail-closed HTTP/1 transition contract. `pingora_delivery` now uses `HttpUpstreamRequestPolicy::deny_upgrades()` instead of the supplier `standard()`/`WebSocketOnly` default while retaining standard hop-by-hop and `Connection`-nomination sanitization. A dedicated peer-level regression prevents a later callback/composition refactor from silently re-enabling WebSocket forwarding below request admission. -- Tightened the generic v1 forwarding trust boundary: request-controlled `Forwarded`, `X-Forwarded-For`, `X-Forwarded-Host`, `X-Forwarded-Port`, `X-Forwarded-Proto`, `X-Forwarded-Server`, and `X-Real-IP` are all stripped before emitting only gateway-owned `Forwarded: proto=http`. -- Kept migration admin parsing side-effect free for custom TLS trust material: deterministic authority validation happens during parse, while peer/trust-bundle materialization occurs once during `build_proxy` before listener creation, avoiding a validate-then-reload trust-file window. -- Added a dedicated compiled pg-erd listener contract: `/livez` and `/readyz` remain process-local, consumer `/healthz` and raw `/api` prefix traffic follow the characterized origins, response policy is replaced at the edge, forwarding identity is rebuilt from accepted client transport plus Host/scheme authority, and declared oversize bodies fail before origin delivery. This is executable candidate evidence, not canary/cutover evidence. -- Added an Ingress Forwarding Policy that removes request-controlled `Forwarded`, `X-Forwarded-*`, and `X-Real-IP` authority before rebuilding only the characterized pg-erd compatibility fields from accepted client transport and validated Host/scheme authority. The listener bind socket is not treated as external-port truth. The current clear-text `web` characterization emits `http`; HTTPS remains a separate listener/TLS contract. -- Added shared payload-free observability for both Pingora adapters: low-cardinality completion status/outcome/body-byte facts plus backpressure counters, with paths, query strings, headers, cookies, credentials, product identifiers, and customer payloads excluded. -- Added dedicated compiled pg-erd payload-free observability acceptance. A routed request carries unique URI/query, Host, Authorization, Cookie, and product-context sentinels; the backend must receive those exact values while shared gateway stderr contains only the bounded completion vocabulary and none of the sentinels. The fixture keeps traffic/metrics listener reservations distinct through config construction, admits the process only after a complete bounded `/readyz` HTTP/1.1 200 response, bounds origin header receipt to five seconds/64 KiB, and uses exact HTTP-field plus Prometheus sample matching to reject lookalike false positives. -- Added a process-wide payload-safe dependency logging policy to both production binaries. Operator `RUST_LOG` still selects levels and targets, but Pingora-family dependency message bodies are replaced with a static marker before formatting so supplier trace/debug/error records cannot expose request URI, Host, Authorization, Cookie, payload, or other request-derived material. The compiled generic regression snapshots redaction activity after readiness probes, requires the secret-bearing request to reach the origin, requires a new redacted Pingora diagnostic after that request, and requires none of the sentinels in process stderr. Product/consumer logging remains outside this gateway-owned boundary. -- Added mandatory positive `max_in_flight_requests` and `upstream_keepalive_pool_size` capacity budgets; Pingora's framework keepalive default is overridden from the validated edge contract. -- Added process-local fail-fast backpressure: non-health requests above the in-flight budget receive HTTP 503, health remains observable, rejection telemetry increments, and capacity is released after request completion or failure. -- Added dedicated compiled pg-erd traffic acceptance for streamed/chunked body overflow and routed in-flight saturation/recovery: the migration process must return 413 above the shared body budget, return 503 in less than one second above the in-flight budget, keep `/readyz` observable, expose the exact single-rejection Prometheus sample, and admit a later routed request after capacity is released. -- Added dedicated compiled pg-erd refused-origin recovery acceptance using a Linux TCP socket bound to the characterized backend address without entering LISTEN state. The fixture first proves direct `ECONNREFUSED` while retaining exclusive port ownership, then requires the migration gateway to return 502 within a conservative one-second envelope around the configured 200/400 ms connection budgets, keep `/readyz` 200, expose the exact single-error Prometheus sample, and allow a later independent frontend route to recover. Connected read stall, TCP reset, partial-response/streaming failure, retry and failover behavior remain separate gaps. -- Added dedicated compiled pg-erd connected read-stall acceptance with `read_ms=100`: the backend accepts the routed request, completes bounded request-header receipt, records that causal point, and remains open without response bytes until the gateway has already failed it. The contract requires 502 no earlier than a conservative 50 ms lower bound from completed origin headers and still inside a one-second outer envelope, preserves `/readyz`, requires the exact single-error Prometheus sample, and proves independent frontend recovery. Pingora `read_timeout` remains a per-read inactivity budget, not a whole-response lifetime; reset, post-commit partial response, slow-drip and whole-response-deadline cases remain open. -- Added dedicated compiled pg-erd pre-header TCP-reset acceptance on Linux. The backend must first receive complete `/api/reset` request headers, then apply abortive `SO_LINGER(0)` before emitting any response bytes. Traffic admission now requires a complete bounded `/readyz` HTTP/1.1 200 response rather than bare TCP connect; origin header receipt remains bounded to five seconds/64 KiB. Acceptance requires exact HTTP/1.1 502 in under two seconds despite `read_ms=5000`, no silent frontend failover, preserved `/readyz`, exactly one low-cardinality request-error sample, and a later independent frontend HTTP 200. Broader streaming/upgraded failure, slow-drip and whole-response lifetime remain separate gaps. -- Added dedicated compiled pg-erd post-commit TCP-reset acceptance on Linux. The backend receives `/api/post-commit-reset` headers under one absolute five-second/64 KiB fixture bound, emits exact HTTP/1.1 200 framing with one `Content-Length: 20` field plus the seven-byte `partial` prefix, and then waits for bounded Linux `SIOCOUTQ == 0` after `write_all`; Linux computes that queue from `write_seq - snd_una`, so the origin releases abortive `SO_LINGER(0)` only after the peer TCP stack has acknowledged every response byte emitted before the reset. This transport acknowledgement is not treated as proof that Pingora application code parsed or forwarded the response. Downstream termination is independently read under one absolute five-second/64 KiB evidence bound. Both repeated-read paths derive each socket timeout from the remaining `Instant` deadline, and dedicated slow-drip regressions prove partial progress cannot renew either total budget. Acceptance requires exact committed HTTP/1.1 200 framing with exactly one `Content-Length: 20`, permits only an exact prefix of `partial` if any body bytes cross, requires the incomplete downstream response to terminate before 20 bytes, forbids a second status or silent failover, records exactly one request error, keeps `/readyz` healthy, and proves independent frontend recovery. Exact status/header/metric oracles reject protocol, field-name, duplicate-sample, and numeric-prefix lookalikes. -- Added dedicated routed pg-erd graceful-drain acceptance: a characterized `/api/held` backend request is held in flight, backend request-header receipt must complete within five seconds and 64 KiB before the request is considered admitted, SIGTERM is sent only after that causal point, the response is released during the shared grace period, the downstream must still receive HTTP 200, and the migration process must exit successfully inside the single absolute external termination budget anchored at signal delivery. The fixture retains traffic and metrics socket reservations through config construction, releases them only at child-bind handoff, and admits the drain case only after bounded `/readyz` HTTP/1.1 200 readiness; generic drain evidence is not transferred to this composition root. -- Added dedicated compiled pg-erd post-header truncation acceptance: the backend commits exact HTTP/1.1 200 framing with one `Content-Length: 20` field and the seven-byte `partial` prefix, remains open until downstream commit evidence is observed, then closes. The fixture requires the committed response to terminate incomplete without a second status or failover, records exactly one request error, preserves `/readyz`, and proves an independent frontend route still works. Status parsing rejects protocol/prefix lookalikes, origin request-header reads are time/size bounded, and content-length parsing rejects lookalike or duplicate/conflicting fields. -- Added routed pg-erd concurrency/latency acceptance with bounded Rust-only measured origins. The load lane format-checks and directly tests the std-only origin, builds an optimized fixture, runs 4 VUs for 400 iterations, requires exact status/body and zero HTTP failures, independently gates aggregate/backend/frontend p95 below 20 ms, and requires at least 198 samples per route. Route selection remains compiled into the pg-erd migration plan rather than becoming operator-configurable CI data. Controlled loopback evidence is not production SLO proof. -- Added optional per-upstream absolute PEM trust-bundle consumption without taking ownership of certificate issuance/rotation; trust material is loaded fail-closed before listeners open. -- Added an executable local-CA TLS test through the compiled generic gateway that holds CA trust constant and proves SNI/hostname mismatch is rejected. -- Added dedicated pg-erd real-listener upstream TLS acceptance without changing production routing/TLS code. An ephemeral one-day local CA plus `backend.test` certificate must allow characterized `/api` traffic only with matching explicit CA/SNI; the same CA with mismatched SNI must fail as exact HTTP 502 without route failover, preserve `/readyz`, and leave an independent `frontend` route usable. Fixture I/O is bounded to five seconds/64 KiB, traffic/metrics ports stay concurrently reserved until startup, and exact HTTP/1.1 status parsing rejects numeric-prefix/protocol-case false GREEN. This remains gateway-to-upstream TLS evidence only; downstream TLS termination, certificate issuance/rotation and representative TLS latency are not claimed. -- Added a focused transport-adapter regression proving an upstream without a custom trust bundle leaves Pingora's platform trust roots selected rather than replacing the CA store. -- Added fail-closed binary startup and real loopback production-path tests, including held-request saturation/recovery at an in-flight budget of one. -- Hardened the generic compiled production-path fixture after instrumented coverage exposed an ephemeral-port/readiness race: traffic and metrics loopback sockets now remain reserved through config construction and are released only at the child-bind handoff, and traffic startup requires a bounded application-level `/readyz` HTTP/1.1 200 carrying `Cache-Control: no-store` rather than a bare TCP accept. This prevents a sibling test or recycled loopback port from manufacturing readiness evidence without changing production gateway semantics. -- Added `/livez` and `/readyz` through the Pingora serving path. -- Added request-body limits and a distrust-by-default forwarded-header policy. -- Added low-cardinality metrics plus credential/cookie-safe access logging through the production path. -- Overrode Pingora framework retry/drain defaults with one total upstream attempt, a 5-second SIGTERM grace period, and a 30-second graceful-shutdown timeout. -- Added non-root/read-only-root OCI packaging with a fail-closed build-time allowlist for the generic and bounded pg-erd process identities; exact-head OCI acceptance builds and starts both profiles under uid/gid 65532, dropped capabilities and `no-new-privileges`, while the supply-chain lane builds and vulnerability-scans both candidate images. -- Extended pg-erd OCI acceptance so promotion also requires the separately published `/metrics` listener to identify the Prometheus service by exact base media type `text/plain`; a bare HTTP 200 or prefix-wildcard media-type match is insufficient. -- Added a committed dependency lock, fail-closed license/source/advisory policy, exact-source SBOM and image-vulnerability evidence. -- Added an exact-head owned-production coverage gate that requires 100% lines and regions without filename/function/branch exclusions; repaired compiler-generated generic startup coverage and structurally impossible literal-header error regions rather than weakening the gate. -- Added missing-public-rustdoc enforcement and documentation builds with warnings denied. -- Added a load-workflow contract that proves the measured loopback origin is ready before gateway startup so fixture races cannot be counted as gateway latency or availability behavior, and separately preserves the primary failure when k6 never produces a summary. -- Added DDD, product, technical, security, threat, test, operability, configuration, migration-gap, and primary-source traceability documentation. +- Bootstrapped an executable Rust Pingora proxy through pull-request governance with strict versioned configuration and explicit network authority. +- Added exact upstream transport/TLS identity, optional pre-activation PEM trust bundles, RFC 6066 SNI hostname admission, explicit connection/read/write/idle budgets, and fail-closed invalid authority. +- Added transport-neutral pg-erd Edge Routing characterization for exact `/healthz`, raw-prefix `/api` semantics including `/apiary`, fallback `/`, explicit precedence, and ambiguous/malformed route rejection. +- Added the separate pg-erd HTTP Policy characterization for the captured response-security fields, case-insensitive field identity, duplicate authority rejection, and RFC 9110-compatible field-value admission. +- Added `EdgeMigrationPlan` and `MigrationDeliveryPlan` so the bounded pg-erd profile can consume only compiled `backend`/`frontend` route authority and prevalidated concrete Pingora peers without becoming product service discovery. +- Added the dedicated `cwl-pingora-pg-erd-migration` composition root and bounded `PgErdMigrationConfig`; generic `GatewayConfig` v1 remains a one-upstream contract. +- Added effective socket-authority validation for traffic, metrics, and upstream endpoints, including port-zero, wildcard/concrete, native/IPv4-mapped, dual-stack ambiguity, recursive self-loop, broadcast, and multicast rejection while preserving valid distinct concrete unicast authority. +- Added a fail-closed HTTP/1 protocol-transition boundary. `Upgrade` or a `Connection: upgrade` token returns HTTP 501 before admission/upstream selection, and immutable Pingora peers use `HttpUpstreamRequestPolicy::deny_upgrades()` so supplier WebSocket defaults cannot bypass the public contract. +- Tightened generic and pg-erd forwarding trust. Request-controlled `Forwarded`, `X-Forwarded-*`, and `X-Real-IP` authority is discarded; generic v1 emits only gateway-owned `Forwarded: proto=http`, while pg-erd rebuilds only the characterized compatibility fields from accepted transport/request authority. +- Kept migration Admin Config parsing side-effect free for custom trust bytes. Deterministic validation completes before `build_proxy()` materializes peer/trust state once and before listener creation. +- Added compiled pg-erd production-path traffic proving process-local `/livez`/`/readyz`, characterized backend/fallback routing, exact response-policy replacement, forwarding reconstruction, declared and streamed request-body rejection, and readiness recovery. +- Added shared process-local request isolation: positive `max_in_flight_requests`, positive `upstream_keepalive_pool_size`, fail-fast HTTP 503 backpressure, lease recovery, and low-cardinality rejection telemetry. +- Added shared payload-free observability for both adapters plus process-wide Pingora-family diagnostic redaction so broad `RUST_LOG` settings cannot expose request URI, Host, Authorization, Cookie, product context, or payload data. +- Added deterministic pg-erd failure traffic for refused origin, connected read stall, orderly post-header truncation, pre-header TCP reset, post-commit TCP reset, and later independent-route recovery. Post-commit cases preserve the committed status/framing and forbid an invented second status or silent failover. +- Hardened readiness and fixture deadlines. Traffic admission requires a complete bounded `/readyz` HTTP/1.1 200 rather than bare TCP connect; repeated connect/write/read/retry loops consume absolute deadlines so slow-drip progress cannot renew fixture evidence windows indefinitely. +- Added dedicated routed pg-erd graceful-drain acceptance with one SIGTERM-relative external deadline and held-request completion during the shared drain grace period. +- Added pg-erd upstream TLS acceptance using an ephemeral local CA and `backend.test` certificate. Matching CA/SNI must succeed; the same CA with mismatched SNI must fail closed as HTTP 502 without route failover while readiness and an independent clear-text route remain usable. +- Added controlled-loopback pg-erd routed load acceptance with bounded Rust origins, exact route/body checks, zero HTTP failures, per-route sample floors, and aggregate/backend/frontend p95 below 20 ms. This remains regression evidence rather than a production SLO. +- Added non-root/read-only-root OCI packaging with a fail-closed build-time process allowlist for the generic and pg-erd binaries. Exact candidate acceptance requires uid/gid 65532, dropped capabilities, `no-new-privileges`, read-only configuration, process health, and the separate pg-erd Prometheus listener identity. +- Added committed dependency lock, license/source/advisory policy, SBOM and image-vulnerability evidence, exact-source candidate identity, and missing-public-rustdoc/documentation builds with warnings denied. +- Added an exact-head owned-production coverage gate requiring 100% lines and regions without filename/function/branch exclusions; coverage findings are repaired causally rather than excluded or threshold-weakened. +- Added versioned pg-erd response-body lifetime control. Admin Config version 2 requires a positive explicit `max_upstream_response_body_ms`; version 1 rejects that field and retains its previous semantics. +- Added `ResponseBodyLifetimeBudget` to Runtime Isolation. It starts at the first non-informational upstream response header, is not reset by later final-header callbacks, ignores empty/end-of-stream bookkeeping, and rejects non-empty body progress at or beyond the configured monotonic lifetime. +- Kept the response-lifetime control separate from Pingora `read_ms`. `read_ms` remains a per-read inactivity timer; continuously progressing body traffic is stopped at the first non-empty body callback after the lifetime expires. The current callback surface does not claim an exact interrupt for a pending read or a deadline for incomplete response headers. +- Made response-lifetime failure phase-aware. Pre-commit upstream failures preserve the existing policy-complete local error response, while `Session::response_written()` suppresses any second local status after a final response is already committed. +- Added realistic pg-erd slow-drip response traffic: the origin commits HTTP 200 framing and sends body bytes frequently enough to stay within each `read_ms`, while the shorter version-2 lifetime must terminate the incomplete response without a second status or failover, record exact request-error telemetry, preserve `/readyz`, and allow independent-route recovery. +- Added DDD, architecture, security, threat, test, operability, configuration, ADR, migration-gap, and primary-source traceability documentation. Dedicated repository-wide product-gap baseline and TRACEABILITY authority remain with their owner lane rather than being copied into migration descendants. -Release remains blocked on a maintainer-integrated and release-qualified disposition of unmaintained `derivative 2.2.0` / `RUSTSEC-2024-0388`, the exact Pingora supplier/protocol gates tracked by the foundation stack, the Rust compiler promotion owned by #56, terminal exact-current CI/supply-chain/security/review evidence for the current TLS child and later descendants, central required-workflow convergence and independent approval on protected promotion, representative pg-erd routed TLS/concurrency/origin-capacity/network-failure/drain and benchmark evidence, an immutable package/image identity with SBOM/provenance/reproducibility and rehearsed rollback, and protected-branch integration. The ordinary pg-erd protocol-transition chain is integrated through #33; the current pg-erd TLS child must independently reacquire exact-head hosted/current-range evidence after reconciliation. No consumer migration, shadow/canary, cutover, rollback, or legacy removal is claimed before those release and traffic-contract gates are satisfied. \ No newline at end of file +Release remains blocked on maintainer-integrated and release-qualified disposition of the supplier `derivative 2.2.0` / `RUSTSEC-2024-0388` admission failure, terminal exact-current CI/Supply/Security/CodeQL/current-review evidence for the active ancestry, the compiler/supplier promotion chain, representative routed TLS/origin-capacity/failure performance, protected-source reproducibility and provenance, immutable package/image identity, rollback rehearsal, and protected integration. Consumer parity, shadow/canary, cutover, rollback, and verified legacy reverse-proxy removal are not claimed before those gates are satisfied. diff --git a/OPERABILITY.md b/OPERABILITY.md index 8df42111..55abebcf 100644 --- a/OPERABILITY.md +++ b/OPERABILITY.md @@ -6,23 +6,29 @@ Run the generic process as `cwl-pingora-gateway --config /path/to/gateway.yaml`. The pg-erd migration candidate is a separate executable: `cwl-pingora-pg-erd-migration --config /path/to/pg-erd-migration.yaml`. It consumes the bounded `PgErdMigrationConfig` profile rather than widening generic `GatewayConfig` v1. The profile admits exactly the compiled `backend` and `frontend` transport identities plus deployment-variable sockets, runtime budgets and upstream transport/TLS values. Route tables, response-policy fields, product authentication/business rules, Keyverse identity, Wardnet/EgressWeave verdicts, service discovery and arbitrary destinations are not operator-configurable. +Pg-erd configuration version 1 preserves the existing unreleased characterization behavior and rejects `max_upstream_response_body_ms`. Version 2 requires an explicit positive `max_upstream_response_body_ms`; zero or omission fails before listener activation. Choose the value from observed application long-response requirements before canary or cutover. Example values are not production SLOs. + Admin parsing is side-effect free with respect to custom trust bytes. It validates the exact transport-authority set and `UpstreamConfig` invariants first. `build_proxy` then materializes Pingora peers and any custom PEM trust bundle once, still before listener registration. This avoids reading mutable trust material in a validation pass and reading it again for activation. Trust bundles are deployment input, not certificate-authority ownership. Mount them read-only from the platform or canonical secret/certificate owner and rotate them by replacing the deployment revision. The gateway does not issue certificates, manage ACME, or write trust material. -## Health and backpressure +## Health, backpressure, and response lifetime Both process identities reserve `/livez` and `/readyz` and return 200 with `Cache-Control: no-store` through the Pingora serving path. Readiness is process/configuration readiness, not upstream reachability. Do not use it as proof that a dependent application is healthy. For the pg-erd migration, consumer `/healthz` remains routed application traffic to `backend`; it is intentionally distinct from the gateway-local probes. `max_in_flight_requests` limits concurrently admitted non-health requests for one gateway process. At capacity the gateway fails new application traffic fast with HTTP 503 and increments `cwl_pingora_gateway_backpressure_rejections_total`; it does not queue unbounded work. Process health probes bypass that admission budget so operators can distinguish process health from traffic saturation. The request lease is released when the Pingora request context ends, including error paths, and a subsequent request is admissible again. +For pg-erd version 2, `max_upstream_response_body_ms` begins at the first non-informational upstream response header. Runtime Isolation checks elapsed monotonic time only on non-empty upstream body progress. `read_ms` remains a separate per-read inactivity timeout and resets after successful reads. A continuously progressing body is stopped at the first non-empty body callback at or beyond the configured lifetime; a quiescent response remains bounded by `read_ms`. + +The response-lifetime guard is not an exact wall-clock interrupt: the pinned Pingora callback surface does not wake an already-pending read only because the lifetime instant elapsed, and slow delivery of an incomplete response header is a separate gap. When a final response is already committed, `fail_to_proxy` observes `Session::response_written()` and does not emit a second status. Pre-commit upstream failures continue to use the existing policy-complete local error response. No route failover is introduced by the lifetime contract. + `upstream_keepalive_pool_size` is copied into Pingora `ServerConf` before bootstrap. Choose it with expected upstream concurrency, origin capacity, instance count and connection reuse in mind. It limits retained reusable upstream connections; it is not a substitute for the downstream in-flight admission limit and does not create product-domain load-balancing semantics. ## Forwarding and protocol boundary The generic v1 adapter and the pg-erd migration adapter have different compatibility forwarding contracts. Generic v1 strips inbound proxy-identity fields and emits only `Forwarded: proto=http`. The pg-erd migration adapter removes request-controlled `Forwarded`, `X-Forwarded-*`, `X-Real-IP` and legacy `X-Forwarded-Server` authority. It rebuilds `X-Forwarded-For` and `X-Real-IP` from the accepted client socket, preserves the original Host authority in `X-Forwarded-Host`, derives `X-Forwarded-Port` from the explicit Host port or the admitted scheme default, and emits the characterized `X-Forwarded-Proto`. Do not substitute the process listener bind port for the original authority: Kubernetes Services, containers, NAT and port publishing may expose a different external port. Product identity, tenant identity and authorization are never derived from these transport fields by the shared gateway. -The captured pg-erd Traefik entryPoint is clear-text `web`, so this migration profile currently uses downstream scheme `http`. Do not deploy it behind a TLS listener and assume `X-Forwarded-Proto: https` parity. Downstream TLS, HTTP/2, HTTP/3, WebSocket/upgrade and streaming behavior require separate executable contracts before they become admitted migration behavior. +The captured pg-erd Traefik entryPoint is clear-text `web`, so this migration profile currently uses downstream scheme `http`. Do not deploy it behind a TLS listener and assume `X-Forwarded-Proto: https` parity. Downstream TLS, HTTP/2, HTTP/3, WebSocket/upgrade and broader streaming behavior require separate executable contracts before they become admitted migration behavior. ## Logging @@ -36,7 +42,7 @@ The runtime does not inherit Pingora's retry, keepalive-pool, or drain defaults. SIGTERM uses Pingora graceful termination with an explicit 5-second request-drain grace period and a 10-second runtime-shutdown timeout. The pinned Pingora server calls Tokio `Runtime::shutdown_timeout` with that timeout and then sleeps for the same timeout while service runtimes are shut down in parallel. The policy therefore requires a 30-second supervisor hard-kill budget: its modeled worst-case Pingora process budget is 25 seconds plus scheduler/process-exit overhead. A Kubernetes-style deployment must set `terminationGracePeriodSeconds` to at least 30 or provide an equivalent supervisor budget; a shorter external kill deadline is not an admitted deployment contract. -`tests/graceful_shutdown.rs` exercises the generic compiled binary with a held upstream response. `tests/production_path.rs` covers generic saturation, health and failure recovery. `tests/pg_erd_production_path.rs` exercises the dedicated pg-erd process with real loopback backend/frontend origins, including process-local health, characterized route/response-header behavior, Host-authority forwarding replacement and declared body rejection. `tests/pg_erd_binary_startup.rs` requires missing/invalid configuration and unreadable trust material to fail before listener activation. `tests/pingora_diagnostic_log_safety.rs` runs the compiled generic process with broad trace diagnostics, proves the secret-bearing URI/Host/Authorization/Cookie request reaches the origin, requires a new Pingora redaction marker after the readiness probes, and requires none of those sentinels in process stderr. These source contracts become evidence only after terminal success on the exact current head; predecessor success never transfers. +`tests/graceful_shutdown.rs` exercises the generic compiled binary with a held upstream response. `tests/production_path.rs` covers generic saturation, health and failure recovery. `tests/pg_erd_production_path.rs` exercises the dedicated pg-erd process with real loopback backend/frontend origins, including process-local health, characterized route/response-header behavior, Host-authority forwarding replacement and declared body rejection. `tests/pg_erd_slow_drip_response_traffic.rs` separately exercises version-2 continuously progressing response bodies that stay inside each `read_ms` interval but cross the explicit lifetime; the downstream must keep its committed status/framing, terminate incomplete, avoid route failover or a second status, preserve `/readyz`, record the bounded error telemetry, and allow independent-route recovery. `tests/pg_erd_binary_startup.rs` requires missing/invalid configuration and unreadable trust material to fail before listener activation. `tests/pingora_diagnostic_log_safety.rs` runs the compiled generic process with broad trace diagnostics, proves the secret-bearing URI/Host/Authorization/Cookie request reaches the origin, requires a new Pingora redaction marker after the readiness probes, and requires none of those sentinels in process stderr. These source contracts become evidence only after terminal success on the exact current head; predecessor success never transfers. ## Container diff --git a/TEST_STRATEGY.md b/TEST_STRATEGY.md index f7678ea5..0a885013 100644 --- a/TEST_STRATEGY.md +++ b/TEST_STRATEGY.md @@ -1,53 +1,59 @@ # Test Strategy -Tests are organized by responsibility rather than by implementation layer. +Tests are organized by responsibility and evidence boundary. A source test is not promotion evidence until the unchanged exact head passes the applicable hosted gates; parent or historical receipts never transfer to a changed child. -`tests/config_contract.rs` proves strict parsing, mandatory positive request/concurrency/keepalive budgets, and baseline Edge Contract invariants. `tests/trust_bundle_contract.rs` proves the additional TLS trust-path invariants. `tests/tls_sni_hostname_contract.rs` freezes RFC 6066 SNI HostName admission before network authority: literal IP addresses, trailing-dot names, raw Unicode, malformed or overlong DNS labels, and overlong names fail closed while ordinary ASCII DNS names and IDNA A-labels remain admitted. `tests/startup_contract.rs` proves explicit fail-closed startup. `tests/pingora_peer_adapter.rs` proves TLS identity, explicit trust-bundle loading failures, timeout, and Pingora hop-by-hop policy mapping. `tests/gateway_proxy.rs` proves admitted config produces the expected peer, missing trust material blocks activation, and the process-local admission budget fails closed and releases capacity. `tests/runtime_policy.rs` locks the no-retry, configured upstream keepalive-pool, and bounded-shutdown process policy. `tests/binary_startup.rs` and `tests/trust_bundle_startup.rs` exercise generic compiled-process startup failures. `tests/production_path.rs` drives the generic compiled binary through a loopback Pingora listener and real local HTTP upstream fixture, covering health, declared and streamed/chunked 413 body-limit rejection, post-rejection readiness recovery, forwarding-identity sanitization, proxy response, and an executable saturation contract: with one in-flight lease held upstream, the next application request must fail fast with 503, `/readyz` must remain available, the rejection counter must advance, and a later request must succeed after the lease is released. Its traffic and metrics loopback sockets remain reserved through config construction and are released only at the child-bind handoff; traffic startup is admitted only after a bounded `/readyz` HTTP/1.1 200 carrying `Cache-Control: no-store`, so a bare TCP accept or a concurrently recycled ephemeral port cannot manufacture readiness evidence. On Unix, `tests/local_ca_tls.rs` generates a short-lived CA/server certificate at test time, proves the compiled gateway can trust that explicitly mounted CA, then changes only SNI and requires hostname verification to fail closed. On Unix, `tests/graceful_shutdown.rs` holds a request in flight, sends SIGTERM, proves the response drains during the configured grace period, and requires clean process exit before the external supervisor hard-kill budget. +## Generic gateway contracts -Migration characterization is executable before transport activation. `tests/pg_erd_route_contract.rs` freezes the observed `pg-erd-cloud` Traefik path precedence, including its literal raw `/api` prefix behavior. `tests/pg_erd_http_policy_contract.rs` freezes the separately owned response-security middleware fields and values, case-insensitive field identity, absent lookup behavior, duplicate-field rejection, empty/invalid name and value rejection, RFC 9110 control-octet and boundary-whitespace rejection, and valid interior SP/HTAB preservation. `tests/pg_erd_migration_plan_contract.rs`, `tests/pg_erd_upstream_binding_contract.rs`, `tests/pg_erd_forwarding_contract.rs`, and `tests/pg_erd_runtime_proxy_contract.rs` then prove the transport-neutral plan, exact upstream-authority binding, forwarding-trust boundary, response-policy composition, failure-response policy, and shared runtime-isolation/observability callback behavior. Characterization alone is not listener parity evidence. +`tests/config_contract.rs`, `tests/trust_bundle_contract.rs`, `tests/tls_sni_hostname_contract.rs`, `tests/startup_contract.rs`, `tests/pingora_peer_adapter.rs`, `tests/gateway_proxy.rs`, and `tests/runtime_policy.rs` cover strict Admin Config, non-zero and non-overlapping network authority, RFC 6066 SNI hostname admission, custom trust material, fail-closed startup, Pingora peer mapping, request/in-flight/keepalive budgets, one-attempt policy, and bounded shutdown. -The bounded Admin Config transition has its own executable contract. `tests/pg_erd_admin_config_contract.rs` rejects unknown/future configuration, listener collision, zero runtime/keepalive budgets, missing/extra/duplicate/renamed authority, and invalid concrete transport configuration; it also proves only `backend` and `frontend` can bind the compiled route profile. `tests/listener_authority_contract.rs` separately freezes generic zero-port and effective traffic/metrics socket-authority overlap, including native/IPv4-mapped and wildcard aliases. `tests/pg_erd_binary_startup.rs` exercises the dedicated compiled process and requires fail-closed behavior for omitted configuration, unreadable configuration, invalid Admin Config, and custom TLS trust material that cannot be materialized before listener activation. +`tests/binary_startup.rs`, `tests/trust_bundle_startup.rs`, `tests/production_path.rs`, `tests/local_ca_tls.rs`, and `tests/graceful_shutdown.rs` exercise the compiled generic process. Production-path readiness requires a complete bounded `/readyz` HTTP/1.1 200 carrying `Cache-Control: no-store`, not a bare TCP accept. Traffic and metrics loopback reservations are retained through config construction and released only at child-bind handoff so port reuse cannot manufacture readiness evidence. -`tests/pg_erd_local_ca_tls.rs` separately proves the bounded migration composition root consumes its own admitted upstream TLS contract rather than inheriting generic evidence. A one-day local CA and `backend.test` certificate are generated at test time. Matching explicit CA plus SNI must carry characterized `/api/tls` traffic to the TLS backend and preserve a distinct clear-text fallback route. With the same valid CA but mismatched SNI, the backend route must fail as exact HTTP 502 without route failover, `/readyz` must remain 200, the TLS backend must receive no HTTP request bytes, and a later independent fallback request must succeed. Traffic/metrics ports remain simultaneously reserved until startup; accepted origin sockets use five-second I/O deadlines; request headers are capped at 64 KiB; exact case-sensitive HTTP/1.1 three-digit status parsing rejects `2000`/protocol-case false GREEN. This is gateway-to-upstream TLS trust/hostname evidence only, not downstream TLS termination, certificate lifecycle ownership or representative TLS performance. +The generic load lane builds the release-mode gateway and bounded Rust origin in `tests/load/load_origin.rs`, format-checks and directly tests that origin, then uses checksum-pinned k6. `tests/load/gateway_smoke.js` requires exact status/body preservation, zero HTTP failures, and controlled-loopback p95 below 20 ms. This is a regression bound, not a production SLO. -`tests/network_authority_self_loop_contract.rs` covers the gateway-owned authority invariant shared by generic and pg-erd configuration: an upstream may not alias either the public traffic listener or the internal metrics listener through exact, wildcard, dual-stack, or mapped/native socket equivalence. `tests/upstream_unicast_authority_contract.rs` separately requires approved upstream TCP authority to remain a unicast destination: IPv4 limited broadcast, IPv4/IPv6 multicast, and IPv4-mapped broadcast/multicast forms fail closed after canonicalization. Distinct concrete non-aliased unicast IP authorities on the same port remain admitted. Public construction paths are revalidated after direct deserialization, and both compiled composition roots must fail closed before listener activation on recursive authority. +## Characterized pg-erd migration contracts -`tests/protocol_transition_policy.rs` freezes the transport-neutral HTTP/1 transition admission rule. Either an `Upgrade` field or a case-insensitive comma-delimited `upgrade` token in any `Connection` field value is an uncharacterized protocol-transition attempt; ordinary HTTP and unrelated tokens such as `x-upgrade` remain admitted. Both generic and pg-erd request filters invoke the shared rejection before request admission and route/upstream selection, so the current v1 behavior is deliberate HTTP 501 non-support rather than inherited supplier WebSocket semantics. +`tests/pg_erd_route_contract.rs` freezes the observed `pg-erd-cloud` route precedence, including literal raw `/api` prefix behavior. `tests/pg_erd_http_policy_contract.rs` freezes the separately owned response-security fields and RFC 9110-compatible header admission. `tests/pg_erd_migration_plan_contract.rs`, `tests/pg_erd_upstream_binding_contract.rs`, `tests/pg_erd_forwarding_contract.rs`, and `tests/pg_erd_runtime_proxy_contract.rs` prove the transport-neutral route/policy composition, exact upstream-authority binding, forwarding-trust boundary, failure-response policy, runtime isolation, and shared observability callbacks. Characterization alone is not listener parity evidence. -`tests/peer_protocol_policy.rs` locks the same non-support invariant at the immutable Pingora transport boundary. A peer returned by public `build_peer()` must use `HttpUpstreamRequestPolicy::deny_upgrades()` rather than the supplier `standard()`/`WebSocketOnly` default. This is a separate defense-in-depth oracle beneath callback admission; it must not be replaced by a mocked callback-only test or interpreted as WebSocket parity. +`tests/pg_erd_admin_config_contract.rs` proves fail-closed bounded Admin Config. Unknown or future versions, listener collision, zero runtime/keepalive budgets, missing/extra/duplicate/renamed upstream authorities, and invalid concrete transport/TLS data are rejected. Only `backend` and `frontend` may bind the compiled route profile. `tests/listener_authority_contract.rs`, `tests/network_authority_self_loop_contract.rs`, and `tests/upstream_unicast_authority_contract.rs` freeze zero-port, wildcard/native/mapped alias, recursive self-loop, broadcast, and multicast rejection while retaining distinct concrete unicast authority where valid. `tests/pg_erd_binary_startup.rs` requires invalid configuration or unusable trust material to fail before listener activation. -`tests/protocol_transition_traffic.rs` verifies that boundary through both compiled composition roots. WebSocket-shaped HTTP/1.1 requests must receive exact status 501, each relevant origin listener must observe no connection during a fixed 500 ms window after the response, and `/readyz` must remain HTTP 200. Traffic and metrics listener reservations remain distinct through config construction and the two reservation-to-child-bind handoffs are serialized inside the test binary; response-header reads are bounded to five seconds and 64 KiB, and status parsing accepts only case-sensitive `HTTP/1.1` with one exact three-digit status token. The 500 ms observation window is bounded traffic evidence, not proof against arbitrarily delayed contact. WebSocket enablement, HTTP/2 Extended CONNECT, and HTTP/3/QUIC require separate versioned contracts and realistic tunnel/concurrency/disconnect/backpressure/drain evidence. +`tests/pg_erd_local_ca_tls.rs` proves the bounded migration composition root consumes its own admitted upstream TLS contract. A short-lived local CA and `backend.test` certificate must succeed only with matching CA/SNI; a mismatched SNI with the same CA must fail as exact HTTP 502 without route failover, preserve `/readyz`, emit no HTTP request bytes to the rejected TLS backend, and leave an independent clear-text fallback route usable. This is upstream TLS trust/hostname evidence only, not downstream TLS termination or certificate-lifecycle authority. -`tests/pg_erd_production_path.rs` is the first dedicated compiled-listener traffic contract. It starts real loopback `backend` and `frontend` origins plus `cwl-pingora-pg-erd-migration`, requires `/livez` and `/readyz` to remain gateway-local, requires consumer `/healthz` and raw `/apiary` to reach `backend`, requires fallback product traffic to reach `frontend`, requires request-controlled forwarding identity to be discarded and rebuilt from accepted loopback client transport plus Host/scheme authority, requires upstream `X-Frame-Options: SAMEORIGIN` to be replaced by the characterized `DENY` policy together with the other three captured response fields, and requires an over-limit declared body to fail with 413 before origin delivery. Presence of this source test is not GREEN evidence: it must compile and pass with formatting, Clippy, rustdoc, and 100% owned-production line/region coverage on the same exact head before listener parity is credited. +## Protocol-transition boundary -`tests/pg_erd_runtime_isolation_traffic.rs` extends the dedicated compiled-listener path rather than borrowing generic-binary evidence. One test streams a chunked request beyond the eight-byte body budget and requires HTTP 413 while `/readyz` remains available. A second holds the sole admitted backend request with `max_in_flight_requests=1`, requires the next routed application request to return HTTP 503 in less than one second—materially before the configured two-second upstream read budget—requires `/readyz` to remain available and the exact Prometheus sample `cwl_pingora_gateway_backpressure_rejections_total 1`, then releases the held request and requires a later routed request to succeed. Exact-line matching prevents larger counter values from false-passing the single-rejection contract. These contracts count only when the unchanged exact head executes them to terminal GREEN. +`tests/protocol_transition_policy.rs` treats either an `Upgrade` field or a case-insensitive comma-delimited `upgrade` token in `Connection` as an uncharacterized HTTP/1 protocol-transition attempt. Ordinary HTTP and unrelated tokens such as `x-upgrade` remain admitted. Both generic and pg-erd request filters reject the transition with HTTP 501 before request admission or route/upstream selection. -`tests/pg_erd_upstream_failure_traffic.rs` adds the next distinct failure phase through the dedicated compiled process. On Linux it binds the characterized backend TCP address without calling `listen(2)`, so the test retains exclusive port ownership while connection attempts receive `ECONNREFUSED`; a direct `TcpStream::connect_timeout` precondition must observe `ConnectionRefused` before gateway traffic begins. If another process has stolen the selected port, fixture setup fails instead of allowing false-GREEN evidence. The migration gateway must then return HTTP 502 within a conservative one-second outer envelope around the configured 200 ms connection / 400 ms total-connection budgets, keep `/readyz` at HTTP 200, expose the exact Prometheus sample `cwl_pingora_gateway_request_errors_total 1`, and route an independent fallback request successfully to `frontend`. Exact-line matching prevents values such as `10` or `11` from false-passing the single-error contract. This contract does not claim connected read-stall, TCP reset, post-commit truncation, slow-drip/whole-response lifetime, retry, or failover behavior. +`tests/peer_protocol_policy.rs` locks the same boundary at immutable peer construction through `HttpUpstreamRequestPolicy::deny_upgrades()`, rather than inheriting the supplier `WebSocketOnly` default. `tests/protocol_transition_traffic.rs` verifies both compiled roots return exact 501, make no dedicated-origin connection during a bounded observation window, and keep `/readyz` available. This is explicit non-support evidence, not WebSocket parity; WebSocket enablement, HTTP/2 Extended CONNECT, and HTTP/3/QUIC require separate versioned contracts and realistic tunnel/concurrency/disconnect/backpressure/drain tests. -`tests/pg_erd_read_stall_traffic.rs` separates connected upstream inactivity from refusal. The characterized backend accepts `/api/read-stall`, reads request headers under a five-second/64 KiB fixture bound, records that causal point, then remains connected and sends no response bytes until the test explicitly releases it after the gateway has already returned. Traffic and metrics loopback sockets remain reserved through config construction and are released only at the child-bind handoff. With `read_ms=100`, the gateway must fail as HTTP 502 no earlier than a conservative 50 ms lower bound measured from completed origin request-header receipt and still inside the one-second outer envelope, preserve `/readyz`, expose the exact Prometheus sample `cwl_pingora_gateway_request_errors_total 1`, and leave an independent `frontend` route usable. Keeping the fixture connection open prevents origin closure from masquerading as the timeout, while the lower bound prevents an unrelated immediate 502 from masquerading as the configured read-inactivity path. Because Pingora's `read_timeout` is per successful `read()` rather than a whole-response lifetime, reset, post-commit partial response, slow-drip and whole-response deadline behavior remain separate contracts. +## pg-erd production and failure traffic -`tests/pg_erd_upstream_reset_traffic.rs` covers the distinct established-connection pre-header RST phase on Linux. Traffic and metrics sockets remain reserved through config construction and are released only at child-bind handoff. Before reset traffic begins, the traffic listener must answer a complete `/readyz` HTTP/1.1 200 under bounded connect/read/write timeouts and a 64 KiB response-header ceiling; a bare TCP accept cannot satisfy application readiness. The backend then receives complete `/api/reset` headers under a five-second/64 KiB bound and only afterward applies abortive `SO_LINGER(0)` without sending response bytes. Acceptance requires exact HTTP/1.1 502 in under two seconds despite a configured five-second read inactivity budget, exact `cwl_pingora_gateway_request_errors_total 1`, preserved `/readyz`, no silent frontend failover, and an independent frontend HTTP 200 recovery. Exact status/sample oracles reject protocol-case, numeric-prefix, and counter-prefix lookalikes. This contract does not claim post-commit reset, upgraded-stream failure, slow-drip/whole-response lifetime, retry, or failover behavior and counts only after terminal evidence on the unchanged exact head. +`tests/pg_erd_production_path.rs` starts real loopback `backend` and `frontend` origins with `cwl-pingora-pg-erd-migration`. It requires gateway-local `/livez` and `/readyz`, routed consumer `/healthz`, raw-prefix `/api` behavior, fallback routing, hostile forwarded-header replacement from accepted transport/request authority, exact characterized response-policy replacement, and pre-origin declared-body rejection. -`tests/pg_erd_post_commit_reset_traffic.rs` covers the distinct established-connection reset after the origin has emitted a valid HTTP/1.1 response header and partial body on Linux. Traffic and metrics sockets remain reserved through config construction and are released only at child-bind handoff. Process admission requires a complete bounded `/readyz` HTTP/1.1 200 rather than a bare TCP handshake; the metrics listener is only treated as bound at startup and is later revalidated through the real `/metrics` service. The backend receives `/api/post-commit-reset` headers under one absolute five-second/64 KiB fixture bound: every repeated origin-side read derives its socket timeout from the remaining `Instant` deadline, so successful partial reads cannot renew the budget. It emits exact HTTP/1.1 200 with one `Content-Length: 20` field plus the seven-byte `partial` prefix. A descendant coverage run exposed that making downstream header delivery the RST-release authority can still form a timing-dependent cycle: the origin waits for downstream forwarding while an intermediary may defer forwarding until the origin terminates. The owner fixture now breaks that cycle at the transport boundary. After `write_all`, it bounded-polls Linux `SIOCOUTQ`; Linux defines this queue value from `write_seq - snd_una`, so zero proves that the peer TCP stack has acknowledged every response byte emitted before the abort. Only then does the origin apply abortive `SO_LINGER(0)`. This transport acknowledgement is not claimed as proof that Pingora application code has parsed or forwarded the header; downstream acceptance remains the independent oracle for that behavior. Downstream termination is also read under one absolute five-second/64 KiB evidence bound rather than a renewable per-read timeout. Dedicated slow-drip oracles exercise both origin-header receipt and downstream termination to prove partial progress cannot extend either total budget. Acceptance requires the first downstream response to contain exact committed HTTP/1.1 200 framing with exactly one `Content-Length: 20`, permits only an exact prefix of `partial` if any body bytes cross before the abort, requires termination before all 20 declared bytes arrive, forbids a second status or silent failover, requires exact `cwl_pingora_gateway_request_errors_total 1`, preserves `/readyz`, and proves independent frontend HTTP 200 recovery. Status, framing and metric oracles reject protocol-case, field-name and numeric-prefix lookalikes. This contract counts only after terminal evidence on the unchanged exact head. +`tests/pg_erd_runtime_isolation_traffic.rs` covers streamed request-body overflow plus process-local in-flight saturation/recovery and exact backpressure telemetry. `tests/pg_erd_upstream_failure_traffic.rs` creates deterministic `ECONNREFUSED`; `tests/pg_erd_read_stall_traffic.rs` keeps an accepted origin connection open without response bytes so per-read inactivity cannot be confused with origin closure. Both require bounded failure, exact error telemetry, preserved readiness, no silent route failover, and later independent-route recovery. -On Unix, `tests/pg_erd_graceful_shutdown.rs` proves the bounded pg-erd composition root consumes the shared drain policy instead of borrowing generic-binary evidence. Traffic and metrics listener reservations survive config construction and are released only at child-bind handoff; readiness requires a bounded complete `/readyz` HTTP/1.1 200. The characterized backend accepts `/api/held`, and request-header receipt itself is bounded to five seconds and 64 KiB before the fixture declares the request in flight. Only then does the test send SIGTERM, release the held response during the shared grace period, require downstream HTTP 200 completion, and require successful migration-process exit before the single absolute external termination deadline anchored at signal delivery. This contract counts only when the unchanged exact head executes it to terminal GREEN. +On Linux, `tests/pg_erd_upstream_reset_traffic.rs` covers abortive reset after request headers but before any response header. `tests/pg_erd_post_commit_reset_traffic.rs` covers abortive reset after HTTP 200 framing and a partial body have crossed the transport boundary. The latter uses bounded origin/downstream deadlines and Linux `SIOCOUTQ` to break timing cycles without pretending TCP acknowledgement proves Pingora application parsing. Acceptance preserves committed status/framing, terminates before the declared body completes, forbids a fabricated second status or silent failover, records exactly one request error, keeps readiness healthy, and proves independent-route recovery. -`tests/pg_erd_partial_response_traffic.rs` covers the distinct post-header failure phase. The backend sends exact HTTP/1.1 200 with one `Content-Length: 20` field and the body prefix `partial`, then remains open until the downstream reader has observed the complete response header block and exact prefix. Only after that acknowledgement does the fixture allow origin close, preventing an immediate FIN from accidentally exercising a pre-commit failure phase. Acceptance requires the committed response to terminate before all 20 bytes arrive, forbids a rewritten second status or silent failover, preserves `/readyz`, records exactly `cwl_pingora_gateway_request_errors_total 1`, and proves an independent `frontend` route still completes. The status oracle accepts only exact HTTP/1.1 plus a three-digit code, the framing oracle rejects `X-Content-Length` and duplicate/conflicting Content-Length fields, origin request-header reads are bounded by five seconds and 64 KiB, and traffic/metrics reservations remain held until child-bind handoff. This contract counts only when the unchanged exact head reaches terminal hosted GREEN. +`tests/pg_erd_partial_response_traffic.rs` covers orderly close after a committed HTTP 200 and partial body. Exact status, `Content-Length`, and Prometheus sample parsing reject protocol, field-name, duplicate-framing, and numeric-prefix lookalikes. `tests/pg_erd_graceful_shutdown.rs` holds a routed request in flight, sends SIGTERM only after bounded origin-header receipt, releases the response during the grace period, and requires downstream completion plus clean process exit inside one absolute external termination deadline. -`tests/pg_erd_payload_free_observability.rs` proves the shared gateway observability boundary through the compiled migration process without borrowing product logging authority. A real routed request carries unique URI/query, Host, Authorization, Cookie, and product-context sentinels; the backend must receive those exact values so the test cannot pass vacuously, while the `cwl_pingora_gateway::observability` target may emit only the bounded completion vocabulary and none of the sentinels. Traffic and metrics listener reservations remain distinct through config construction and are released only at child-bind handoff; process admission requires a complete `/readyz` HTTP/1.1 200 response under bounded read/write timeouts and a 64 KiB response-header ceiling rather than a bare TCP accept. Origin header reads are bounded by five seconds and 64 KiB, HTTP field identity is matched case-insensitively by exact name with OWS trimming, and the Prometheus oracle requires a complete exact sample line so a counter value such as `10` cannot satisfy an expected value of `1`. This proves only the shared gateway observability contract; it does not claim authority over product-owned logging, tracing, identity, or third-party logger configuration. +`tests/pg_erd_payload_free_observability.rs` sends unique URI/query, Host, Authorization, Cookie, and product-context sentinels through the compiled migration process. The backend must receive them so the test is non-vacuous, while shared gateway stderr and low-cardinality metrics must not expose those sentinels. `tests/pingora_diagnostic_log_safety.rs` separately covers broad Pingora-family diagnostics under `RUST_LOG=trace`, requiring a new redaction marker attributable to the characterized request and forbidding request-derived secrets in process stderr. -## Process-wide diagnostic logging +## Version-2 response-body lifetime -`tests/pingora_diagnostic_log_safety.rs` covers the dependency-log boundary that the shared `observability` target cannot prove. It starts the compiled generic runtime with `RUST_LOG=trace`, holds traffic and metrics reservations simultaneously until startup, and sends a request with unique URI/query, Host, Authorization and Cookie sentinels. The origin must receive the exact request target and semantically exact header values, with origin accept/header acquisition bounded to five seconds and 64 KiB, so a reject/strip path cannot manufacture GREEN. +`tests/pg_erd_response_lifetime_config.rs` freezes the version transition. Pg-erd version 2 requires a positive explicit `max_upstream_response_body_ms`; version 2 rejects zero or omission, and version 1 rejects the field so timing semantics cannot change silently. -Readiness probes occur before the characterized request and can themselves cause supplier diagnostics. The test therefore snapshots the static Pingora-redaction marker count only after both probes, requires the count to increase after the secret-bearing request reaches the origin and returns, waits for an exact bounded completion-log line suffix, and then requires none of the sentinels anywhere in captured process stderr. This proves that the characterized request traversed a redacted Pingora-family diagnostic path rather than merely observing an unrelated startup marker. Product/application loggers remain outside this process boundary. +Unit coverage in `src/runtime_isolation.rs` and `src/migration_proxy.rs` proves the monotonic lifetime begins only at the first non-informational upstream response header, is not reset by later final-header callbacks, ignores empty/end-of-stream bookkeeping, and rejects non-empty body progress at or beyond the configured limit. A lifetime error is upstream-scoped. `fail_to_proxy` must preserve the existing local error response for pre-commit failures, but once `Session::response_written()` reports a final response it must return error code 0 and emit no second status. -The `oci-runtime` job validates artifact composition separately from compiled-process tests. The Dockerfile admits only `cwl-pingora-gateway` and `cwl-pingora-pg-erd-migration` as build-time process identities, normalizes the selected executable to one fixed distroless runtime path, and CI builds both image profiles. Each exact candidate must declare uid/gid `65532`, start under a read-only root filesystem with all capabilities dropped and `no-new-privileges`, and consume only a read-only configuration mount. The generic profile must expose local `/livez`. The pg-erd profile must expose local `/livez` on the traffic listener and its separately published `/metrics` listener must identify the Prometheus service by exact base media type `text/plain` after stripping only optional semicolon parameters. `tests/pg_erd_oci_metrics_workflow_contract.rs` prevents a bare HTTP 200 or prefix wildcard from false-greening a mistakenly bound proxy service while deliberately avoiding a metric-family requirement before routed application traffic has emitted one. `examples/pg-erd-migration.yaml` is an origin-independent OCI smoke fixture, not routed parity evidence. The supply-chain job additionally builds and vulnerability-scans both profiles and binds both local image IDs plus per-image scan outputs to the exact source SHA. These workflow contracts count only after terminal success on the unchanged exact head. +`tests/pg_erd_slow_drip_response_traffic.rs` is the realistic wire oracle. The backend commits HTTP 200 with `Content-Length: 20` and emits body progress frequently enough that each read remains within `read_ms`, while a shorter `max_upstream_response_body_ms` expires. The downstream must retain the committed status/framing, terminate before all declared bytes arrive, receive no fabricated second status or silent failover, observe exact request-error telemetry, keep `/readyz` available, and allow a later independent frontend request. This proves continuous body slow-drip is bounded. It does not claim an exact timer interrupt for an already-pending read or an absolute deadline for incomplete response headers; those remain separate gaps. -`tests/load/gateway_smoke.js` is a separate concurrent traffic contract executed with checksum-pinned k6 2.2.0 against the release-mode generic gateway binary and the bounded Rust origin in `tests/load/load_origin.rs`. The load lane format-checks and directly tests the std-only origin, builds an optimized fixture, then sends 400 requests across four virtual users, requires every response to preserve the expected status/body, requires zero HTTP failures, and gates loopback `http_req_duration` p95 below 20 ms. `tests/load_origin_readiness_workflow_contract.rs` requires the measured origin to be ready before gateway startup. `tests/load_evidence_workflow_contract.rs` requires a successful k6 path to produce a nonempty summary while ensuring a pre-k6 failure cannot be replaced by a secondary missing-artifact failure. This is a regression bound for the minimal local generic path, not evidence that an Internet, TLS, multi-hop, consumer production path satisfies a 20 ms p95 SLO. +## Routed performance acceptance -`tests/load/pg_erd_gateway_smoke.js` adds the dedicated routed pg-erd controlled-loopback concurrency/latency contract. Separate bounded Rust origins represent the characterized `backend` and `frontend` authorities while route selection remains compiled into the migration plan. The test runs four VUs for 400 total iterations, alternates `/api/load-contract` and `/load-contract`, requires exact HTTP 200/body identity and zero HTTP failures, independently gates aggregate, backend and frontend p95 below 20 ms, and requires at least 198 measured requests per route. `tests/pg_erd_routed_latency_contract.rs` freezes route/path/body/tag bindings plus the per-route thresholds and sample floors, while `tests/rust_load_origin_workflow_contract.rs` prevents regression to interpreted measured-origin execution. `k6-pg-erd-summary.json` is exact-SHA evidence only after the unchanged child head executes the hosted load lane. This loopback gate is not representative TLS, multi-hop, container/Kubernetes scheduling, origin-capacity or production SLO evidence; sample reduction, route omission, threshold removal, or unrealistic warm-up is not an admissible performance repair. +`tests/load/pg_erd_gateway_smoke.js` and `tests/pg_erd_routed_latency_contract.rs` exercise characterized backend/frontend routes through bounded Rust origins. The measured lane requires exact status/body identity, zero HTTP failures, aggregate/backend/frontend p95 below 20 ms, and the configured minimum samples per route. `tests/rust_load_origin_workflow_contract.rs` prevents regression to interpreted measured-origin execution. Controlled loopback evidence is not representative TLS, multi-hop, container/Kubernetes scheduling, origin-capacity, or production-SLO evidence; sample reduction, route omission, threshold removal, or unrealistic warm-up is not an admissible repair. -Every behavioral migration should begin with characterization against the old owned edge behavior, then add equivalent production-path evidence for Pingora. Static-serving consumers must cover route precedence, SPA fallback, MIME, ETag/cache, Range/HEAD/304/416, redirects/security headers, and compression as applicable. Proxy consumers must cover Host/SNI/TLS, forwarding trust, WebSocket/upgrade, limits, timeout/retry behavior, streaming/uploads, saturation/backpressure, errors, health/readiness, and drain. Characterized response policies must additionally be exercised through the compiled proxy path before any parity or canary claim. +## OCI, supply chain, coverage, and release evidence -Release-quality gaps remain: no property/fuzz tests yet; no downstream TLS listener contract; no HTTP/2 or HTTP/3 parity evidence; no tracing evidence; no immutable published registry digest/provenance and rehearsed rollback; and no benchmark against replaced Nginx/Traefik traffic. For the pg-erd candidate specifically, dedicated source now covers streamed body overflow, in-flight saturation/recovery, refused-origin recovery, connected silent-origin read timeout, pre-header and post-commit TCP reset, routed graceful drain, orderly post-header truncation, routed controlled-loopback concurrency/latency, payload-free shared-observability behavior, fail-closed HTTP/1 transition rejection, plus OCI process/Prometheus-listener identity. The post-commit fixture now also proves its own origin/downstream evidence budgets cannot be extended by slow-drip partial progress; this is fixture-boundedness evidence, not yet a whole-response lifetime contract for arbitrary routed traffic. Pg-erd upstream TLS trust/hostname success/failure now has dedicated source acceptance but must close on the unchanged current exact head before it is credited. Broader streaming/upgraded failure, application-level slow-drip/whole-response lifetime, representative routed TLS/origin-capacity load, terminal current-head traffic/OCI/supply-chain execution, shadow/canary, and rollback still lack evidence. OCI non-root/read-only-root source acceptance covers both admitted process images, while owned-production 100% line/region coverage, public rustdoc, SBOM/image vulnerability, local-CA upstream TLS, explicit backpressure/recovery, and the k6 loopback paths remain exact-head gates. None can be transferred to a changed or consumer-specific head. Production performance claims remain forbidden until representative deployment measurements exist. \ No newline at end of file +The `oci-runtime` job builds the generic and pg-erd images separately. Both must declare uid/gid `65532`, run with a read-only root filesystem, all capabilities dropped, `no-new-privileges`, and a read-only versioned configuration mount. The generic image must expose `/livez`; pg-erd must expose its traffic `/livez` and a separately published Prometheus `/metrics` listener whose base media type is `text/plain`. + +Supply-chain evidence must remain exact-source-bound and include dependency audit, both candidate image builds, SPDX SBOM evidence, and image vulnerability scans. Owned production code requires 100% line and region coverage without exclusions, missing-public-rustdoc enforcement, warning-denied documentation, formatting, all-target compile/test, and strict Clippy. Release credit additionally requires immutable registry/package identity, signing/attestation/provenance, reproducibility evidence, and rollback rehearsal; local image IDs or Draft PR checks are insufficient. + +## Remaining gaps + +Open acceptance includes downstream TLS/H2, H2-to-H1 Cookie normalization, versioned WebSocket/Extended CONNECT, explicit H3/QUIC disposition, incomplete-response-header slow-drip/absolute header-read handling, broader admitted long-lived-stream semantics where consumer evidence requires them, dynamic reload, tracing, property/fuzz testing, representative routed TLS/origin-capacity load, shadow/canary, rollback, cutover, and verified legacy proxy removal. Every changed descendant must reacquire applicable exact-head evidence; neither source presence nor predecessor GREEN is release evidence. diff --git a/TRD.md b/TRD.md index 107cf254..ec70dcb4 100644 --- a/TRD.md +++ b/TRD.md @@ -1,58 +1,78 @@ -# Technical Requirements +# Technical Requirements Document -## Runtime +## Runtime and composition roots -This branch uses Rust edition 2021 with `rust-version = "1.98.0"`. The Pingora crates are pinned to Cloudflare Pingora `0.8.0` at exact upstream Git revision `09696b51bc59315353d96686355861604d0bb48c`; mutable branch, tag, or contributor-PR resolution is not release authority. +This branch uses Rust edition 2021 with manifest MSRV `1.98.0`. Cloudflare Pingora crates remain pinned to the exact upstream revision admitted by the parent stack; mutable branches, tags, contributor PRs, or local forks are not release authority. -There are two composition roots with deliberately different contracts: +There are two composition roots with intentionally different public contracts: -- `src/bin/cwl-pingora-gateway.rs` activates generic version-1 `GatewayConfig` and the one-upstream `GatewayProxy`. -- `src/bin/cwl-pingora-pg-erd-migration.rs` activates only the bounded `PgErdMigrationConfig` profile and `MigrationGatewayProxy` for the characterized pg-erd edge surface. +- `src/bin/cwl-pingora-gateway.rs` activates generic version-1 `GatewayConfig` and the single-upstream `GatewayProxy`. +- `src/bin/cwl-pingora-pg-erd-migration.rs` activates only bounded `PgErdMigrationConfig` and `MigrationGatewayProxy` for the characterized pg-erd edge surface. -Both parse an explicit `--config` path before creating listeners and delegate process lifecycle to Pingora's server lifecycle. The separate binaries prevent the generic v1 configuration language from being widened implicitly by one consumer migration. +Both parse and validate an explicit configuration before creating listeners and delegate serving/shutdown to Pingora. The pg-erd binary does not widen generic v1 into a product routing language. Product authentication/authorization, business logic, certificate issuance/ACME, Keyverse identity, Wardnet/EgressWeave policy, and consumer service discovery remain outside this runtime. -## Generic edge contract +## Network and Admin Config authority -Generic configuration version 1 is strict YAML with unknown fields denied. It admits one explicit non-zero listener, a distinct non-zero metrics authority, exactly one non-zero upstream socket, a positive request-body budget, positive process in-flight and upstream keepalive budgets, and explicit positive upstream I/O budgets. Listener/metrics validation rejects effective socket-authority overlap: equal sockets, same-family wildcard aliases, native/IPv4-mapped IPv4 aliases, and IPv6-wildcard/IPv4 same-port ambiguity. Distinct concrete non-aliased addresses remain independent. HTTPS upstreams require SNI and certificate/hostname verification; cleartext upstreams must not carry SNI or trust-bundle data. No request may select arbitrary upstream authority dynamically. +Generic v1 is strict YAML with unknown fields denied. It admits one explicit non-zero traffic listener, one distinct non-zero metrics listener, exactly one non-zero upstream authority, positive request-body/in-flight/keepalive budgets, and positive upstream I/O budgets. Listener and metrics validation rejects equal sockets, same-family wildcard/concrete aliases, native/IPv4-mapped IPv4 aliases, native or mapped wildcard aliases, and platform-dependent same-port IPv6-wildcard/IPv4 ambiguity while preserving distinct concrete non-aliased authority. -Requests with a parseable `Content-Length` above the configured limit fail with 413 before upstream selection. Streamed body bytes are counted against the same bound. Saturated process admission fails with 503 and the lease is released when the request completes or aborts. +TLS upstreams require an admitted RFC 6066 DNS hostname for SNI/hostname verification and may optionally consume one absolute PEM trust-bundle path before listeners open. Clear-text upstreams may define neither SNI nor a trust bundle. The gateway does not issue, renew, rotate, or persist trust material. -## Bounded pg-erd Admin Config +`PgErdMigrationConfig` is a bounded Admin Config contract, not a generic route DSL. Operators may provide only the traffic/metrics sockets, positive runtime and keepalive budgets, and concrete transport/TLS values for the compiled `backend` and `frontend` identities. Route precedence, response-security policy, product auth, business routing, and arbitrary destinations are not operator-configurable. Missing, duplicate, extra, renamed, zero-port, recursive, multicast/broadcast, or otherwise invalid transport authority fails before listener activation. -`PgErdMigrationConfig` is not a generic route language. Operators may provide only the traffic listener, metrics listener, positive body/in-flight/keepalive budgets, and concrete transport/TLS values for the already characterized `backend` and `frontend` identities. Route precedence, response-security fields and admitted upstream names remain compiled migration semantics. Product authentication/authorization, business routing, domain response semantics, Keyverse identity and Wardnet/EgressWeave verdicts remain outside this bounded context. +The type is publicly deserializable, so `build_proxy()` revalidates the complete deterministic contract before creating delivery peers or infallible runtime limits. Direct deserialization cannot bypass version, listener, runtime, keepalive, or transport-authority invariants. Custom trust bytes are materialized exactly once during peer construction after deterministic validation and before listener registration. -Configuration validation fails before listener activation on unsupported versions, zero ports or runtime budgets, overlapping traffic/metrics socket authority, zero keepalive capacity, missing/duplicate/extra/renamed upstream authority, or invalid upstream TLS/transport data. The migration profile consumes the same shared effective socket-authority invariant as generic v1 while preserving its characterized zero-transport-authority error surface. +## Version-2 response-body lifetime -`PgErdMigrationConfig` derives public Serde `Deserialize`, so callers are not forced to enter through `PgErdMigrationConfig::from_yaml`. The public activation boundary therefore revalidates the complete deterministic configuration in `build_proxy()` before delivery peers or runtime limits are materialized. Only after that check may `RuntimeIsolationLimits::from_validated` reuse the proven positive budgets. Direct deserialization cannot bypass version, listener-authority, runtime, keepalive or transport-authority invariants. +Pg-erd version 1 preserves the existing unreleased characterization semantics and rejects `max_upstream_response_body_ms`. Version 2 requires an explicit positive `max_upstream_response_body_ms`; omission or zero is invalid. This version boundary prevents a timing policy from appearing through a hidden default. -Custom upstream trust-bundle bytes are not preloaded during YAML parsing. Peer/trust materialization happens once during `build_proxy()` before the composition root creates listeners, reducing validate-then-reload drift for operator-supplied trust material. +`RuntimeIsolationLimits` carries the optional response-body lifetime only for the version-2 profile. `MigrationRequestContext` creates a `ResponseBodyLifetimeBudget` from those limits. `upstream_response_filter` starts the monotonic budget at the first non-informational upstream response header. Informational headers do not start it, and later final-header callbacks do not reset it. `upstream_response_body_filter` checks elapsed time only when a non-empty body chunk is observed; empty/end-of-stream bookkeeping cannot manufacture expiry. -## Request and forwarding policy +A body-progress callback at or beyond the configured lifetime becomes an upstream-scoped fatal error. The lifetime is deliberately independent of Pingora peer `read_ms`: `read_ms` remains a per-read inactivity timer that resets after successful reads, while the version-2 lifetime bounds continuously progressing bodies at their first non-empty callback after the absolute lifetime is reached. With the pinned callback surface this is not an exact interrupt of an already-pending read, and slow delivery of an incomplete response header remains a separate gap. -Every immutable Pingora upstream peer uses `HttpUpstreamRequestPolicy::deny_upgrades()`. This retains the pinned supplier's standard hop-by-hop and `Connection`-nomination sanitization but changes its HTTP/1 upgrade policy from the default `WebSocketOnly` behavior to `Deny`. The separate transport-neutral admission guard remains authoritative for returning HTTP 501 before application admission or origin selection. Keeping both boundaries aligned prevents callback/composition changes from implicitly enabling a supplier protocol capability that the versioned gateway contract does not admit. +Failure handling is response-phase-aware. Before a final downstream response has been written, existing `fail_to_proxy` behavior may still emit the policy-complete local error response for an upstream failure. After `Session::response_written()` reports a final response, the runtime returns error code 0 and writes no second status. A post-commit lifetime breach therefore terminates the incomplete response instead of rewriting it to 502 or routing to another pg-erd origin. The ordinary request context then releases its in-flight admission lease. -The generic gateway additionally removes request-controlled `Forwarded`, `X-Forwarded-For`, `X-Forwarded-Host`, `X-Forwarded-Port`, `X-Forwarded-Proto`, `X-Forwarded-Server`, and `X-Real-IP`, then emits only gateway-owned `Forwarded: proto=http` for its current cleartext downstream contract. Generic v1 deliberately makes no client-IP identity or downstream proxy-provenance claim. +## Request, forwarding, and protocol policy -The pg-erd migration callback uses the separate Ingress Forwarding Policy. Request-controlled `Forwarded`, `X-Forwarded-*`, `X-Real-IP` and legacy `X-Forwarded-Server` values are discarded. `X-Forwarded-For` and `X-Real-IP` are rebuilt from the accepted client socket, `X-Forwarded-Host` preserves original Host authority, and `X-Forwarded-Port` comes from an explicit Host port or the admitted scheme default rather than the process listener bind. The currently characterized legacy entry point is cleartext `web`, so downstream scheme is explicitly `http`; HTTPS forwarding semantics require a separate downstream-TLS contract. +Every immutable Pingora peer uses `HttpUpstreamRequestPolicy::deny_upgrades()`. A separate transport-neutral request guard returns HTTP 501 for an HTTP/1 request carrying either `Upgrade` or a case-insensitive `upgrade` token in `Connection`, before request admission or upstream selection. This is deliberate protocol non-support, not WebSocket parity. -Failure handling is phase-aware. Before an upstream response header is committed downstream, transport failure may still be represented by the gateway's fail-closed error response under the one-attempt policy. After a valid response header has been committed, a later upstream framing/body failure cannot be rewritten into a second HTTP status or silently failed over: the incomplete downstream response terminates, low-cardinality error telemetry records the failed request, process readiness remains available, and independent routes must remain usable. This is an edge transport invariant, not product retry authority. +Generic v1 strips request-controlled `Forwarded`, `X-Forwarded-*`, `X-Real-IP`, and related proxy-identity fields before emitting only gateway-owned `Forwarded: proto=http`. It makes no client-identity or downstream proxy-provenance assertion. -## Health, observability and graceful lifecycle +The pg-erd migration adapter also removes request-controlled forwarding identity. It rebuilds only the characterized compatibility fields from accepted client transport plus validated request authority: `X-Forwarded-For`, `X-Real-IP`, `X-Forwarded-Host`, `X-Forwarded-Port`, and `X-Forwarded-Proto`. The currently characterized consumer entry point is clear-text `web`, so the admitted downstream scheme is `http`; HTTPS forwarding semantics require a separate downstream-TLS contract. -`/livez` and `/readyz` are gateway process endpoints served locally through the production Pingora path and do not become consumer routes. Pg-erd `/healthz` remains characterized product traffic to `backend`. Shared observability is low-cardinality and payload-free: request path/query, headers, cookies, credentials, customer payloads and product identifiers are outside the shared telemetry contract. +Non-health traffic acquires the process `max_in_flight_requests` lease before upstream selection. Saturation fails fast with HTTP 503 and increments bounded telemetry. A declared `Content-Length` above `max_request_body_bytes` fails with HTTP 413 before origin selection, and streamed body bytes are counted against the same limit. `/livez` and `/readyz` bypass application admission so process health remains observable under saturation. -Both composition roots use the shared Pingora server policy and bounded graceful shutdown. Process tests terminate successful children through the graceful path so LLVM coverage profiles can flush; emergency cleanup remains a test-harness fallback rather than the normal lifecycle. +## Routing, response policy, and delivery -## Packaging and release boundary +`EdgeMigrationPlan` owns the characterized transport-neutral pg-erd route/policy composition. The route contract admits exact `/healthz -> backend`, raw `/api` prefix behavior including `/apiary -> backend`, and fallback `/ -> frontend` according to the captured consumer edge semantics. The response-policy contract owns only the characterized gateway response fields; it does not take product-domain response ownership. -The Docker builder is digest-pinned Rust 1.98.0 Bookworm and the final image is digest-pinned distroless Debian 13 `base-nossl` non-root. `CWL_GATEWAY_BIN` is a build-time-only fail-closed allowlist of exactly `cwl-pingora-gateway` and `cwl-pingora-pg-erd-migration`; the selected executable is normalized to one fixed runtime path, so the final image contains one admitted process identity and no runtime shell selector. +`MigrationDeliveryPlan` binds each admitted upstream identity to exactly one prevalidated Pingora `HttpPeer`. Missing, duplicate, undeclared, recursive, or invalid transport bindings fail closed. Request data cannot select an arbitrary destination or create service-discovery authority. -Exact-head OCI acceptance builds both profiles and starts each as uid/gid `65532` under read-only-root, all-capabilities-dropped and `no-new-privileges` restrictions with a read-only configuration mount. The supply-chain lane builds and vulnerability-scans both candidate images, binds both local image IDs and per-image scan outputs to the exact source SHA, and keeps failure diagnostics distinct from promotion-shaped success evidence. These are unreleased candidate receipts only. +The migration proxy applies response-security fields through replacement semantics. Upstream transport failures before response commitment use the bounded local error mapping; post-commit truncation, reset, or lifetime failure preserves the committed response rather than inventing a second status or silent failover. -A protected release remains blocked until exact-head CI, strict Clippy, warning-denied rustdoc, 100% owned production line/region coverage, load/runtime and supply-chain evidence are terminal GREEN; protected review/governance is satisfied without bypass; supplier/advisory policy is clean; and an immutable image digest, release-bound SBOM/provenance/reproducibility and rollback evidence exist. Source capability, predecessor GREEN or a mutable image/tag does not establish release, parity, shadow, canary, cutover or legacy-removal state. +## Health, observability, and graceful lifecycle -## Protocol and migration limits +`GET /livez` and `/readyz` are process-local and return non-cacheable HTTP 200 through Pingora. They indicate validated process/configuration readiness, not consumer dependency health. Pg-erd consumer `/healthz` remains routed application traffic to `backend`. -Generic v1 remains a cleartext downstream HTTP proxy with one explicit upstream per process. HTTP/1 Upgrade is explicitly denied both before request admission and at immutable peer construction; that is non-support evidence, not WebSocket parity. Downstream TLS termination, HTTP/2 admission, H2-to-H1 Cookie normalization, HTTP/3/QUIC, versioned WebSocket/Extended CONNECT, dynamic reload, Kubernetes Gateway API, and consumer-specific multi-route behavior are separate increments with realistic RED-to-GREEN evidence. +Shared observability is low-cardinality and payload-free. The gateway records bounded request completion/error/body-byte/backpressure facts but excludes request paths, query strings, credentials, cookies, customer payloads, product identifiers, and unbounded labels. Pingora-family dependency diagnostics pass through the process-wide payload-safe logger so broad `RUST_LOG` settings cannot bypass this boundary. -The pg-erd migration stack is a bounded consumer-characterization adapter and does not widen generic v1. Promotion still requires unchanged exact-head formatting, compile/test, strict Clippy, rustdoc, owned-production coverage, routed traffic/load/failure evidence, immutable release identity, consumer deployment pin, shadow/canary, rollback rehearsal, protected cutover, and verified legacy removal. +The shared server policy makes one total upstream attempt, configures the admitted keepalive pool, applies an explicit five-second drain grace period, and uses the bounded runtime shutdown policy documented in `OPERABILITY.md`. Consumer retry/failover requires idempotency and product knowledge and is not inferred by the gateway. + +## Packaging and supply chain + +The Docker build admits only `cwl-pingora-gateway` or `cwl-pingora-pg-erd-migration` as build-time process identities and normalizes the selected executable into one distroless non-root runtime image. Exact-head OCI acceptance requires uid/gid `65532`, read-only root, dropped capabilities, `no-new-privileges`, and read-only configuration mounts. The pg-erd image must expose the traffic health endpoint and the separately published Prometheus metrics listener. + +Supply-chain evidence remains exact-source-bound: committed lockfile, dependency/advisory policy, SBOM, both candidate image builds, and image vulnerability scans. These are candidate receipts, not immutable release identity. Protected promotion additionally requires immutable package/image digests, release-bound SBOM/provenance/attestation, reproducibility, rollback rehearsal, and current protected ancestry. + +## Test and performance requirements + +Every changed exact head must pass formatting, all-target compile/test, strict Clippy, warning-denied rustdoc, 100% owned-production line and region coverage without exclusions, realistic production-path traffic, OCI, supply-chain, and current-head review. Parent or historical GREEN never transfers after source, parent, or evidence changes. + +The version-2 lifetime path requires both focused unit/config coverage and real slow-drip traffic. The real fixture must keep each upstream read inside `read_ms` while total body progress crosses `max_upstream_response_body_ms`, preserve committed 200/framing, terminate before the declared body completes, emit no second status or route failover, record exact error telemetry, retain `/readyz`, and prove an independent route can recover. + +Applicable routed buyer paths use bounded Rust origins and k6/E2E with p95 below 20 ms on controlled loopback. Such loopback evidence is a regression gate only. Representative TLS, multi-hop, container/Kubernetes scheduling, origin-capacity, failure contention, and deployment traffic are required before production p95 credit; sample reduction, route omission, threshold weakening, or unrealistic warm-up are not repairs. + +## Known limits and migration boundary + +Generic v1 remains a clear-text downstream HTTP proxy with one explicit upstream per process. The pg-erd binary remains a bounded consumer-characterization adapter. Downstream TLS/H2, H2-to-H1 Cookie normalization, versioned WebSocket/Extended CONNECT, H3/QUIC, incomplete-response-header slow-drip, broader long-lived-stream semantics, dynamic reload, tracing, property/fuzz testing, and consumer-specific cutover behavior remain separate increments with their own RED-to-GREEN evidence. + +Source capability is not parity or deployment state. Promotion still requires exact protected lineage, normal integration, immutable release identity, representative traffic/security/performance evidence, consumer shadow/canary, observed rollback, cutover, and verified legacy reverse-proxy removal. diff --git a/docs/adr/0010-version-pg-erd-response-body-lifetime.md b/docs/adr/0010-version-pg-erd-response-body-lifetime.md new file mode 100644 index 00000000..d116b405 --- /dev/null +++ b/docs/adr/0010-version-pg-erd-response-body-lifetime.md @@ -0,0 +1,62 @@ +# ADR 0010: Version the pg-erd upstream response-body lifetime + +- Status: Proposed +- Date: 2026-09-03 +- Owners: Runtime Isolation / bounded pg-erd migration + +## Problem + +The characterized pg-erd candidate configures Pingora `read_timeout`, but at the pinned Pingora revision `09696b51bc59315353d96686355861604d0bb48c` that setting applies to each individual upstream read and resets after a successful read. A backend can therefore retain an admitted request indefinitely by continuing to send response-body progress before each inactivity timeout. This can consume the process-local in-flight budget even though no single read stalls. + +The existing version-1 pg-erd Admin Config has no whole-response field. Reinterpreting `read_ms`, deriving a hidden multiplier, or reusing `total_connection_ms` would change an existing contract without an explicit version boundary; `total_connection_ms` also belongs to connection establishment, including TLS, rather than response delivery. + +## Constraints + +- Generic gateway v1 must not silently gain product-specific long-response semantics. +- The pg-erd migration may add only edge/runtime authority. Product retry, idempotency, authentication, business routing, Wardnet/EgressWeave verdicts, and Keyverse identity remain outside the gateway. +- Existing pg-erd version-1 fixtures and predecessor PRs must remain readable while the stack is unreleased. +- A failure after the downstream response header is committed cannot be converted into a second HTTP status or silently failed over. +- The pinned Pingora callback API exposes response-header and response-body progress callbacks but does not expose an exact absolute-deadline interrupt for a currently pending upstream read. + +## Options considered + +### Keep only `read_ms` + +Rejected. It protects inactivity, not total response-body lifetime, so continuous slow-drip remains unbounded. + +### Reinterpret `read_ms` or derive a multiplier + +Rejected. This creates undocumented semantics, couples two distinct failure modes, and makes operator intent impossible to audit. + +### Reuse `total_connection_ms` + +Rejected. Pingora applies that budget to connection establishment; changing its meaning at the gateway layer would conflict with supplier semantics. + +### Add an explicit pg-erd version-2 body-lifetime budget + +Selected. Version 2 requires positive `max_upstream_response_body_ms`. Version 1 rejects that field and otherwise retains its existing behavior. Runtime Isolation owns monotonic elapsed-time accounting, while the migration adapter starts the budget on the first non-informational upstream response header and checks it only on non-empty upstream response-body progress callbacks. + +## Decision + +Introduce pg-erd Admin Config version 2 with mandatory positive `max_upstream_response_body_ms`. Preserve version 1 without a hidden lifetime. On a version-2 request, start the response-body budget at the first non-informational upstream response header. If a later body-progress callback arrives at or beyond the budget, raise an upstream-scoped fatal error. Preserve the existing post-commit invariant: terminate the incomplete downstream response, record low-cardinality request-error telemetry, release the request admission lease with its context, and do not invent retry or failover. + +This is a progress-driven bound, not an exact timer interrupt. A continuously progressing body is stopped at the first body callback at or after the configured lifetime. A quiescent upstream is still bounded by its independent per-read `read_ms`. Slow-drip of an incomplete response header remains a separate gap because the current callback guard has not yet established an absolute header-read deadline. + +## Effects and risks + +The selected design makes slow-drip body retention explicit and auditable and keeps it in the Runtime Isolation bounded context. It avoids changing generic v1 and preserves the pg-erd v1 characterization stack. Operators must choose the version-2 value from observed long-response requirements before canary or cutover; example values are not production SLOs. + +The remaining limitation is timer precision: scheduler delay and Pingora read cadence can move actual termination past the configured instant until the next body callback. A future supplier/runtime capability may justify an exact timer boundary, but this ADR does not claim one. + +## Verification + +- deterministic Runtime Isolation tests cover absent, dormant, active, and expired response-body budgets; +- versioned Admin Config tests cover v1 compatibility, v1 rejection of v2 fields, v2 explicit positive values, zero rejection, and incomplete v2 schema rejection; +- real-listener pg-erd traffic commits HTTP 200 and drips one body byte every 60 ms while `read_ms` is 500 ms and `max_upstream_response_body_ms` is 300 ms; the response must terminate before the declared 20-byte body completes, without a second status or route failover, while error telemetry, readiness, and an independent route recover; +- exact-head format, compile, clippy, rustdoc, 100% owned-production coverage, load, OCI, security, supply-chain, and independent-review evidence remain required before this ADR may move from Proposed to Accepted. + +## References + +Cloudflare. (2026). *Pingora peer configuration and timeout semantics* (revision `09696b51bc59315353d96686355861604d0bb48c`). GitHub. + +Cloudflare. (2026). *ProxyHttp response filtering callbacks* (revision `09696b51bc59315353d96686355861604d0bb48c`). GitHub. diff --git a/src/migration_admin.rs b/src/migration_admin.rs index c610482d..8237cf1d 100644 --- a/src/migration_admin.rs +++ b/src/migration_admin.rs @@ -22,9 +22,12 @@ use crate::migration_plan::EdgeMigrationPlan; use crate::migration_proxy::MigrationGatewayProxy; use crate::runtime_isolation::{RuntimeIsolationConfigError, RuntimeIsolationLimits}; -/// Version of the bounded `pg-erd-cloud` migration admin configuration. +/// Original bounded `pg-erd-cloud` migration configuration version without a body-lifetime budget. pub const PG_ERD_MIGRATION_CONFIG_VERSION: u32 = 1; +/// Opt-in pg-erd configuration version that requires an explicit response-body lifetime budget. +pub const PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION: u32 = 2; + /// Fail-closed admin configuration for the characterized `pg-erd-cloud` migration runtime. #[derive(Debug, Clone, Deserialize, PartialEq, Eq)] #[serde(deny_unknown_fields)] @@ -34,6 +37,8 @@ pub struct PgErdMigrationConfig { metrics_listener: SocketAddr, max_request_body_bytes: u64, max_in_flight_requests: usize, + #[serde(default)] + max_upstream_response_body_ms: Option, upstream_keepalive_pool_size: usize, upstreams: Vec, } @@ -47,6 +52,12 @@ pub enum PgErdMigrationConfigError { /// The configuration requests a migration-admin version this binary does not implement. #[error("unsupported pg-erd migration configuration version {0}")] UnsupportedVersion(u32), + /// Version 2 must state the response-body lifetime rather than inheriting a hidden default. + #[error("pg-erd migration config version 2 requires max_upstream_response_body_ms")] + MissingUpstreamResponseBodyLifetime, + /// Version 1 cannot silently acquire semantics introduced by the version-2 contract. + #[error("max_upstream_response_body_ms requires pg-erd migration config version 2")] + ResponseBodyLifetimeRequiresVersion2, /// A port-zero traffic listener would delegate the public authority to an ephemeral OS port. #[error("listener must use a non-zero port")] ZeroListenerPort, @@ -121,6 +132,11 @@ impl PgErdMigrationConfig { self.metrics_listener } + /// Returns the explicit response-body lifetime when the version-2 contract is active. + pub fn max_upstream_response_body_ms(&self) -> Option { + self.max_upstream_response_body_ms + } + /// Returns the validated Pingora upstream keepalive-pool budget. pub fn upstream_keepalive_pool_size(&self) -> usize { self.upstream_keepalive_pool_size @@ -136,17 +152,24 @@ impl PgErdMigrationConfig { pub fn build_proxy(&self) -> Result { self.validate()?; let delivery = self.build_delivery()?; - let limits = RuntimeIsolationLimits::from_validated( - self.max_request_body_bytes, - self.max_in_flight_requests, - ); + let limits = self.runtime_isolation_limits_from_validated(); Ok(MigrationGatewayProxy::new(delivery, limits)) } /// Rejects configuration that could fail after listener authority has already been granted. fn validate(&self) -> Result<(), PgErdMigrationConfigError> { - if self.version != PG_ERD_MIGRATION_CONFIG_VERSION { - return Err(PgErdMigrationConfigError::UnsupportedVersion(self.version)); + match self.version { + PG_ERD_MIGRATION_CONFIG_VERSION => { + if self.max_upstream_response_body_ms.is_some() { + return Err(PgErdMigrationConfigError::ResponseBodyLifetimeRequiresVersion2); + } + } + PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION => { + if self.max_upstream_response_body_ms.is_none() { + return Err(PgErdMigrationConfigError::MissingUpstreamResponseBodyLifetime); + } + } + unsupported => return Err(PgErdMigrationConfigError::UnsupportedVersion(unsupported)), } if self.listener.port() == 0 { return Err(PgErdMigrationConfigError::ZeroListenerPort); @@ -161,10 +184,42 @@ impl PgErdMigrationConfig { return Err(PgErdMigrationConfigError::InvalidUpstreamKeepalivePoolSize); } - RuntimeIsolationLimits::try_new(self.max_request_body_bytes, self.max_in_flight_requests)?; + match self.max_upstream_response_body_ms { + Some(max_upstream_response_body_ms) => { + RuntimeIsolationLimits::try_new_with_response_body_limit( + self.max_request_body_bytes, + self.max_in_flight_requests, + max_upstream_response_body_ms, + )?; + } + None => { + RuntimeIsolationLimits::try_new( + self.max_request_body_bytes, + self.max_in_flight_requests, + )?; + } + } self.validate_transport_authority(&pg_erd_migration_plan()) } + fn runtime_isolation_limits_from_validated(&self) -> RuntimeIsolationLimits { + self.max_upstream_response_body_ms.map_or_else( + || { + RuntimeIsolationLimits::from_validated( + self.max_request_body_bytes, + self.max_in_flight_requests, + ) + }, + |max_upstream_response_body_ms| { + RuntimeIsolationLimits::from_validated_with_response_body_limit( + self.max_request_body_bytes, + self.max_in_flight_requests, + max_upstream_response_body_ms, + ) + }, + ) + } + /// Enforces a complete one-to-one binding of operator transport data to compiled upstream names. fn validate_transport_authority( &self, diff --git a/src/migration_proxy.rs b/src/migration_proxy.rs index d28dbe58..a87f7022 100644 --- a/src/migration_proxy.rs +++ b/src/migration_proxy.rs @@ -4,6 +4,8 @@ //! isolation, trusted forwarding metadata, and shared transport observability. It does not //! introduce product authorization, service discovery, or business logic. +use std::time::{Duration, Instant}; + use async_trait::async_trait; use bytes::Bytes; use log::error; @@ -22,7 +24,7 @@ use crate::pingora_delivery::reject_uncharacterized_http1_protocol_transition; use crate::process_health::{respond_healthy, LIVENESS_PATH, READINESS_PATH}; use crate::runtime_isolation::{ BodyLimitExceeded, RequestAdmission, RequestAdmissionBudget, RequestBodyBudget, - RuntimeIsolationLimits, + ResponseBodyLifetimeBudget, ResponseBodyLifetimeExceeded, RuntimeIsolationLimits, }; /// Fail-closed callback errors for a characterized migration runtime. @@ -40,6 +42,7 @@ pub enum MigrationGatewayProxyError { #[derive(Debug)] pub struct MigrationRequestContext { request_body: RequestBodyBudget, + response_body_lifetime: ResponseBodyLifetimeBudget, admission: Option, } @@ -47,6 +50,7 @@ impl MigrationRequestContext { fn new(limits: RuntimeIsolationLimits) -> Self { Self { request_body: RequestBodyBudget::new(limits), + response_body_lifetime: ResponseBodyLifetimeBudget::new(limits), admission: None, } } @@ -161,6 +165,34 @@ fn body_rejection_to_pingora(rejection: BodyLimitExceeded) -> Box { ) } +fn response_body_lifetime_to_pingora(rejection: ResponseBodyLifetimeExceeded) -> Box { + let _ = (rejection.elapsed, rejection.limit); + Error::new_up(ErrorType::Custom("UpstreamResponseBodyLifetimeExceeded")) +} + +fn start_response_body_lifetime( + is_informational: bool, + ctx: &mut MigrationRequestContext, + now: Instant, +) { + if !is_informational { + ctx.response_body_lifetime.start(now); + } +} + +fn enforce_response_body_lifetime( + body: &Option, + ctx: &MigrationRequestContext, + now: Instant, +) -> pingora::Result> { + if body.as_ref().is_some_and(|chunk| !chunk.is_empty()) { + ctx.response_body_lifetime + .reject_if_expired(now) + .map_err(response_body_lifetime_to_pingora)?; + } + Ok(None) +} + fn unmatched_route_to_pingora(_error: MigrationGatewayProxyError) -> Box { Error::explain( ErrorType::HTTPStatus(404), @@ -183,6 +215,13 @@ fn proxy_error_status(error: &Error) -> u16 { } } +fn proxy_error_response_status(error: &Error, response_already_written: bool) -> u16 { + if response_already_written { + return 0; + } + proxy_error_status(error) +} + #[async_trait] impl ProxyHttp for MigrationGatewayProxy { type CTX = MigrationRequestContext; @@ -252,6 +291,23 @@ impl ProxyHttp for MigrationGatewayProxy { self.apply_upstream_request_policy(upstream_request, &forwarding) } + async fn upstream_response_filter( + &self, + _session: &mut Session, + upstream_response: &mut ResponseHeader, + ctx: &mut Self::CTX, + ) -> pingora::Result<()> + where + Self::CTX: Send + Sync, + { + start_response_body_lifetime( + upstream_response.status.is_informational(), + ctx, + Instant::now(), + ); + Ok(()) + } + async fn response_filter( &self, _session: &mut Session, @@ -264,13 +320,23 @@ impl ProxyHttp for MigrationGatewayProxy { self.apply_response_headers(upstream_response) } + fn upstream_response_body_filter( + &self, + _session: &mut Session, + body: &mut Option, + _end_of_stream: bool, + ctx: &mut Self::CTX, + ) -> pingora::Result> { + enforce_response_body_lifetime(body, ctx, Instant::now()) + } + async fn fail_to_proxy( &self, session: &mut Session, error_value: &Error, _ctx: &mut Self::CTX, ) -> FailToProxy { - let status = proxy_error_status(error_value); + let status = proxy_error_response_status(error_value, session.response_written().is_some()); if status > 0 { let mut response = ServerSession::generate_error(status); if let Err(policy_error) = self.apply_response_headers(&mut response) { @@ -305,20 +371,26 @@ impl ProxyHttp for MigrationGatewayProxy { #[cfg(test)] mod tests { use std::net::SocketAddr; + use std::time::{Duration, Instant}; + use bytes::Bytes; use pingora::prelude::{Error, ErrorType}; use pingora::ErrorSource; use super::{ - body_rejection_to_pingora, proxy_error_status, unmatched_route_to_pingora, - MigrationGatewayProxy, MigrationGatewayProxyError, MigrationRequestContext, + body_rejection_to_pingora, enforce_response_body_lifetime, proxy_error_response_status, + proxy_error_status, response_body_lifetime_to_pingora, start_response_body_lifetime, + unmatched_route_to_pingora, MigrationGatewayProxy, MigrationGatewayProxyError, + MigrationRequestContext, }; use crate::edge_contract::{UpstreamConfig, UpstreamTimeouts}; use crate::edge_routing::{RouteMatch, RouteRule}; use crate::http_policy::ResponseHeaderRule; use crate::migration_delivery::MigrationDeliveryPlan; use crate::migration_plan::EdgeMigrationPlan; - use crate::runtime_isolation::{BodyLimitExceeded, RuntimeIsolationLimits}; + use crate::runtime_isolation::{ + BodyLimitExceeded, ResponseBodyLifetimeExceeded, RuntimeIsolationLimits, + }; fn admission_proxy() -> (MigrationGatewayProxy, RuntimeIsolationLimits) { let plan = EdgeMigrationPlan::try_new( @@ -390,6 +462,54 @@ mod tests { assert!(rejected.admission.is_some()); } + #[test] + fn informational_headers_do_not_start_or_reset_the_response_body_lifetime() { + let limits = RuntimeIsolationLimits::try_new_with_response_body_limit(8, 1, 300) + .expect("fixture limits are valid"); + let mut ctx = MigrationRequestContext::new(limits); + let now = Instant::now(); + + start_response_body_lifetime(true, &mut ctx, now); + assert!(ctx + .response_body_lifetime + .reject_if_expired(now + Duration::from_secs(1)) + .is_ok()); + + start_response_body_lifetime(false, &mut ctx, now); + start_response_body_lifetime(false, &mut ctx, now + Duration::from_millis(250)); + assert!(ctx + .response_body_lifetime + .reject_if_expired(now + Duration::from_millis(299)) + .is_ok()); + assert!(ctx + .response_body_lifetime + .reject_if_expired(now + Duration::from_millis(300)) + .is_err()); + } + + #[test] + fn response_lifetime_ignores_empty_callbacks_but_rejects_expired_body_progress() { + let limits = RuntimeIsolationLimits::try_new_with_response_body_limit(8, 1, 300) + .expect("fixture limits are valid"); + let mut ctx = MigrationRequestContext::new(limits); + let started = Instant::now(); + start_response_body_lifetime(false, &mut ctx, started); + let before_expiry = started + Duration::from_millis(299); + let expired = started + Duration::from_millis(300); + + assert!(enforce_response_body_lifetime(&None, &ctx, expired).is_ok()); + assert!(enforce_response_body_lifetime(&Some(Bytes::new()), &ctx, expired).is_ok()); + assert!(enforce_response_body_lifetime( + &Some(Bytes::from_static(b"x")), + &ctx, + before_expiry + ) + .is_ok()); + assert!( + enforce_response_body_lifetime(&Some(Bytes::from_static(b"x")), &ctx, expired).is_err() + ); + } + #[test] fn delivery_errors_map_to_fail_closed_http_errors() { let body_error = body_rejection_to_pingora(BodyLimitExceeded { @@ -402,6 +522,18 @@ mod tests { request_path: "/missing".to_string(), }); assert_eq!(route_error.etype, ErrorType::HTTPStatus(404)); + + let lifetime_error = response_body_lifetime_to_pingora(ResponseBodyLifetimeExceeded { + elapsed: Duration::from_millis(301), + limit: Duration::from_millis(300), + }); + assert_eq!( + lifetime_error.etype, + ErrorType::Custom("UpstreamResponseBodyLifetimeExceeded") + ); + assert_eq!(lifetime_error.esource, ErrorSource::Upstream); + assert_eq!(proxy_error_response_status(&lifetime_error, true), 0); + assert_eq!(proxy_error_response_status(&lifetime_error, false), 502); } #[test] diff --git a/src/runtime_isolation.rs b/src/runtime_isolation.rs index c81ec372..95b4873c 100644 --- a/src/runtime_isolation.rs +++ b/src/runtime_isolation.rs @@ -1,10 +1,12 @@ //! Transport-neutral runtime-isolation budgets shared by Pingora delivery adapters. //! -//! This bounded context owns request-body and concurrent-request admission limits. It does not -//! select routes, mutate HTTP policy, authenticate callers, or make product-domain decisions. +//! This bounded context owns request-body, concurrent-request admission, and optional upstream +//! response-body lifetime limits. It does not select routes, mutate HTTP policy, authenticate +//! callers, or make product-domain decisions. use std::sync::atomic::{AtomicUsize, Ordering}; use std::sync::Arc; +use std::time::{Duration, Instant}; use thiserror::Error; @@ -17,6 +19,9 @@ pub enum RuntimeIsolationConfigError { /// A zero in-flight limit would reject every proxied request. #[error("max_in_flight_requests must be greater than zero")] ZeroMaxInFlightRequests, + /// A zero response-body lifetime would reject every non-empty upstream response immediately. + #[error("max_upstream_response_body_ms must be greater than zero")] + ZeroMaxUpstreamResponseBodyMs, } /// Immutable request-isolation limits shared by gateway delivery adapters. @@ -24,13 +29,38 @@ pub enum RuntimeIsolationConfigError { pub struct RuntimeIsolationLimits { max_request_body_bytes: u64, max_in_flight_requests: usize, + max_upstream_response_body_ms: Option, } impl RuntimeIsolationLimits { /// Validates explicit non-zero body and in-flight request budgets. + /// + /// The generic v1 contract has no response-body lifetime field, so this constructor preserves + /// that versioned behavior rather than inventing a hidden timeout. pub fn try_new( max_request_body_bytes: u64, max_in_flight_requests: usize, + ) -> Result { + Self::try_new_internal(max_request_body_bytes, max_in_flight_requests, None) + } + + /// Validates request isolation plus an explicit upstream response-body lifetime budget. + pub(crate) fn try_new_with_response_body_limit( + max_request_body_bytes: u64, + max_in_flight_requests: usize, + max_upstream_response_body_ms: u64, + ) -> Result { + Self::try_new_internal( + max_request_body_bytes, + max_in_flight_requests, + Some(max_upstream_response_body_ms), + ) + } + + fn try_new_internal( + max_request_body_bytes: u64, + max_in_flight_requests: usize, + max_upstream_response_body_ms: Option, ) -> Result { if max_request_body_bytes == 0 { return Err(RuntimeIsolationConfigError::ZeroMaxRequestBodyBytes); @@ -38,9 +68,13 @@ impl RuntimeIsolationLimits { if max_in_flight_requests == 0 { return Err(RuntimeIsolationConfigError::ZeroMaxInFlightRequests); } + if max_upstream_response_body_ms == Some(0) { + return Err(RuntimeIsolationConfigError::ZeroMaxUpstreamResponseBodyMs); + } Ok(Self { max_request_body_bytes, max_in_flight_requests, + max_upstream_response_body_ms, }) } @@ -51,6 +85,19 @@ impl RuntimeIsolationLimits { Self { max_request_body_bytes, max_in_flight_requests, + max_upstream_response_body_ms: None, + } + } + + pub(crate) fn from_validated_with_response_body_limit( + max_request_body_bytes: u64, + max_in_flight_requests: usize, + max_upstream_response_body_ms: u64, + ) -> Self { + Self { + max_request_body_bytes, + max_in_flight_requests, + max_upstream_response_body_ms: Some(max_upstream_response_body_ms), } } @@ -63,6 +110,11 @@ impl RuntimeIsolationLimits { pub fn max_in_flight_requests(self) -> usize { self.max_in_flight_requests } + + /// Returns the configured upstream response-body lifetime, when the active contract owns one. + pub fn max_upstream_response_body_ms(self) -> Option { + self.max_upstream_response_body_ms + } } #[derive(Debug, Clone, PartialEq, Eq)] @@ -148,11 +200,61 @@ impl RequestBodyBudget { } } +/// Elapsed-time guard for the body phase of one admitted upstream response. +#[derive(Debug)] +pub(crate) struct ResponseBodyLifetimeBudget { + limit: Option, + started_at: Option, +} + +/// Evidence that response-body progress arrived after its configured lifetime. +#[derive(Debug, Clone, PartialEq, Eq)] +pub(crate) struct ResponseBodyLifetimeExceeded { + pub(crate) elapsed: Duration, + pub(crate) limit: Duration, +} + +impl ResponseBodyLifetimeBudget { + /// Creates a dormant response-body budget from the active runtime-isolation contract. + pub(crate) fn new(limits: RuntimeIsolationLimits) -> Self { + Self { + limit: limits + .max_upstream_response_body_ms() + .map(Duration::from_millis), + started_at: None, + } + } + + /// Starts the body lifetime once; repeated response-header callbacks cannot reset the deadline. + pub(crate) fn start(&mut self, now: Instant) { + if self.limit.is_some() && self.started_at.is_none() { + self.started_at = Some(now); + } + } + + /// Rejects the first observed body-progress boundary at or beyond the configured lifetime. + pub(crate) fn reject_if_expired( + &self, + now: Instant, + ) -> Result<(), ResponseBodyLifetimeExceeded> { + let (Some(limit), Some(started_at)) = (self.limit, self.started_at) else { + return Ok(()); + }; + let elapsed = now.saturating_duration_since(started_at); + if elapsed >= limit { + return Err(ResponseBodyLifetimeExceeded { elapsed, limit }); + } + Ok(()) + } +} + #[cfg(test)] mod tests { + use std::time::{Duration, Instant}; + use super::{ - BodyLimitExceeded, RequestAdmissionBudget, RequestBodyBudget, RuntimeIsolationConfigError, - RuntimeIsolationLimits, + BodyLimitExceeded, RequestAdmissionBudget, RequestBodyBudget, ResponseBodyLifetimeBudget, + ResponseBodyLifetimeExceeded, RuntimeIsolationConfigError, RuntimeIsolationLimits, }; #[test] @@ -165,10 +267,27 @@ mod tests { RuntimeIsolationLimits::try_new(1, 0), Err(RuntimeIsolationConfigError::ZeroMaxInFlightRequests) ); + assert_eq!( + RuntimeIsolationLimits::try_new_with_response_body_limit(1, 1, 0), + Err(RuntimeIsolationConfigError::ZeroMaxUpstreamResponseBodyMs) + ); let limits = RuntimeIsolationLimits::try_new(1024, 2).expect("non-zero limits are valid"); assert_eq!(limits.max_request_body_bytes(), 1024); assert_eq!(limits.max_in_flight_requests(), 2); + assert_eq!(limits.max_upstream_response_body_ms(), None); + + let bounded = RuntimeIsolationLimits::try_new_with_response_body_limit(2048, 3, 750) + .expect("explicit positive response lifetime must be valid"); + assert_eq!(bounded.max_request_body_bytes(), 2048); + assert_eq!(bounded.max_in_flight_requests(), 3); + assert_eq!(bounded.max_upstream_response_body_ms(), Some(750)); + + let validated = + RuntimeIsolationLimits::from_validated_with_response_body_limit(4096, 4, 900); + assert_eq!(validated.max_request_body_bytes(), 4096); + assert_eq!(validated.max_in_flight_requests(), 4); + assert_eq!(validated.max_upstream_response_body_ms(), Some(900)); } #[test] @@ -207,4 +326,43 @@ mod tests { } ); } + + #[test] + fn response_body_lifetime_is_dormant_without_a_versioned_budget() { + let limits = RuntimeIsolationLimits::try_new(4, 1).expect("fixture limits are valid"); + let mut budget = ResponseBodyLifetimeBudget::new(limits); + let now = Instant::now(); + + budget.start(now); + assert!(budget + .reject_if_expired(now + Duration::from_secs(60)) + .is_ok()); + } + + #[test] + fn response_body_lifetime_starts_once_and_rejects_at_the_limit() { + let limits = RuntimeIsolationLimits::try_new_with_response_body_limit(4, 1, 300) + .expect("fixture limits are valid"); + let mut budget = ResponseBodyLifetimeBudget::new(limits); + let started = Instant::now(); + assert!(budget + .reject_if_expired(started + Duration::from_secs(1)) + .is_ok()); + + budget.start(started); + budget.start(started + Duration::from_millis(250)); + + assert!(budget + .reject_if_expired(started + Duration::from_millis(299)) + .is_ok()); + assert_eq!( + budget + .reject_if_expired(started + Duration::from_millis(300)) + .unwrap_err(), + ResponseBodyLifetimeExceeded { + elapsed: Duration::from_millis(300), + limit: Duration::from_millis(300), + } + ); + } } diff --git a/tests/pg_erd_admin_config_contract.rs b/tests/pg_erd_admin_config_contract.rs index 0ba9b244..d157325a 100644 --- a/tests/pg_erd_admin_config_contract.rs +++ b/tests/pg_erd_admin_config_contract.rs @@ -1,5 +1,6 @@ use cwl_pingora_gateway::migration_admin::{ PgErdMigrationConfig, PgErdMigrationConfigError, PG_ERD_MIGRATION_CONFIG_VERSION, + PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION, }; use pingora::upstreams::peer::Peer; @@ -180,6 +181,17 @@ fn pg_erd_admin_config_build_proxy_revalidates_direct_deserialization() { "public build_proxy must revalidate directly deserialized runtime budgets" ); } + + let incomplete_v2 = valid_yaml().replace( + &format!("version: {PG_ERD_MIGRATION_CONFIG_VERSION}"), + &format!("version: {PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION}"), + ); + let config: PgErdMigrationConfig = serde_yaml::from_str(&incomplete_v2) + .expect("direct deserialization may construct an incomplete version-2 value"); + assert!(matches!( + config.build_proxy(), + Err(PgErdMigrationConfigError::MissingUpstreamResponseBodyLifetime) + )); } #[test] @@ -273,7 +285,7 @@ fn pg_erd_admin_config_rejects_invalid_concrete_transport_contract() { } #[test] -fn pg_erd_admin_config_rejects_unknown_fields_and_future_versions() { +fn pg_erd_admin_config_rejects_unknown_incomplete_and_future_versions() { let unknown = valid_yaml().replace( "max_request_body_bytes: 1048576", "max_request_body_bytes: 1048576\nproduct_auth_mode: embedded", @@ -283,12 +295,21 @@ fn pg_erd_admin_config_rejects_unknown_fields_and_future_versions() { Err(PgErdMigrationConfigError::Parse(_)) )); + let incomplete_v2 = valid_yaml().replace( + &format!("version: {PG_ERD_MIGRATION_CONFIG_VERSION}"), + &format!("version: {PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION}"), + ); + assert_eq!( + PgErdMigrationConfig::from_yaml(&incomplete_v2), + Err(PgErdMigrationConfigError::MissingUpstreamResponseBodyLifetime) + ); + let future = valid_yaml().replace( &format!("version: {PG_ERD_MIGRATION_CONFIG_VERSION}"), - "version: 2", + "version: 3", ); assert_eq!( PgErdMigrationConfig::from_yaml(&future), - Err(PgErdMigrationConfigError::UnsupportedVersion(2)) + Err(PgErdMigrationConfigError::UnsupportedVersion(3)) ); } diff --git a/tests/pg_erd_response_lifetime_config.rs b/tests/pg_erd_response_lifetime_config.rs new file mode 100644 index 00000000..82b560f8 --- /dev/null +++ b/tests/pg_erd_response_lifetime_config.rs @@ -0,0 +1,54 @@ +use cwl_pingora_gateway::migration_admin::{ + PgErdMigrationConfig, PgErdMigrationConfigError, PG_ERD_MIGRATION_CONFIG_VERSION, + PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION, +}; +use cwl_pingora_gateway::runtime_isolation::RuntimeIsolationConfigError; + +fn config_yaml(version: u32, lifetime_line: &str) -> String { + format!( + "version: {version}\nlistener: 127.0.0.1:18080\nmetrics_listener: 127.0.0.1:19090\nmax_request_body_bytes: 1024\nmax_in_flight_requests: 8\n{lifetime_line}upstream_keepalive_pool_size: 4\nupstreams:\n - name: backend\n address: 127.0.0.1:18000\n tls: false\n timeouts:\n connection_ms: 100\n total_connection_ms: 200\n read_ms: 300\n write_ms: 400\n idle_ms: 500\n - name: frontend\n address: 127.0.0.1:13000\n tls: false\n timeouts:\n connection_ms: 100\n total_connection_ms: 200\n read_ms: 300\n write_ms: 400\n idle_ms: 500\n" + ) +} + +#[test] +fn version_two_requires_and_preserves_an_explicit_positive_response_body_lifetime() { + let configured = PgErdMigrationConfig::from_yaml(&config_yaml( + PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION, + "max_upstream_response_body_ms: 750\n", + )) + .expect("version 2 with an explicit positive lifetime must parse"); + assert_eq!(configured.max_upstream_response_body_ms(), Some(750)); + configured + .build_proxy() + .expect("version-2 runtime budgets must materialize before listeners open"); + + assert_eq!( + PgErdMigrationConfig::from_yaml(&config_yaml( + PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION, + "max_upstream_response_body_ms: 0\n", + )), + Err(PgErdMigrationConfigError::RuntimeIsolation( + RuntimeIsolationConfigError::ZeroMaxUpstreamResponseBodyMs, + )) + ); + + assert_eq!( + PgErdMigrationConfig::from_yaml(&config_yaml(PG_ERD_RESPONSE_LIFETIME_CONFIG_VERSION, "",)), + Err(PgErdMigrationConfigError::MissingUpstreamResponseBodyLifetime) + ); +} + +#[test] +fn version_one_cannot_silently_acquire_version_two_response_semantics() { + let legacy = PgErdMigrationConfig::from_yaml(&config_yaml(PG_ERD_MIGRATION_CONFIG_VERSION, "")) + .expect("legacy version 1 remains readable without a hidden response lifetime"); + assert_eq!(legacy.max_upstream_response_body_ms(), None); + + assert_eq!( + PgErdMigrationConfig::from_yaml(&config_yaml( + PG_ERD_MIGRATION_CONFIG_VERSION, + "max_upstream_response_body_ms: 750\n", + )), + Err(PgErdMigrationConfigError::ResponseBodyLifetimeRequiresVersion2) + ); +} diff --git a/tests/pg_erd_slow_drip_response_traffic.rs b/tests/pg_erd_slow_drip_response_traffic.rs new file mode 100644 index 00000000..25e1029e --- /dev/null +++ b/tests/pg_erd_slow_drip_response_traffic.rs @@ -0,0 +1,419 @@ +//! Real-listener RED→GREEN contract for a pg-erd upstream that continuously drips response-body bytes. +//! +//! Pingora's peer `read_timeout` is an inactivity timeout that resets after each successful read. +//! This fixture therefore keeps each origin write well inside `read_ms` while extending the response +//! beyond an explicit migration-owned response-body lifetime. Fixture readiness and I/O use absolute +//! deadlines so partial progress cannot renew the evidence window. + +use std::io::{ErrorKind, Read, Write}; +use std::net::{SocketAddr, TcpListener, TcpStream}; +use std::process::{Child, Command, Stdio}; +use std::thread; +use std::time::{Duration, Instant}; + +use tempfile::NamedTempFile; + +const MAX_HEADER_BYTES: usize = 64 * 1024; +const FIXTURE_IO_TIMEOUT: Duration = Duration::from_secs(5); +const STARTUP_TIMEOUT: Duration = Duration::from_secs(10); +const TERMINATION_TIMEOUT: Duration = Duration::from_secs(2); +const PRE_RESPONSE_HEADER_DELAY: Duration = Duration::from_millis(150); +const RESPONSE_BODY_LIFETIME: Duration = Duration::from_millis(300); + +struct GatewayProcess(Child); + +impl Drop for GatewayProcess { + fn drop(&mut self) { + let _ = self.0.kill(); + let _ = self.0.wait(); + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum DownstreamTermination { + Eof, + ConnectionReset, +} + +fn reserve_distinct_loopback_listeners() -> (TcpListener, TcpListener) { + let traffic = TcpListener::bind("127.0.0.1:0").expect("traffic port should be reservable"); + let metrics = TcpListener::bind("127.0.0.1:0").expect("metrics port should be reservable"); + assert_ne!( + traffic.local_addr().expect("traffic reservation address"), + metrics.local_addr().expect("metrics reservation address") + ); + (traffic, metrics) +} + +fn write_config( + listener: SocketAddr, + metrics_listener: SocketAddr, + backend: SocketAddr, + frontend: SocketAddr, +) -> NamedTempFile { + let mut file = NamedTempFile::new().expect("temporary config should be writable"); + writeln!( + file, + "version: 2\nlistener: {listener}\nmetrics_listener: {metrics_listener}\nmax_request_body_bytes: 8\nmax_in_flight_requests: 8\nmax_upstream_response_body_ms: 300\nupstream_keepalive_pool_size: 4\nupstreams:\n - name: backend\n address: {backend}\n tls: false\n timeouts:\n connection_ms: 200\n total_connection_ms: 400\n read_ms: 500\n write_ms: 1000\n idle_ms: 5000\n - name: frontend\n address: {frontend}\n tls: false\n timeouts:\n connection_ms: 200\n total_connection_ms: 400\n read_ms: 1000\n write_ms: 1000\n idle_ms: 5000" + ) + .expect("migration config should be written"); + file +} + +fn remaining(deadline: Instant, context: &str) -> Duration { + deadline + .checked_duration_since(Instant::now()) + .filter(|duration| !duration.is_zero()) + .unwrap_or_else(|| panic!("absolute fixture deadline expired while {context}")) +} + +fn set_read_timeout_to_remaining(stream: &TcpStream, deadline: Instant, context: &str) { + stream + .set_read_timeout(Some(remaining(deadline, context))) + .expect("read timeout should be configurable"); +} + +fn set_write_timeout_to_remaining(stream: &TcpStream, deadline: Instant, context: &str) { + stream + .set_write_timeout(Some(remaining(deadline, context))) + .expect("write timeout should be configurable"); +} + +fn read_header_block( + stream: &mut TcpStream, + deadline: Instant, + context: &str, +) -> std::io::Result> { + let mut bytes = Vec::new(); + let mut buffer = [0_u8; 1024]; + loop { + set_read_timeout_to_remaining(stream, deadline, context); + let read = stream.read(&mut buffer)?; + if read == 0 { + return Err(std::io::Error::new( + ErrorKind::UnexpectedEof, + "connection closed before response headers completed", + )); + } + bytes.extend_from_slice(&buffer[..read]); + if bytes.len() > MAX_HEADER_BYTES { + return Err(std::io::Error::new( + ErrorKind::InvalidData, + "header block exceeded 64 KiB fixture bound", + )); + } + if bytes.windows(4).any(|window| window == b"\r\n\r\n") { + return Ok(bytes); + } + } +} + +fn probe_http_status(address: SocketAddr, path: &str, deadline: Instant) -> Option { + let connect_timeout = + remaining(deadline, "connecting readiness probe").min(Duration::from_millis(100)); + let mut stream = TcpStream::connect_timeout(&address, connect_timeout).ok()?; + let request = + format!("GET {path} HTTP/1.1\r\nHost: readiness.invalid\r\nConnection: close\r\n\r\n"); + set_write_timeout_to_remaining(&stream, deadline, "writing readiness probe"); + stream.write_all(request.as_bytes()).ok()?; + let headers = read_header_block(&mut stream, deadline, "reading readiness response").ok()?; + http11_status(&String::from_utf8_lossy(&headers)) +} + +fn wait_until_http_status(address: SocketAddr, path: &str, process: &mut Child, expected: u16) { + let deadline = Instant::now() + STARTUP_TIMEOUT; + loop { + if let Some(status) = process + .try_wait() + .expect("gateway process state should be readable") + { + panic!("gateway exited before HTTP readiness: {status}"); + } + + if Instant::now() >= deadline { + panic!("gateway did not return HTTP {expected} for {path} within 10s"); + } + + if probe_http_status(address, path, deadline) == Some(expected) { + return; + } + + let sleep_for = + remaining(deadline, "waiting to retry readiness").min(Duration::from_millis(25)); + thread::sleep(sleep_for); + } +} + +fn start_gateway( + config: &NamedTempFile, + gateway_address: SocketAddr, + metrics_address: SocketAddr, +) -> GatewayProcess { + let mut child = Command::new(env!("CARGO_BIN_EXE_cwl-pingora-pg-erd-migration")) + .args(["--config", config.path().to_str().expect("UTF-8 temp path")]) + .stdin(Stdio::null()) + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .spawn() + .expect("compiled pg-erd migration binary should start"); + + wait_until_http_status(gateway_address, "/readyz", &mut child, 200); + wait_until_http_status(metrics_address, "/metrics", &mut child, 200); + GatewayProcess(child) +} + +fn raw_request(address: SocketAddr, request: &[u8]) -> String { + let deadline = Instant::now() + FIXTURE_IO_TIMEOUT; + let mut downstream = + TcpStream::connect_timeout(&address, remaining(deadline, "connecting bounded request")) + .expect("gateway should accept traffic"); + set_write_timeout_to_remaining(&downstream, deadline, "writing bounded request"); + downstream + .write_all(request) + .expect("downstream request should be writable"); + + let mut response = Vec::new(); + let mut buffer = [0_u8; 1024]; + loop { + set_read_timeout_to_remaining(&downstream, deadline, "reading bounded response"); + match downstream.read(&mut buffer) { + Ok(0) => break, + Ok(read) => response.extend_from_slice(&buffer[..read]), + Err(error) if error.kind() == ErrorKind::ConnectionReset => break, + Err(error) => { + panic!("gateway response should complete inside absolute deadline: {error}") + } + } + } + String::from_utf8_lossy(&response).into_owned() +} + +fn raw_request_until_terminal( + address: SocketAddr, + request: &[u8], +) -> (Vec, DownstreamTermination, Duration) { + let started = Instant::now(); + let deadline = started + TERMINATION_TIMEOUT; + let mut downstream = TcpStream::connect_timeout( + &address, + remaining(deadline, "connecting slow-drip downstream"), + ) + .expect("gateway should accept traffic"); + set_write_timeout_to_remaining(&downstream, deadline, "writing slow-drip request"); + downstream + .write_all(request) + .expect("downstream request should be writable"); + + let mut response = Vec::new(); + let mut buffer = [0_u8; 1024]; + loop { + set_read_timeout_to_remaining(&downstream, deadline, "reading slow-drip termination"); + match downstream.read(&mut buffer) { + Ok(0) => return (response, DownstreamTermination::Eof, started.elapsed()), + Ok(read) => response.extend_from_slice(&buffer[..read]), + Err(error) if error.kind() == ErrorKind::ConnectionReset => { + return ( + response, + DownstreamTermination::ConnectionReset, + started.elapsed(), + ); + } + Err(error) => panic!( + "slow-drip downstream response must terminate inside one absolute 2s deadline: {error}" + ), + } + } +} + +fn get(address: SocketAddr, path: &str) -> String { + raw_request( + address, + format!("GET {path} HTTP/1.1\r\nHost: app.example:8080\r\nConnection: close\r\n\r\n") + .as_bytes(), + ) +} + +fn read_request_headers(stream: &mut TcpStream) -> String { + let deadline = Instant::now() + FIXTURE_IO_TIMEOUT; + let bytes = read_header_block(stream, deadline, "reading origin request headers") + .expect("origin request headers should complete inside one absolute 5s deadline"); + String::from_utf8_lossy(&bytes).into_owned() +} + +fn http11_status(response: &str) -> Option { + let mut tokens = response.lines().next()?.split_ascii_whitespace(); + if tokens.next()? != "HTTP/1.1" { + return None; + } + let code = tokens.next()?; + if code.len() != 3 || !code.bytes().all(|byte| byte.is_ascii_digit()) { + return None; + } + code.parse().ok() +} + +fn header_values(headers: &str, field_name: &str) -> Vec { + headers + .lines() + .skip(1) + .filter_map(|line| line.split_once(':')) + .filter(|(name, _)| name.eq_ignore_ascii_case(field_name)) + .map(|(_, value)| value.trim().to_string()) + .collect() +} + +fn has_exact_metric_sample(metrics: &str, sample: &str) -> bool { + metrics.lines().any(|line| line.trim() == sample) +} + +#[test] +fn compiled_pg_erd_terminates_continuous_response_drip_without_poisoning_other_routes() { + let backend = TcpListener::bind("127.0.0.1:0").expect("backend fixture should bind"); + let backend_address = backend.local_addr().expect("backend address should exist"); + let backend_origin = thread::spawn(move || { + let (mut stream, _) = backend + .accept() + .expect("routed request should reach the characterized backend authority"); + stream + .set_write_timeout(Some(FIXTURE_IO_TIMEOUT)) + .expect("origin write timeout should be configurable"); + let request = read_request_headers(&mut stream); + assert!(request.starts_with("GET /api/slow-drip HTTP/1.1\r\n")); + + thread::sleep(PRE_RESPONSE_HEADER_DELAY); + stream + .write_all(b"HTTP/1.1 200 OK\r\nContent-Length: 20\r\nConnection: close\r\n\r\n") + .expect("backend response header should be writable"); + + for _ in 0..20 { + match stream.write_all(b"x") { + Ok(()) => {} + Err(error) + if matches!( + error.kind(), + ErrorKind::BrokenPipe + | ErrorKind::ConnectionReset + | ErrorKind::NotConnected + ) => + { + break; + } + Err(error) => panic!("unexpected slow-drip origin write failure: {error}"), + } + thread::sleep(Duration::from_millis(60)); + } + }); + + let frontend = TcpListener::bind("127.0.0.1:0").expect("frontend fixture should bind"); + let frontend_address = frontend + .local_addr() + .expect("frontend address should exist"); + let frontend_origin = thread::spawn(move || { + let (mut stream, _) = frontend + .accept() + .expect("fallback request should reach the independent frontend authority"); + stream + .set_write_timeout(Some(FIXTURE_IO_TIMEOUT)) + .expect("origin write timeout should be configurable"); + let request = read_request_headers(&mut stream); + assert!(request.starts_with("GET /after-slow-drip HTTP/1.1\r\n")); + stream + .write_all( + b"HTTP/1.1 200 OK\r\nContent-Length: 9\r\nConnection: close\r\n\r\nrecovered", + ) + .expect("frontend recovery response should be writable"); + }); + + let (traffic_reservation, metrics_reservation) = reserve_distinct_loopback_listeners(); + let gateway_address = traffic_reservation + .local_addr() + .expect("traffic reservation address"); + let metrics_address = metrics_reservation + .local_addr() + .expect("metrics reservation address"); + let config = write_config( + gateway_address, + metrics_address, + backend_address, + frontend_address, + ); + drop(traffic_reservation); + drop(metrics_reservation); + let _process = start_gateway(&config, gateway_address, metrics_address); + + let (partial, termination, elapsed) = raw_request_until_terminal( + gateway_address, + b"GET /api/slow-drip HTTP/1.1\r\nHost: app.example:8080\r\nConnection: close\r\n\r\n", + ); + assert!(matches!( + termination, + DownstreamTermination::Eof | DownstreamTermination::ConnectionReset + )); + assert!( + elapsed < Duration::from_secs(1), + "response-body lifetime must stop the continuous drip instead of allowing completion: {elapsed:?}" + ); + assert!( + elapsed >= PRE_RESPONSE_HEADER_DELAY + RESPONSE_BODY_LIFETIME, + "lifetime must start at the final response header rather than request start: {elapsed:?}" + ); + + let header_end = partial + .windows(4) + .position(|window| window == b"\r\n\r\n") + .map(|position| position + 4) + .expect("slow-drip response must commit a complete header block before termination"); + let headers = String::from_utf8_lossy(&partial[..header_end]); + assert_eq!( + http11_status(&headers), + Some(200), + "post-commit lifetime failure cannot be rewritten as a second status: {headers:?}" + ); + assert_eq!(header_values(&headers, "Content-Length"), vec!["20"]); + let body = &partial[header_end..]; + assert!( + !body.is_empty(), + "body progress must cross the callback boundary" + ); + assert!( + body.len() < 20, + "configured lifetime must terminate before the declared body completes" + ); + + let readiness = get(gateway_address, "/readyz"); + assert_eq!(http11_status(&readiness), Some(200)); + let metrics = get(metrics_address, "/metrics"); + assert!( + has_exact_metric_sample(&metrics, "cwl_pingora_gateway_request_errors_total 1"), + "lifetime enforcement must remain visible through bounded error telemetry: {metrics:?}" + ); + let recovered = get(gateway_address, "/after-slow-drip"); + assert_eq!(http11_status(&recovered), Some(200)); + assert!(recovered.ends_with("\r\n\r\nrecovered")); + + frontend_origin + .join() + .expect("frontend recovery fixture should complete"); + backend_origin + .join() + .expect("slow-drip backend fixture should complete"); +} + +#[test] +fn slow_drip_evidence_parsers_reject_lookalikes() { + assert_eq!(http11_status("HTTP/1.1 200 OK\r\n\r\n"), Some(200)); + assert_eq!(http11_status("HTTP/1.1 2000 Weird\r\n\r\n"), None); + assert_eq!(http11_status("http/1.1 200 OK\r\n\r\n"), None); + + let headers = "HTTP/1.1 200 OK\r\nX-Content-Length: 20\r\ncontent-length: 20\r\n\r\n"; + assert_eq!(header_values(headers, "Content-Length"), vec!["20"]); + + assert!(has_exact_metric_sample( + "cwl_pingora_gateway_request_errors_total 1\n", + "cwl_pingora_gateway_request_errors_total 1" + )); + assert!(!has_exact_metric_sample( + "cwl_pingora_gateway_request_errors_total 10\n", + "cwl_pingora_gateway_request_errors_total 1" + )); +}