Skip to content

Make Brainstore rollout controls configurable - #99

Closed
Alexey Soldatchenko (soldatchenko) wants to merge 4 commits into
mainfrom
codex/configurable-brainstore-rollouts
Closed

Make Brainstore rollout controls configurable#99
Alexey Soldatchenko (soldatchenko) wants to merge 4 commits into
mainfrom
codex/configurable-brainstore-rollouts

Conversation

@soldatchenko

@soldatchenko Alexey Soldatchenko (soldatchenko) commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Brainstore readers, fast readers, and writers maintain local caches and can perform substantial startup and processing work. During an upgrade, deployments may need different trade-offs between rollout speed, available capacity, and the number of replacement pods starting at once.

The chart currently hardcodes the Deployment rollout behavior for all three Brainstore roles. This change exposes standard Kubernetes rollout controls so operators can tune each role for their environment.

What changes

  • add independently configurable strategy, minReadySeconds, and progressDeadlineSeconds values for Brainstore readers, fast readers, and writers
  • validate that each progress deadline is greater than its readiness dwell so invalid Deployments fail during Helm rendering
  • render rolling-update settings only when the selected strategy is RollingUpdate
  • document rollout-only scope, Recreate outages, mixed-version upgrade considerations, and provide a conservative one-at-a-time rollout example
  • cover default, custom, invalid deadline, and Recreate rendering with unit tests

Upgrade impact

This change is backward-neutral for customers who do not configure the new values. The defaults preserve the effective behavior from earlier chart versions:

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 100%
    maxUnavailable: 0
minReadySeconds: 0
progressDeadlineSeconds: 600

Upgrading the chart with existing values does not change Brainstore rollout pacing. These settings are outside the pod template, so this feature does not restart Brainstore pods by itself. Custom settings apply when a later image, configuration, resource, or other pod-template change triggers a rollout.

Why operators may tune this

  • maxSurge limits the additional pods that can start during a rollout.
  • maxUnavailable controls how much existing capacity Kubernetes may remove while replacing pods.
  • minReadySeconds requires a replacement to remain Ready for a minimum period before Kubernetes considers it available and continues the rollout.
  • progressDeadlineSeconds defines how long Kubernetes allows the Deployment to make progress and must be greater than minReadySeconds.

For example, a cache-heavy deployment can use maxSurge: 1, maxUnavailable: 0, and a longer minReadySeconds to bound simultaneous cache warm-up and processing pressure while preserving its desired available replicas. This intentionally makes the rollout slower and requires capacity for the configured surge.

These settings pace Deployment-managed rollouts. They do not create Brainstore PodDisruptionBudgets or protect against node drains, evictions, or failures. Slower rollouts also extend the mixed-version window, so operators should follow version-specific upgrade guidance and avoid combining image changes with incompatible storage-format changes.

Validation

  • ./test.sh
    • 30 suites passed
    • 315 tests passed
    • AWS, Azure, GKE, and minimal chart rendering passed
    • strict Helm lint passed

@soldatchenko
Alexey Soldatchenko (soldatchenko) marked this pull request as ready for review September 3, 2026 03:39
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 3, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-09-03T03:42:24.418270Z 27ed1ed Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 27ed1edebd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".

Comment thread braintrust/templates/brainstore-reader-deployment.yaml

@erikdw Erik Weathers (erikdw) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additive and backward-neutral — defaults match the previous hardcoded rollout, the three roles stay in sync, and this does not restart pods by itself.

The one thing I would like fixed: pair minReadySeconds with progressDeadlineSeconds, or fail/document the 600s ceiling. Please also keep the rollout values/render shape compatible with API (and later gateway) — same Kubernetes field names, not a Brainstore-only wrapper. Fine not to convert the other services in this PR.

Inline notes on Recreate-on-writer, mixed-version windows, and the lack of a Brainstore PDB.

Nits, non-blocking: Recreate rendering is only tested on the writer; the strategy block is copied three times.

Comment thread braintrust/values.yaml
replicas: 2
# Configure rollout pacing for this Brainstore role. Defaults preserve the
# rollout behavior from earlier chart versions.
minReadySeconds: 0

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs a paired progressDeadlineSeconds on all three Brainstore roles.

Kubernetes defaults an omitted deadline to 600 and requires it to be strictly greater than minReadySeconds, so a dwell of 600+ is rejected at admission. Ready resets the progress clock, so the 60s readiness initialDelaySeconds and the dwell are sequential windows — the 300s example is fine — but the dwell still has to fit in one deadline window. Readiness flaps during the dwell restart minReadySeconds without counting as progress, which is an argument for margin, not a product ceiling.

The README sells longer dwell for cache warmup, so I would expose progressDeadlineSeconds next to this. Alternatively, fail when minReadySeconds >= 600 and document that dwell must stay below the deadline. fail / required is how this chart validates; there is no values.schema.json.

minReadySeconds: {{ .Values.brainstore.writer.minReadySeconds }}
strategy:
type: RollingUpdate
type: {{ .Values.brainstore.writer.strategy.type }}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recreate on the writer is a full role outage at the default replicas: 1 — every writer goes down before the replacement starts. If we expose strategy.type, the README should say that explicitly.

Comment thread braintrust/README.md
Comment on lines +308 to +311
Upgrading the chart without overriding these values does not change rollout
pacing. This feature also does not change the pod template, so it does not
restart Brainstore pods by itself. Custom settings take effect the next time a
pod-template change triggers a rollout.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A conservative maxSurge + dwell also widens the mixed-version window. Worth a short caveat not to combine a slow Brainstore rollout with a WAL-footer or image-compat bump — old nodes still rolling out cannot read a new WAL format.

Comment thread braintrust/README.md Outdated
Comment on lines +340 to +344
The same `strategy` and `minReadySeconds` settings are available under each
Brainstore role. A longer dwell reduces rollout pressure but does not prove that
a pod's local cache is fully warm; monitor workload health until the rollout has
converged. `maxUnavailable: 0` also requires enough cluster capacity for the
configured surge.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These knobs only pace Deployment rollouts. Brainstore has no PodDisruptionBudget (api-pdb.yaml is the only PDB in the chart), so minReadySeconds does not protect against node drains or evictions. Worth saying that so operators do not read this as disruption protection.

@soldatchenko
Alexey Soldatchenko (soldatchenko) deleted the codex/configurable-brainstore-rollouts branch September 4, 2026 15:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants