Skip to content

Chore/migrate keycloakx formation sdan - #1107

Closed
iliesmrf wants to merge 16 commits into
mainfrom
chore/migrate-keycloakx-formation-sdan
Closed

iliesmrf wants to merge 16 commits into
mainfrom
chore/migrate-keycloakx-formation-sdan

Conversation

@iliesmrf

Copy link
Copy Markdown
Contributor

Issues liées

Issues numéro:


Quel est le comportement actuel ?

Quel est le nouveau comportement ?

Cette PR introduit-elle un breaking change ?

Autres informations

@iliesmrf
iliesmrf force-pushed the chore/migrate-keycloakx-formation-sdan branch from 4f8bc2e to 189069f Compare September 14, 2026 14:50
Direct 26.1.x -> 26.7.x fails on keycloak/keycloak#51304
(LazyInitializationException in MigrateTo26_7_0.addOrganizationAdminRoles).
Land on 26.6.4 first, verify healthy, then bump to 26.7.2 separately.
dsc.keycloak.* config (config.yaml/releases.yaml) kept intentionally -
keycloakx still reads several fields from it. Do NOT touch
keycloakx.cluster.enabled until this is regenerated, pushed, synced,
and the keycloak Application confirmed gone on the live cluster.
preserveResourcesOnDeletion failed to protect pg-cluster-keycloak when
the keycloak app entry was removed (second confirmed occurrence, same
failure mode as cpin-hp). Bootstrap via mode: recovery from the
pre-step7-manual-backup taken ~1h before the incident, instead of a
blank initdb. Ongoing backup path deliberately distinct from the
recovery source to avoid the "Expected empty archive" WAL check.
Claude Code project config/skills are local tooling, not meant to be
tracked in this shared repo.
The keycloakx block only exposed `replicas` as overridable on the live
DsoSocleConfig CR, unlike keycloak which exposes its whole config
surface (repoSocle, cnpg, postgres sizing, plugin/provider URLs,
custom values). Bringing it to parity is a prerequisite for repointing
the keycloakx templates off dsc.keycloak.* and onto dsc.keycloakx.*.
@iliesmrf
iliesmrf force-pushed the chore/migrate-keycloakx-formation-sdan branch 3 times, most recently from 40ca2e7 to 6203894 Compare September 23, 2026 21:29
Brings this branch up to the same repo-level state as chore/migrate-keycloakx,
adapted to this env's specifics. Live behavior intentionally unchanged:
ansibleJobEnabled stays false, replicas stays 1, dsc.keycloakx.subDomain
stays "keycloakx" - this is a code-parity port, not a cutover.

- Add the entire roles/gitops/post-install/keycloakx/ role (client/realm/
  user provisioning) and post-install/keycloakx.yaml, ported wholesale -
  didn't exist on this branch at all yet.
- Bring config.yaml's keycloakx block to CRD-field parity (repoSocle,
  postgresPvcSize/postgresWalPvcSize - set to 15Gi/15Gi to match the live
  CNPG cluster's actual PVC size, verified via kubectl - pluginDownloadUrl,
  providerDownloadUrl, cnpg.imageName, usersGitOpsEnabled).
- Repoint dsc.keycloak.* -> dsc.keycloakx.* everywhere except where this
  env's subDomain values actually diverge (legacy dsc.keycloak.subDomain
  was "keycloak", dsc.keycloakx.subDomain is "keycloakx" - unlike
  chore/migrate-keycloakx where both matched). fullnameOverride,
  keycloak_domain and every consumer that resolves to the live production
  hostname/service name are pinned to the literal "keycloak" instead, to
  keep their resolved value unchanged now that the legacy block is gone.
  dsc.keycloakx.subDomain itself is left alone (separate, temporary/testing
  subdomain, not touched by this port).
- Fix the same apps/keycloak/values -> apps/keycloakx/values stale Vault
  path bug in every consumer (grafana, sonarqube, console, keycloakx's own
  plugin/provider download).
- Fix the same RBAC resourceNames bug (post-conf-job.yaml.j2: "keycloak" ->
  "keycloakx-admin-secret"), the same dead k8s_info secret-name lookup in
  vault-secrets/tasks/main.yaml, and the same argocd_app bug in
  vault-secrets/vars/keycloak-admin-password.yaml.
- Port the admin-tools fixes: gitops_local_repo crash (post_conf_job unset),
  post_vault_update override, kc.sh full path (official image has no
  /opt/keycloak/bin on PATH), keycloakx-admin-secret name/key, and the
  app.kubernetes.io/name=keycloakx pod label selector (verified live on
  this cluster, matches keycloak-0's actual labels).
- Add keycloak-admin-password-reset-isolated-pod.yaml (runs kc.sh
  bootstrap-admin as its own throwaway Pod instead of exec'ing into a live
  replica, avoiding the port-conflict crash the in-place exec hits).
- Remove the now-permanently-blocking transitional guard in template.yml.
- Delete dead legacy code: templates/keycloak/ chart dir,
  roles/gitops/post-install/keycloak/ role, post-install/keycloak.yaml.
- Drop the keycloak: block from config.yaml/releases.yaml.

Verified live on this cluster (formation-console context): CNPG cluster
healthy (3/3), live PVC sizes match what's now in config.yaml, keycloak-0's
actual pod labels/StatefulSet name confirmed via kubectl before pinning any
literal.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@iliesmrf
iliesmrf force-pushed the chore/migrate-keycloakx-formation-sdan branch from 6203894 to d81d560 Compare September 23, 2026 21:30
iliesmrf and others added 4 commits September 23, 2026 23:32
…over

Flips the flag gating keycloakx's post-install role (realm/client/user
provisioning, RBAC, secret-ansible.yaml.j2) and its cpn-ansible-job. Verified
before flipping: CNPG cluster healthy (3/3), most recent backup 3h31m old and
completed. This will likely rotate the keycloak-client-secret-* secrets
(currently 279d old, Bitnami-era) that harbor/gitlab/vault/sonarqube/awx/
grafana/argo/console clients use - those apps may need a restart to pick up
new secret values, same as any Kubernetes secretKeyRef (env vars don't
hot-reload into already-running containers).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Root cause of the recurring "invalid_grant: Invalid user credentials" on
the post-install job, on both cpin-console-hp and formation-sdan: Keycloak
only ever reads its bootstrap admin password once, at first boot - the
running instance never re-consumes it after that. But envs.yaml generated
it fresh via lookup('ansible.builtin.password', '/dev/null', ...) on every
single full-deploy run (no persistent disk to cache a password file on
across ephemeral job pods), and write.yml writes any non-null value
straight to Vault -> AVP -> keycloakx-admin-secret. Net effect: every full
deploy silently overwrote the stored password with one that no longer
matched Keycloak's real DB value, undoing any password reset the moment
another full deploy ran.

Fix: fetch the current Vault value for apps/keycloakx/values first
(main.yaml, same community.hashi_vault.vault_kv2_get module write.yml
already uses successfully), and have envs.yaml prefer that over generating
fresh - only falls back to a new random password when Vault genuinely has
none yet (true first bootstrap, or the read errors out, e.g. path doesn't
exist). Verified all three paths (existing value reused, Vault-read
failure falls back to fresh, value present but no auth key yet falls back
to fresh) against a standalone ansible-playbook run before applying.

A real password rotation (via keycloak-admin-password-reset*.yaml, which
writes to this same Vault path) now sticks across subsequent deploys
instead of being silently clobbered.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ord to Vault

Lost in translation when this script was copied from
keycloak-admin-password-reset.yaml: it never got post_vault_update: true
on its "Post-install vault-secrets run" task, so write.yml's "Update vault
secret" task silently skipped every time (post_conf_job is set, and
without this override that's enough to block the write even when a real
diff is detected). Confirmed live: a run with a valid Vault token still
produced zero writes until this was added.

Net effect before this fix: the direct k8s Secret patch worked, dsoadmin
could log in immediately - then the next ArgoCD sync re-rendered
keycloakx-admin-secret from the still-stale Vault value via its AVP
<path:...> annotation, silently reverting the fix. Confirmed live on
formation-sdan.

Also dropped the `when: dsc.global.gitOps.deploymentEnabled` guard that
keycloak-admin-password-reset.yaml has on this same block: on this env
that flag is false, which would skip the Vault write entirely regardless
of post_vault_update - same revert-on-next-sync problem. The Secret is
AVP-rendered from Vault either way, so the write needs to happen
unconditionally here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
usersGitOpsEnabled: true activates the CSV-based user sync path for the
first time on this environment, exposing a long-dormant bug present on
all three environment branches: sync_users.yml referenced
keycloak_current_vault_values (undefined - nothing sets this name), with
a .data.data.auth.userTmpPassword shape that doesn't match what
community.hashi_vault.vault_kv2_get actually returns anyway.

vault-secrets/write.yml already sets {{ argocd_app }}_current_vault_values
as a side effect of the "Call vault-secrets role" step earlier in
post-install/keycloakx.yaml - for keycloakx that's
keycloakx_current_vault_values, with the KV data under .secret (the
module's own result shape, not a raw Vault API response).
@iliesmrf iliesmrf closed this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant