Skip to content

Azure IPAM v4.0.0 Release - #377

Open
DCMattyG wants to merge 257 commits into
mainfrom
ipam-3.7.0
Open

Azure IPAM v4.0.0 Release #377
DCMattyG wants to merge 257 commits into
mainfrom
ipam-3.7.0

Conversation

@DCMattyG

@DCMattyG DCMattyG commented Mar 11, 2026

Copy link
Copy Markdown
Contributor

Azure IPAM v4.0.0

This is a major release that delivers significant framework upgrades, a data grid migration, authentication modernization, comprehensive documentation overhaul, revamped examples, and numerous bug fixes.


Major Framework & Dependency Upgrades

  • React 18 → 19: Full migration including removal of forwardRef wrappers (ref-as-prop), removal of PropTypes, explicit null for useRef() calls, and migration of LoadingButton to Button
  • MSAL v4 → v5 (@azure/msal-browser 4.x → 5.x, @azure/msal-react 3.x → 5.x): Removed obsolete config, consolidated event types, fixed silent token timeout recovery (timed_out error code), and prevented iframe fallback timeout loops
  • Inovua React Data Grid → AG Grid (ag-grid-community / ag-grid-react 36.x): Complete migration to AG Grid including centralized DataGrid component, custom styling, column state persistence, unified data loading overlays, and custom cell renderers (drill-down, info, progress)
  • Vite 7 → 8 (vite 7.x → 8.x, @vitejs/plugin-react 5.x → 6.x): Migrated to Vite 8 which replaces Rollup with Rolldown and esbuild with Oxc for bundling, transforms, and minification. Removed vite-plugin-eslint2 (redundant with editor-based linting and incompatible with Vite 8)
  • ESLint 10: Upgraded from ESLint 9 to 10 with modern React linting plugins (@eslint-react/eslint-plugin v5, eslint-plugin-react-hooks v7 — which now consolidates the React Compiler lint rules per the React Compiler 1.0 release), replacing eslint-plugin-react. Added dist/ ignore, fixed no-useless-assignment violations, and removed unused eslint-plugin-jest. Resolved all remaining lint warnings as part of the React 19 modernization — useContextuse, <Context.Provider><Context>, ref naming conventions, stable list keys, hoisted static styled components, migration of SnackbarUtils to notistack's standalone enqueueSnackbar, and moved "ref assigned during render" patterns into useEffect
  • MUI v7 → v9 (@mui/material 7.3.x → 9.0.x, @mui/icons-material 7.3.x → 9.0.x): Major two-version jump. Migrated deprecated component props to the unified slots/slotProps API (largely via @mui/codemod), moved deprecated system props into sx, replaced the removed Unstable_Grid2 with the new default Grid (using size={{ xs: N }}), renamed removed Outline (no "d") icon exports to their Outlined counterparts, and removed @mui/lab (no longer needed — LoadingButton's loading prop is now native to Button)
  • React Router 7 → 8 (react-router 7.x → 8.x): Major version upgrade. The UI uses declarative-mode routing (BrowserRouter / Routes / Route) with all imports already sourced from react-router, so no application code changes were required. v8 raises the minimum runtime to Node.js 22.22.0 (and React 19.2.7+)
  • Updated NPM packages across the board (Vite 7.3.x, React Router 7.13.x, MUI 7.3.x, etc.)

Engine & Backend

  • Azure Function Blueprints: Implemented Blueprint-based function naming for improved clarity
  • Python dependency cleanup: Removed msal, azure-common, azure-keyvault-secrets, and six; added azure-mgmt-resource-subscriptions to address Azure SDK module separation
  • Python linting: Added pyproject.toml with Ruff linter configuration (pycodestyle, pyflakes, isort). Resolved all lint violations across 18 engine files — bare except clauses, wildcard imports replaced with explicit names, unused imports/variables, invalid escape sequences, import sorting and grouping
  • Reservation logic hardening: Fixed CIDR overlap detection during auto-fulfillment, added validation for all in-block vNet prefix overlap checks, and made auto-fulfillment idempotent by deduping existing block vNet associations
  • Endpoint fix: Standalone NICs are now included in the Endpoint list (fixes Standalone NICs not listed in Endpoint node #371)
  • Network associations fix: Resolved improper handling of missing vNETs and vHUBs (fixes Error fetching available IP Block networks #350)
  • Next available network fix: nextAvailableVNet aborted its entire Block list search when the first Block could not satisfy the requested size under smallest_cidr, returning a 500 instead of evaluating the remaining Blocks. max() was called on an empty candidate list, and the resulting ValueError escaped the loop. The same missing guard in nextAvailableSubnet and in single-Block Reservation creation replaced their intended error messages with an unhandled exception (fixes nextAvailableVNet fails when using multiple blocks and smallest_cidr option true if no IP range is available in first block #384)
  • HTTP status codes corrected for client and capacity errors: Conditions the API explicitly anticipates were reported as 500, making them indistinguishable from genuine server faults, prompting pointless client and proxy retries, and counting against server-error metrics. Requests that cannot be satisfied by current allocation state now return 409, as does a JSON patch that conflicts with the stored document; a malformed JSON patch now returns 400. Both already matched the codes used by neighboring checks in the same handlers. Errors originating from Azure, Cosmos DB, or the site configuration are unchanged and still return 500
  • Development reload flag removed from production startup: uvicorn was started with --reload in both production init scripts. Under a read-only run-from-package mount its watcher tracked thousands of files that can never change, and its supervisor process stayed alive when the application crashed — so the container kept running while serving nothing, and health checks failed against an apparently healthy process. A startup failure now exits, so the platform restarts the application and the fault is visible. engine/Dockerfile.dev retains the flag, where hot reload is intended
  • TLS validation moved to the OS trust store: The JWKS fetch backing token validation used requests, which verifies against the CA bundle shipped inside certifi rather than the operating system trust store. WEBSITES_INCLUDE_CLOUD_CERTS populates the OS store, so sovereign cloud roots were present but invisible to that code path — secret cloud (IL6) deployments rejected every token with CERTIFICATE_VERIFY_FAILED while passing in every commercial cloud, where certifi covers the endpoints. Every other outbound call already used aiohttp, which reads the OS store; the JWKS fetch was the lone exception and now matches. The metrics heartbeat moved off requests as well, and a Ruff banned-api rule now fails the build on import requests or import httpx, since the defect is invisible to commercial-cloud testing and would otherwise return unnoticed
  • Entra signing keys cached: Token validation fetched the tenant JWKS document and rebuilt the RSA public key from its JWK on every authenticated request, spending a TLS handshake, an HTTP round trip and a key conversion per API call. Keys are now cached in process following Microsoft's signing key rollover guidance — a one hour refresh with jitter, a twenty four hour hard expiry, and a forced refresh on an unknown kid floored at once per five minutes. Concurrent refreshes collapse to a single fetch, and a failed refresh serves the last known good keys rather than failing closed. The floor also closes an amplification vector, where a flood of tokens bearing bogus kid values previously produced one outbound fetch each on behalf of unauthenticated callers
  • Address space checks no longer pay for full network detail: Every occupancy and utilization check loaded each network in the tenant complete with its subnets, peerings and Cosmos DB parentage, then read nothing but the resource ID and address prefixes. On a large tenant that is a subnet mv-expand of up to 1,024 rows per virtual network, two peering joins, one ARM call per virtual hub and two Cosmos DB queries — on every request — and a direct contributor to the Resource Graph throttling that surfaces as failed allocations. A prefix-only query now backs the thirteen call sites that need nothing further. Most expand responses still use the detailed query because they return whole network objects, but the available-networks endpoint is an exception: its expanded response is modelled on NetworkExpand, whose six fields are exactly what the prefix-only query already returns, so every additional field the detailed query produced was discarded field-for-field. That endpoint backs the network association picker in the UI, which always requests expansion, so the most frequently exercised expanded path now costs one Resource Graph query rather than two. Nothing is cached: results are still read live on every request, so allocation decisions remain based on the current state of Azure. Verified against a live tenant of 158 networks as returning an identical network set, identical prefixes and an identical total address count, with utilization responses byte-for-byte unchanged
  • Azure Resource Graph failures are now classified instead of reported as access denied: Every non-success response from Resource Graph was caught as a bare HttpResponseError and re-raised as 403 Access denied, so throttling, upstream outages and expired credentials all presented as authorization failures — sending users to audit permissions for a fault that had nothing to do with them. Because ClientAuthenticationError subclasses HttpResponseError, the Token has expired handlers in the three Resource Graph wrappers were unreachable. Failures are now classified at a single choke point: throttling returns 429 with Retry-After, upstream and unexpected statuses return 502 rather than being blamed on the caller, authentication failures return 401 — or 500 when the Azure IPAM service principal itself failed to authenticate — and a genuine Forbidden still returns 403. The HTTP exception handler now propagates response headers so Retry-After reaches the caller. Percent-style placeholders in loguru calls, which silently dropped their arguments, were converted at the same time
  • Data Factory discovery now honours subscription exclusions: Data Factory lookup built its own Resource Graph client around a hardcoded query instead of going through the shared query path, so it never applied the configured subscription exclusions and returned factories from subscriptions an administrator had deliberately excluded. It now uses the shared path with a DATA_FACTORY query carrying the standard exclusion filter, which also removes a duplicated Resource Graph client and its per-call setup
  • Virtual Machine Scale Set discovery now skips excluded subscriptions: Scale Set discovery runs through the Azure SDK rather than Resource Graph, and enumerated every subscription visible to the credentials — so excluded subscriptions were still queried, returning resources the administrator had excluded and spending an SDK call per subscription to do it. Subscription enumeration now accepts the exclusion list and filters before any Scale Set call is made
  • Scale Set addresses no longer fail on the sovereign cloud discovery path: The Scale Set endpoint expanded each record over a private_ips array, which is the shape Resource Graph returns. The SDK path used where Resource Graph coverage is incomplete — sovereign clouds such as IL6 — returns a singular private_ip, and those records raised a KeyError instead of being returned. Both shapes are now handled
  • vWAN transit network matching no longer depends on a regular expression: When a Virtual Network connects to a vWAN hub, Azure peers it to a generated transit network named HV_<hub>_<suffix>, which Azure IPAM substitutes back to the hub itself for display. The hub name was interpolated into a pattern without escaping, and Azure permits periods in hub names — so a hub named my.hub produced a pattern in which the period matched any character, and myXhub would have matched too. The pattern was also case-sensitive against a Resource Graph identifier. The transit network is now matched as a lowercased substring, which is all the comparison ever needed and removes both problems
  • Azure resource IDs are now compared case-insensitively: Azure returns resource IDs with inconsistent casing — Microsoft documents this and directs callers to "always perform a case-insensitive comparison of names" — and Azure IPAM stores whichever casing the client supplied when a network was associated, since association validates case-insensitively but saves the request body verbatim. Fifteen comparisons matched those two values exactly, so a casing difference made a managed network invisible: utilization under-reported the address space in use, Block CIDR and external network validation skipped the network entirely, reconciliation marked it inactive, and the CIDR check reported it as belonging to no Space or Block. Every one of these failed in the same direction — occupied space reading as free, which is the failure mode an IPAM can least afford. One of them compared a Resource Graph ID against an ARM SDK ID, where the two APIs genuinely disagree on casing. Reservation identifiers, which Azure IPAM generates itself, remain case-sensitive
  • Scale Set discovery no longer depends on resource ID casing: Six expressions pulled the resource group, virtual network, subnet and instance number out of ARM resource IDs by matching the resourceGroups/, providers/, virtualNetworks/, subnets/ and virtualMachines/ segments literally. Unlike the comparison bugs above these failed loudly rather than silently — a non-matching expression returned no match and the immediate .group(0) raised, so the whole Scale Set request became a 500 — but the trigger is the same casing inconsistency, and the same resource can be returned with different casing by different Azure APIs. This is the SDK-based Scale Set path, which runs only outside Azure Public; commercial deployments answer the same request from Resource Graph and were never affected. That makes it a sovereign and air-gapped cloud fix, which is also where Azure IPAM depends on the SDK paths most, because Resource Graph coverage there is incomplete. The segment matches are now case-insensitive, which is what the equivalent expression in the Reservation path already did
  • Block utilization is computed in one place: The arithmetic that populates size and used on a Block, its networks and their subnets was duplicated across the four endpoints that report utilization — roughly thirty lines repeated four times, differing only in local variable names and in whether the running Space totals were accumulated alongside. That duplication is what allowed the Space utilization defect above to exist in exactly one of the four copies. The four copies collapse to a single add_block_utilization helper, with the Space totals now summed from each Block's own figures, which is arithmetically the same value since the Space totals were only ever the sum of their Blocks. Around seventy lines are removed. The behaviour is deliberately unchanged, including two long-standing quirks in the per-network figures that are documented and fixed separately, and equivalence was verified both against the original logic across every network shape — fully in-Block, partially in-Block, entirely outside the Block, virtual hubs, networks without subnets, unmatched associations and empty Blocks — and byte-for-byte against a live tenant across all twenty-eight utilization responses
  • External networks now report utilization the way Azure networks do: A Block counted each external network's full address range toward its own utilization but reported nothing about how that range was allocated internally, even though external subnets and their endpoints were added to the data model later. External networks and their subnets now carry size and used alongside the Azure networks they sit beside. An external network's used is the address space assigned to its subnets, mirroring how a Virtual Network's used is the space assigned to its subnets; an external subnet's used is the number of endpoints defined within it, mirroring the consumed addresses reported for an Azure subnet. Endpoints are counted as they are, without the five addresses Azure reserves in every subnet, because an external network is not Azure and reserves nothing. The Block's own used is deliberately unchanged and still counts each external network's full range exactly once, since its subnets sit inside that range and counting both would double count. This required utilization variants of the external network and subnet models, so the utilization responses gain two fields at each level and lose none
  • A network's utilization within a Block is measured on one basis: For an expanded network, size counted only the prefixes falling inside the Block while used counted the subnets of every prefix the network owns, including those outside it. The two figures answered different questions and were presented as a ratio, so a network whose address space straddles a Block boundary could report more space used than it has — on the verification tenant one reported size 256 against used 384. Separately, used was reset inside the loop over in-Block prefixes, so a network with no prefixes inside the Block never had it reset at all and reported the entire network's subnet total against a size of zero. used is now initialized once and counts only subnets that fall inside the Block, so both figures describe the same address space and used can no longer exceed size. Block and Space totals are unaffected because they never consulted subnets, and each subnet still reports its own size regardless of where it sits. The Azure IPAM interface never requested these figures, since it does not expand Spaces or Blocks, so this corrects the documented API for automation consumers rather than anything visible in the product
  • Unreadable Reservation tags are reported rather than silently ignored: The helper that splits a Reservation tag into IDs wrapped its work in a bare except returning an empty list, which conflated two unrelated situations — a network carrying no tag at all, which is the normal case for nearly every network, and a tag whose value could not be read. The Resource Graph query parses tag values as JSON so that an absent tag resolves to null, with the side effect that a value resembling JSON arrives as a list, object or number instead of text. Such a network was then treated as having no Reservation, so the Reservation waited to be fulfilled indefinitely and nothing was written to the log. The type check is now explicit, and reconciliation reports the network ID and the offending value once per run rather than once per comparison. Behaviour is otherwise unchanged and was verified identical across every input type the tag can produce: unreadable values are still skipped rather than guessed at, and no attempt is made to strip quotes or interpret JSON, since Resource Graph already removes surrounding double quotes and interpreting the rest would mean acting on a tag the user did not write as an ID
  • Azure Resource Graph throttling is now retried by the SDK rather than failing outright: Resource Graph reports throttling on a POST and states the wait in x-ms-user-quota-resets-after instead of the standard Retry-After. The Azure SDK's retry policy therefore never retried it — POST sits outside its method allowlist, and the Retry-After short circuit that would have overridden that never fires — so a throttled query failed on its first attempt even though 429 is in the SDK's own retryable set. The SDK's policy is now taught both facts rather than a second retry loop running alongside it, so it waits exactly the interval Resource Graph reported instead of backing off blindly, and a hand-picked backoff ceiling gives way to the SDK's own settings. One retry setting is deliberately overridden: Resource Graph quota is shared by every caller using the same credentials rather than allocated per user, and replenishes roughly every five seconds, so the SDK's default of three status retries could be spent in about fifteen seconds while a burst of requests was still draining. That allowance is raised to six, which bounds a throttled request at roughly thirty seconds of waiting rather than a failure. It is not raised further on purpose — beyond ten the total retry budget silently becomes the real limit, and a longer wait encourages users to reload the page, which adds load to the very quota the request is waiting on
  • Truncated Resource Graph results are no longer treated as complete: Resource Graph withholds the continuation token when it truncates a result set, so the missing rows cannot be paged for and the response is silently short while still reporting success. For an IPAM that is the most dangerous shape of failure, since an unseen network reads as free address space. The condition is now detected and reported as an error rather than returned as though whole
  • Space utilization no longer fails when a Block holds an external network: GET /api/spaces/{space}?utilization=true returned 500 for any Space with an external network in one of its Blocks. External address space was accumulated onto the space path parameter — a string — rather than the Space document, raising TypeError: string indices must be integers. The line was correct in the equivalent list handler, where space is the loop variable, and was copied into the single-Space handler without renaming. The list endpoint and the Block-level endpoints were never affected, which is why the fault went unseen
  • Virtual hubs now expand in Space and Block responses: Requesting expand=true on a Space or Block containing a vWAN hub silently returned unexpanded network references — just an ID and active flag — with an HTTP 200 and no error. The expanded network model requires a subnets field, which virtual networks carry and virtual hubs never had, so validation of the expanded shape failed and the Union response model quietly fell through to the plain reference shape. The effect was not limited to the hub: a single hub in a Block collapsed every network in that Block back to a reference. Virtual hubs now carry an empty subnet list, which is what they genuinely have, so the expanded shape validates. The Block networks and available-networks endpoints were never affected, as their response model does not require subnets
  • Virtual hub listing no longer fails once a hub is associated to a Block: GET /api/azure/vhub returned 500 as soon as any vWAN hub was associated to a Block, reporting Input should be a valid string for parent_block. A network can belong to Blocks in more than one Space, so the engine has always produced a list of Block names here; the response model declared a single string. The model was corrected to match the data rather than the reverse, since collapsing to one name would discard a real association. The defect was invisible on /api/azure/network, which returns the same hub data but declares no response model, and on the virtual network equivalent for the same reason

UI & UX Improvements

  • Drill-down navigation in Discover (fixes Navigation Improvements #183): Added hierarchical drill-down with multi-filter pass-through and hidden column auto-reveal when filters are active
  • Centralized authentication handling: New AuthHandler for MSAL error handling, centralized token acquisition via tokenService
  • Custom DraggablePaper component: Replaced react-draggable package with a purpose-built component
  • Associations UX: Shifted API work to Redux thunks, optimized data refresh, fixed infinite update loops by separating initial grid selection from user selection state
  • Planner data loading: Addressed issue where Planner would not load when data was incomplete (fixes Planner will not load #366)
  • Reservation UX: Fixed view reversion after cancelling a Reservation, fixed column sorting, and streamlined the interface
  • Unified grid experience: Centralized DataGrid and ConfigureGrid components with shared filter utilities, consistent loading overlays, and AG Grid custom styling (brightness filter for row hover)
  • Search bar: Fixed messaging when no resources are found in Azure IPAM scope
  • API errors without the standard envelope are now surfaced: The Axios response interceptor read error.response.data.error unconditionally, so any response that did not use the { error } envelope produced an Error with an empty message — and an error snackbar with no text at all. Unhandled engine exceptions return plain text, request validation failures return { detail }, and proxy errors return HTML; each is now resolved to a message, falling back to the Axios status text. This also fixes silent failures on any request rejected by model validation, which had always returned 422

Notifications & Service Management

  • New notification framework (GET /api/notifications, POST /api/notifications/{id}/resolve): A self-describing, API-first advisory system. Server-side detectors emit notifications that the UI — or any API / IaC consumer — can read, while remediation is resolve-by-reference: the client asks the backend to resolve a notification by id and the backend owns all the logic (admin-gated, with an active-notification guard). Adding a new advisory is a single detector module registered in one list
  • Registry-migration detector: Flags deployments still pulling container images from the legacy public registry (azureipam.azurecr.io, critical) or the development registry (azureipamdev.azurecr.io, warning) and offers a one-click remediation that repoints the App Service / Function LinuxFxVersion to registry.azureipam.com. The target image is validated as anonymously pullable before any change is applied, then the app restarts to pull it — complementing the update script's auto-migration with an in-app path
  • Update-available detector: Compares the running IPAM_VERSION against the latest published GitHub release and surfaces an informational notice linking to the update guide. Fails safe (no notification, no error) on network or rate-limit failures
  • Platform-migration detector: Flags deployments still running on the legacy Docker Compose App Service model (DEPLOYMENT_STACK == "LegacyCompose", critical) and links to the migration guide ahead of Microsoft's March 31, 2027 retirement of Docker Compose support for Azure App Service. Link-only guidance (no in-app remediation) since migration is a scripted, multi-step process run from the operator's workstation
  • Notification center (UI): AppBar bell with a severity-colored unread badge, a severity-ordered list (popover on desktop, full-screen sheet on mobile), a detail dialog with server-driven actions and early admin gating, per-item dismiss, mark-all-read, and a manual refresh that resyncs to the server's truth
  • Service restart gate (UI): A full-screen overlay shown while the backend restarts (e.g. after a registry switch). The live SPA stays in memory, polls /api/status, and confirms recovery via a changed service start time before reloading — with calm, time-keyed messaging and a deliberate manual-reload escape hatch if recovery runs long. Background polling is paused while the gate is up

Deployment, Build & Infrastructure

  • Configuration drift detection and remediation in update: The update script now compares an existing deployment against what a fresh deployment would produce today, presents the differences, and converges them on approval. Detection is based on live Azure resource state rather than the version originally deployed, so each difference is evaluated independently and a partially-current deployment only sees what it actually needs. Covers the container registry endpoint, Python runtime version, App Service startup command, health check, baseline app settings, creation of the staging slot, and Function App slot-sticky content-share settings. Existing values and user-added settings are never overwritten, and removal is restricted to settings Azure IPAM owns that no longer apply to the deployment's shape. Production is converged before the staging slot so a newly created slot inherits corrected values, and the two are kept in sync thereafter so a future swap cannot regress production. When an archive is supplied explicitly, the Python runtime version is read from that archive rather than from the local checkout, so LinuxFxVersion is retargeted to match the wheels actually bundled in it — those wheels are ABI-specific, and a mismatch leaves every native module un-importable at startup
  • Version-aware update decisions: Version comparisons across all three update paths (public registry image label, private ACR repository build, and native GitHub release) now use semantic versioning rather than string equality, so prerelease suffixes order correctly. Downgrades are no longer silently applied and reported as updates — the update stops and points to -Force. Images built without a version stamp report the 0.0.0 Dockerfile default and are treated as an unknown version rather than a downgrade
  • Updates against stopped or unreachable deployments: The update script now detects application state up front and reports it. Configuration changes, image builds, and ZIP deployments all continue to work, restarts that would achieve nothing are skipped, and a stopped application is reported as remaining stopped. A new -ContainerType (Debian|RHEL) override on update matches the existing migrate switch, so a private ACR deployment can still be rebuilt when its container distro can't be probed. The distro probe is also now bounded by a timeout, so an application whose container fails to start no longer stalls the update indefinitely
  • Guidance instead of failures for anticipated conditions: update and migrate no longer raise exceptions for situations with a known remedy, such as an undetectable container distro. These now print actionable guidance under the relevant phase heading and exit cleanly, matching how legacy Compose deployments and out-of-resource-group registries were already handled. Genuinely unexpected errors now surface their message on screen rather than only in the log
  • DOCKER_REGISTRY_SERVER_URL removed from all deployment templates: This setting is only required for registries authenticating with stored credentials. Azure IPAM pulls anonymously from the public registry or with a managed identity from a private ACR, so it was never needed, and App Service removes it on its own whenever the container registry is reconfigured — which caused the update script to repeatedly report it as configuration drift
  • Azure PowerShell requirements aligned to 11.4.0: The deployment, update, and migration guides now state a single Azure PowerShell rollup version, and each script's granular #Requires module pins match the versions packaged in that rollup. Previously the update guide understated its requirement (Az.Resources 6.16.0 is only available from Az 11.4.0), and deploy.ps1 pinned the Az 10.3.0 module set while its documentation stated Az 11.0.0
  • Fixed Debian container build port: update.ps1 built Debian images with PORT=80 while deploy.ps1, migrate.ps1, and the Dockerfile default all use 8080
  • Public container registry migrated: The publicly hosted Azure IPAM container images have moved from azureipam.azurecr.io to the new registry.azureipam.com endpoint. The deployment and migration Bicep templates now reference the new registry by default, while the update script continues to recognize the legacy azureipam.azurecr.io endpoint so existing deployments keep working. The Docker Compose migration tooling intentionally still targets the legacy endpoint, as it only ever processes pre-existing legacy deployments
  • Migration robustness: The Docker Compose migrate script now halts on a non-standard or unresolvable container registry with guidance to re-run using the -JsonFile override, instead of silently falling back to the public registry. A new -ContainerType (Debian|RHEL) override handles cases where the source app is stopped/unreachable and its distro can't be auto-detected
  • RHEL images updated to UBI9: Including the container-image overrides in the deploy, update, and migrate PowerShell scripts
  • Node.js minimum raised to 22.22.0: Required by React Router v8; enforced in the build.ps1 version gate and the UI package.json engines field
  • Dockerfiles optimized: Improved layering, removed unneeded steps across all container images. Resolved Hadolint lint violations (ADDCOPY, JSON notation for CMD/ENTRYPOINT, pipefail for piped RUN commands). Added centralized .hadolint.yaml for rule suppressions
  • KeyVault Soft Delete added to deployment and migration Bicep templates (fixes Enable soft delete for Key Vault #373)
  • Build flexibility: Added support for building with either current or latest NPM/Python packages
  • CI/CD path exclusions: Added exclusions to avoid unnecessary test/build runs for documentation-only changes
  • GitHub Actions modernized: Bumped all workflow actions to their latest versions (checkout@v7, setup-node@v7, setup-python@v7, github-script@v9, create-github-app-token@v3, azure/login@v3, hadolint-action@v3.5.0), initially to clear the Node.js 20 runner deprecation and then to the current majors so the release does not ship already behind. The v7 line of the actions/* set is an ESM migration on the same node24 runtime, and none of the inputs these workflows pass were changed or removed. The fork-PR checkout restriction added in checkout@v7 does not apply here, as no workflow uses pull_request_target or workflow_run
  • GitHub Actions pinned to commit SHAs: Every action reference across all four workflows is now pinned to a full-length commit SHA with a trailing version comment, rather than a mutable tag such as @v7. A tag can be retargeted to arbitrary code by anyone able to push to the action's repository — the mechanism behind the tj-actions/changed-files and codfish/semantic-release-action compromises — and every one of these workflows holds credentials, whether an OIDC federated identity or a GitHub App private key. A new .github/dependabot.yml tracks the github-actions ecosystem weekly as a single grouped PR, with a seven day cooldown so a compromised release has a window to be reported before it is proposed here. Dependabot updates the SHA and the version comment together, so the pins do not go stale. Verified with the GitHub REST API that each pinned SHA is the commit the corresponding release tag resolves to
  • Azure PowerShell SDK v14 compatibility: Fixed Get-AzAccessToken breaking changes (fixes Breaking changes to Get-AzAccessToken #343); updated deploy & migrate scripts
  • OCI version labels on container images: All production images now carry standard OCI labels (org.opencontainers.image.version, .title, .source), stamped at build time via a new IPAM_VERSION build arg. This lets you determine the exact version behind a floating tag like latest by inspecting the registry — no pull or run required (e.g. docker buildx imagetools inspect or az acr manifest show). Covers all deb, rhel, and func variants across the root, engine, ui, and lb images, with the build workflow passing --build-arg IPAM_VERSION to every az acr build. Labels begin with the v4.0.0 release images
  • Release tag parsing hardened: Version extraction now strips only a leading v (^v) from the release tag, preserving suffixes such as -preview
  • Internet-restricted cloud deployments fixed: The ZIP-deploy path used by internet-restricted clouds (AZURE_US_GOV_SECRET) never resolved its bundled Python packages. init.sh referenced an APP_PATH variable that is defined nowhere in the repository — App Service exposes it only to interactive SSH sessions via ~/.bashrc, which a non-interactive startup command never reads — so PYTHONPATH expanded to a non-existent path at the filesystem root and no bundled dependency could be imported. The application root is now derived from the script's own location, matching what function_app.py already did for the Function App entry point. The accompanying PATH export was removed: it pointed at packages while pip installs console scripts to packages/bin, and nothing invokes them
  • Deployment archive now targets the App Service runtime rather than the build host: pip install --target resolves wheels for the machine running pip. When the GitHub Actions runner moved to Ubuntu 24.04, cryptography began resolving to a manylinux_2_34 wheel that cannot load on the App Service Python 3.11 image (Debian bullseye, glibc 2.31), producing GLIBC_2.33 not found at startup. Wheel resolution is now pinned explicitly (--only-binary=:all:, --platform manylinux2014_x86_64, --implementation cp, --python-version, --abi), with the Python tag and ABI derived from engine/app/version.json. manylinux2014 (glibc 2.17) is targeted deliberately, since sovereign clouds can run older stamps than commercial Azure
  • Native module verification gate: The build now fails if any bundled shared object targets the wrong Python ABI, requires a newer glibc than the platform tag allows, or is a Windows .pyd. Both defects above were invisible at build time and only surfaced after deployment — and the ABI mismatch presented as ModuleNotFoundError, which reads like a missing dependency rather than a build fault. The gate also inspects each distribution's *.dist-info/WHEEL metadata and rejects any wheel whose platform tag exceeds the target, so a locally compiled linux_x86_64 build or an over-new manylinux_2_28/manylinux_2_34 variant is caught by name before the binary scan has to find it. The highest glibc symbol required by any bundled module is now reported on every build, not only on failure
  • Deployment archives identify themselves: Every archive now carries a build.json at its root recording the build timestamp, the app and Python versions, the pip platform, the glibc ceiling and the highest glibc actually required, the native module count, and the resolved wheel tag for every package. In an air-gapped cloud a build log cannot be copied out, so establishing which build is actually deployed previously meant inference from indirect evidence; it is now a single cat
  • Sovereign cloud root certificates now trusted: Secret cloud deployments (AZURE_US_GOV_SECRET, IL6) run against endpoints whose certificate chains are issued by that cloud's own roots, which are not present in the App Service image's default trust store. Every outbound TLS call the engine makes — Key Vault references, Cosmos DB, ARM, Microsoft Graph — therefore failed certificate validation. WEBSITES_INCLUDE_CLOUD_CERTS is now set for that cloud in the deployment, update, and migration templates, so the platform injects the cloud's root certificates. The update script's configuration drift detection also adds it to existing secret cloud deployments. This setting is necessary but was not sufficient on its own — see the TLS validation fix under Engine & Backend, without which the engine still could not validate tokens in that cloud. The setting is gated on the cloud rather than on the deployment shape, so it now also reaches container deployments in that cloud, which the previous nesting excluded
  • Managed identity role assignments no longer collide on redeploy: Deployments that pin resource names via -ResourceNames failed with RoleAssignmentUpdateNotPermittedTenant ID, application ID, principal ID, and scope are not allowed to be updated — whenever the managed identity had been deleted and recreated. Role assignment names are globally unique GUIDs, and the Contributor and Managed Identity Operator grants seeded theirs from the identity's resource ID, which survives a delete and recreate unchanged. The name therefore matched an existing assignment while the principal behind it had changed, which Azure treats as an illegal update rather than a create. Every other module already seeded from the principal ID and was never exposed; managedIdentity.bicep could not, because a role assignment name must resolve at the start of deployment (BCP120) and the principal ID does not exist until the identity is created. Both grants moved into a new managedIdentityRoles.bicep module that receives the principal ID as a parameter, following Microsoft's documented guid(scope, principalId, roleDefinitionId) pattern — a recreated identity now yields a new assignment name and deploys cleanly, with no manual cleanup. Because the names change, a first re-run of deploy.ps1 against an intact deployment created by an earlier version may report RoleAssignmentExists; removing the two superseded assignments clears it. Deployments that let Azure IPAM generate resource names are unaffected either way, as every run produces fresh names. update and migrate needed no equivalent change — update deploys no role assignments at all, and migrate references the identity as existing and already seeds from the principal ID
  • Build script failure reporting: Seven prerequisite checks ended in a bare exit, which returns 0. A build that could not find NodeJS, or that rejected the installed Python version, printed errors and then reported success to CI. All now exit non-zero. The exception handler also printed an unassigned variable in place of the log path
  • Faster archive creation: Compress-Archive took 353 seconds to package the 6,404-file deploy archive; ZipFile.CreateFromDirectory produces an identically sized result in 11 seconds. The new API also preserves Unix file modes, which Compress-Archive flattened to 0644
  • CI Python version single-sourced: All three workflows now read the interpreter version from engine/app/version.json instead of pinning 3.11 by hand. This matters most in the versioning workflow, which regenerates requirements.lock.txt — dependency resolution is Python-version sensitive, so a lock file produced on the wrong interpreter can omit packages the runtime requires

Documentation Overhaul

  • Massively revamped How-To docs: authentication, exclusions, Discover, Reservations, External Networks, and Virtual Network Associations sections
  • New documentation: Comprehensive automation docs, API docs for vNet Associations, initial External Networks docs, detailed Reservations feature docs
  • 50+ screenshots added or replaced to reflect the current UI
  • Fixed all markdown warnings/errors and added .markdownlint.json configuration
  • Doc cleanup: Fixed stale links, grammar, spelling issues, deprecated folder descriptions, and undocumented switches across all sections
  • New troubleshooting entry: Deployments in internet-restricted clouds that report success without changing the running application — how to confirm which package is actually mounted, package retention and rollback, and the expected virtual-environment warning

Examples

  • Terraform example revamped: Migrated from Shell scripts to the official Azure IPAM Terraform provider
  • Azure ESLZ example modernized: Updated resource API references, modern coding standards, and improved parameterization
  • Script examples reorganized: PowerShell and Shell scripts moved into a dedicated examples/scripts/ folder with new helper scripts and README
  • Token helper function: Standardized access token generation; removed legacy Microsoft Graph SDK v1 support

Testing

  • Added tests for Virtual Network Association permutations
  • Expanded overall testing coverage with numerous additional Pester tests
  • Updated test expectations to align with additionally created resources
  • Added Tools coverage for Block list evaluation — falling through to a later Block when an earlier one cannot satisfy the requested size (with and without smallest_cidr), and the genuinely exhausted case
  • CI lint gate: Added pre-deployment lint job to the testing workflow (ESLint, Vite build verification, Ruff for Python, PSScriptAnalyzer for the PowerShell scripts, Bicep template validation, Hadolint for Dockerfiles). Deploy is skipped if any check fails, preventing wasted Azure resources
  • PowerShell scripts are analyzed in CI: No PowerShell file had ever been linted automatically, despite the repository carrying a PSScriptAnalyzerSettings.psd1, which is how the test suite came to be the only script not following the house conventions and the only one carrying an analyzer finding. The lint job now runs PSScriptAnalyzer over every script and fails on a finding at any severity, and also fails if it matches no scripts at all, so the check cannot pass green having analyzed nothing. The step is deliberately scoped to the four rules the settings file configures: running PSUseCorrectCasing alongside the full default rule set trips a thread-safety defect in the analyzer's command cache and aborts the run roughly a third of the time, which would fail the build on a clean tree. That is PSScriptAnalyzer #1708, open since 2021 and reproducible in both 1.24.0 and 1.25.0, so pinning an older analyzer does not avoid it. The scoping is annotated in the workflow with the single change needed to undo it once the defect is fixed
  • PowerShell conventions applied uniformly: The Pester suite used capitalized Function and Param, making it the only script not following the lowercase keyword style the settings file describes and every other script applies, and it carried the repository's only analyzer error — a false positive from normalizing an Azure access token to a SecureString, now suppressed inline with the same justification already used in the deployment scripts. version.ps1 declared pipeline binding on all four parameters without a process block, so piping several objects would have silently processed only the last; nothing pipes to it, and no comparable script declares that binding, so the unused bindings were removed. Every script also declared $logPath and then referenced $logpath on the following line, a mismatch copied into all five. Every PowerShell file in the repository now reports no analyzer findings at any severity
  • Transient failure handling in the API test suite: The integration suite failed outright on the first transient error, so an Azure Resource Graph throttle during a run looked like a product regression rather than an environmental hiccup. The shared request helpers now retry only 429 and 5xx responses, honouring Retry-After; the client errors the suite deliberately asserts are never retried, and transport failures are never replayed, since a POST may already have applied server-side. Access tokens are cached and refreshed ahead of expiry, so a long run no longer fails partway through on an expired token
  • Utilization and network expansion coverage: Added a Utilization & Expansion context exercising utilization=true and expand=true across the Spaces, Space, Blocks and Block endpoints. Neither parameter had any coverage at all, which is precisely why the Space utilization defect above survived. The context asserts that the utilization reported without expansion matches the expansion path exactly — the two are answered by different Resource Graph queries, so this guards the prefix-only query against divergence; that a Block's used address count equals its networks plus its external networks; and that expanded responses genuinely carry the expanded fields, since a Union response model degrades silently rather than erroring. Every assertion was confirmed to fail against the unfixed engine before the corresponding fix landed, so none of them can pass vacuously. All are read-only, leaving the suite's ordered state untouched
  • External network and per-network utilization coverage: The utilization tests asserted Block and Space arithmetic but nothing verified a network's own size and used, nor that external networks reported utilization at all, so both defects fixed in this release would have passed the suite. Four assertions were added: that an expanded network never reports more space used than its size, that an external network's used is the space assigned to its subnets, that an external subnet's used is its endpoint count, and that a Block counts an external network once rather than counting its subnets again. The first three were confirmed to fail against the code that carried each defect; the fourth is a guard that would catch an external network being double counted if the accounting were ever changed

Bug Fixes

[major]

DCMattyG and others added 30 commits August 24, 2026 10:13
Upated NPM packages to latest versions.
Pin every action reference to an immutable commit SHA with a trailing
version comment, and add a Dependabot configuration for the
github-actions ecosystem so the pins stay current. Pinning by SHA
prevents supply-chain attacks in which a mutable tag is retargeted to
malicious code, per the GitHub Actions security hardening guidance. The
7 day Dependabot cooldown leaves a window for compromised releases to be
reported before they are proposed as updates.

Also move every action to its latest GA release so the branch does not
ship already behind: checkout v7.0.1, setup-node v7.0.0 and setup-python
v7.0.0 (all ESM migrations, both majors already on the node24 runtime),
plus hadolint-action v3.5.0. The inputs in use are unchanged across the
major boundary, and the fork-PR checkout restriction added in checkout
v7 does not apply here as no workflow uses pull_request_target or
workflow_run.

hadolint-action v3.5.0 upgrades the linter from v2.14.0 to v2.15.1,
which adds DL3066 and flags the intentional `USER root` directives in
the RHEL images. Ignore DL3066 so the lint job stays green; this affects
linting only and leaves the built images unchanged.

Supersedes #386, which was authored against main and pinned to action
versions that predate the OIDC migration and action bumps on this
branch.

Co-authored-by: Dan Fiedler <151573964+danfiedler-msft@users.noreply.github.com>
Handle both aggregated ARG private_ips arrays and singular private_ip
records returned by the sovereign-cloud SDK path.
…ng them as 403

Every non-success response from Azure Resource Graph was caught as a bare
HttpResponseError and re-raised as "403 Access denied.", so throttling,
upstream outages and expired credentials were all reported as authorization
failures. Because ClientAuthenticationError subclasses HttpResponseError, the
"Token has expired." handlers in the three ARG wrappers were unreachable.

Failures are now classified at the single choke point in arg_query_helper:

  - 429 is retried with jittered backoff that honours Retry-After and
    x-ms-user-quota-resets-after, then surfaced as 429 with Retry-After when
    the wait would exceed the retry budget
  - 5xx and unexpected upstream statuses become 502 rather than being relayed
    and blamed on the caller
  - authentication failures become 401, or 500 when the IPAM service principal
    itself failed to authenticate
  - genuine Forbidden responses remain 403

Only the individual page request is retried, so paged results are never
re-appended. Remaining Resource Graph quota is logged per page, and the HTTP
exception handler now propagates response headers so Retry-After reaches the
caller.

Also converts percent-style placeholders in loguru calls, which silently
dropped their arguments.

BREAKING CHANGE: Azure Resource Graph throttling now returns 429, upstream
failures return 502, and authentication failures return 401 or 500. These
previously returned 403.
…okens

The Pester suite issued every request once and treated any failure as final, so a
single Azure Resource Graph throttle surfaced as an unrelated assertion failure and
aborted the whole run.

Requests now route through a shared pipeline that retries only definitive transient
responses (429, 500, 502, 503, 504), honouring the Retry-After header the engine
sends. Client errors the suite asserts deliberately (400/403/404/409/422) are never
retried, and transport failures are never replayed because a POST may already have
been applied server-side. Retries are logged as warnings so throttling that was
silently recovered from is still visible in CI output.

The access token is now cached and re-acquired on expiry rather than fetched once in
BeforeAll. The suite waits about twelve minutes on Resource Graph propagation by
design, so a slow run could previously outlive its token and fail with an
unexplained 401. A 401 is deliberately not retried, so genuine authentication
regressions still fail loudly.

Helper signatures and all test bodies are unchanged.
get_network, get_vnet and get_vhub served as both FastAPI routes and internal
helpers, so their Depends() defaults made a missing positional argument silently
truthy rather than an error. Thirteen internal callers passed only two arguments,
which put the admin flag into tenant_id and left admin as an unresolved Depends
object. Every one of those calls therefore queried across the whole tenant by
accident, and the Cosmos lookup in get_vnet hit a nonexistent partition, so
parent_space and parent_block always resolved to None.

The query bodies are now plain functions with required arguments (fetch_vnets,
fetch_vhubs, fetch_networks) and the routes are thin delegates, so an omitted
argument is a TypeError. The third parameter is named all_networks because the
question at an internal call site is whether the operation needs every network in
the tenant, not whether the caller is an admin.

Occupancy questions -- reservations, next-available searches, CIDR overlap
validation and utilization totals -- deliberately ask for every network:
answering them from the caller's scope would let IPAM allocate over networks the
caller cannot see. Resource enumeration stays scoped, so /available continues to
offer only networks the caller can actually associate.

Behaviour is unchanged apart from parent_space and parent_block now resolving
correctly and one redundant Cosmos query per call being removed.
Occupancy and utilization paths read only a network's ID and address
prefixes, but called fetch_networks, which expands every subnet, joins
peerings, resolves vHub route tables via ARM and enriches from Cosmos DB.
On large tenants that is the dominant source of Azure Resource Graph
quota pressure and a contributor to 429 throttling.

Add fetch_network_prefixes, backed by a single NET_BASIC query, and use
it at the thirteen call sites that need nothing further. Paths that
return whole network objects still use fetch_networks, selected on the
expand flag.

Nothing is cached. Both queries are read live on every request, so
allocation decisions remain based on the current state of Azure.
…orks

GET /api/spaces/{space}?utilization=true returned 500 for any space with
an external network in one of its blocks. External address space was
accumulated onto the space path parameter, which is a string, rather than
the space document:

  File "app/routers/space.py", line 744, in get_space
      space['used'] += IPNetwork(ext['cidr']).size
  TypeError: string indices must be integers, not 'str'

The line is correct in get_spaces, where `space` is the loop variable, and
was copied into get_space without renaming. The list endpoint and the
block-level endpoints were never affected, which is why this went unseen.
Requesting expand=true on a space or block containing a vWAN hub silently
returned unexpanded network references, with an HTTP 200 and no error.

VNetExpand requires a subnets field. Virtual networks carry one, virtual
hubs never did, so validation of the expanded shape failed and the Union
response model fell through to the plain reference shape. The effect was
not confined to the hub: a single hub in a block collapsed every network
in that block back to {id, active}.

Virtual hubs now carry an empty subnet list, which is what they actually
have, so the expanded shape validates. The block networks and available
endpoints were unaffected, as NetworkExpand does not require subnets.
GET /api/azure/vhub returned 500 as soon as any vWAN hub was associated
to a block:

  ResponseValidationError
    {'loc': ('response', 0, 'parent_block'),
     'msg': 'Input should be a valid string', 'input': ['BlockHub']}

A network can belong to blocks in more than one space, so fetch_vhubs has
always assigned a list of block names. The model declared a single string.

Correct the model to match the data rather than the reverse, since
collapsing to one name would discard a real association. parent_space is
unchanged; it resolves to a single value.

The defect was masked on /api/azure/network, which returns the same hub
data but declares no response model.
Neither utilization=true nor expand=true had any coverage, which is why
the space utilization defect and the silent virtual hub expansion
fallback both went unnoticed.

Add a read-only Utilization & Expansion context covering both parameters
across the spaces, space, blocks and block endpoints:

- utilization reported without expansion must match the expansion path
  exactly. The two are answered by different Resource Graph queries, so
  this guards the prefix-only query against divergence
- a block's used address count must equal its networks plus its externals
- expanded responses must actually carry the expanded fields, since a
  Union response model degrades silently rather than erroring

Every assertion was confirmed to fail against the unfixed engine before
the fixes landed, so none can pass vacuously. All are GETs, leaving the
suite's ordered, stateful sequence untouched.
Azure Resource Graph throttles a POST and states the wait as
x-ms-user-quota-resets-after rather than Retry-After. azure-core therefore
never retried it: POST sits outside the method allowlist, and the
Retry-After short circuit that would have overridden that never fires.

Teach the SDK's own retry policy both facts rather than hand-rolling a
retry loop beside it. It now waits exactly as long as Resource Graph asked
instead of backing off blindly, and the fixed attempt count and backoff
ceiling give way to the SDK's own settings.

Also guard result completeness. Resource Graph withholds the continuation
token when it truncates a result set, so the missing rows cannot be paged
and the response is silently short. That now raises rather than returning
partial data.
Azure returns resource IDs with inconsistent casing between APIs, and
Microsoft's naming rules direct callers to always compare names
case-insensitively. Azure IPAM stores whichever casing the client supplied
when associating a network, because association validates with a
case-insensitive lookup but persists the request body verbatim.

Fifteen comparisons matched a stored ID against an Azure-returned ID
exactly, so a casing difference made a managed network invisible:

  - utilization under-reported the address space in use
  - block CIDR and external network validation skipped the network
  - reconciliation marked the network inactive
  - the CIDR check reported it as belonging to no space or block
  - removing a network by ID rejected a valid ID as invalid

Every one failed in the same direction, reporting occupied space as free.

One case compared a Resource Graph ID against an ARM SDK ID, where the two
APIs genuinely disagree on casing. Reservation identifiers are generated by
Azure IPAM rather than Azure, so those stay case-sensitive.
When a virtual network connects to a vWAN hub, Azure peers it to a
generated transit network named HV_<hub>_<suffix>, which is substituted
back to the hub for display. The hub name was interpolated into a pattern
without escaping, and Azure permits periods in hub names, so a hub named
"my.hub" produced a pattern where the period matched any character and
"myXhub" matched as well.

The pattern was also case-sensitive against a Resource Graph identifier.

Match the transit network as a lowercased substring instead. That is all
the comparison ever needed, and it removes both the escaping hazard and
the case sensitivity.
…query

The available-networks endpoint used the detailed network query whenever
expansion was requested, then returned a NetworkExpand response whose six
fields are exactly what the prefix-only query already provides. Every
additional field the detailed query produced was discarded field-for-field,
at the cost of a second Resource Graph query, a subnet expansion of up to
1,024 rows per virtual network, two peering joins, an ARM call per virtual
hub and two Cosmos DB queries. The UI association picker always requests
expansion, so this was the most frequently exercised expanded path.
Scale Set discovery extracted the resource group, virtual network, subnet
and instance number from ARM resource IDs by matching the resourceGroups/,
providers/, virtualNetworks/, subnets/ and virtualMachines/ segments
literally. Azure does not guarantee the casing of those segments, and a
non-matching expression returned None, whose immediate .group(0) raised and
failed the whole request with a 500.

This is the SDK-based path, reached only when AZURE_ENV is not AZURE_PUBLIC,
so it affects sovereign and air-gapped clouds; commercial deployments answer
the same request from Resource Graph. The Reservation path already matched
case-insensitively; these six now do the same.
The helper splitting a reservation tag into IDs wrapped its work in a bare
except returning an empty list, conflating a network with no tag at all --
the normal case for nearly every network -- with a tag whose value could
not be read.

The Resource Graph query parses tag values as JSON so an absent tag resolves
to null, which also means a value resembling JSON arrives as a list, object
or number rather than text. That network was then treated as having no
reservation, so the reservation waited to be fulfilled indefinitely with
nothing logged.

The type check is now explicit, and reconciliation reports the network ID
and offending value once per run rather than once per comparison. Values
that cannot be read are still skipped rather than guessed at.
Resource Graph quota is shared by every caller using the same credentials
rather than allocated per user, and replenishes roughly every five seconds.
The SDK's default of three status retries could therefore be spent in about
fifteen seconds while a burst of requests was still draining, surfacing a
429 to the caller.

Raise the status allowance to six, bounding a throttled request at roughly
thirty seconds of waiting instead. It is not raised further deliberately:
past ten the total retry budget becomes the real limit, and a longer wait
encourages a page reload, which adds load to the same quota.
The arithmetic populating size and used on a block, its networks and their
subnets was duplicated across the four endpoints reporting utilization --
about thirty lines repeated four times, differing only in local variable
names and whether the running space totals were accumulated alongside. That
duplication is why the space utilization defect existed in exactly one copy.

Collapse the four into add_block_utilization, summing the space totals from
each block's own figures, which is the same value since the space totals
were only ever the sum of their blocks.

Behaviour is unchanged, including two quirks in the per-network figures that
are addressed separately.
A block counted each external network's full address range toward its own
utilization but reported nothing about how that range was allocated, even
though external subnets and endpoints were added to the model later.

External networks and their subnets now carry size and used alongside the
Azure networks beside them. An external network's used is the space assigned
to its subnets, mirroring a virtual network; an external subnet's used is
its endpoint count, mirroring an Azure subnet's consumed addresses, without
the five addresses Azure reserves since an external network reserves none.

The block's own used is unchanged and still counts each external range once,
as its subnets sit inside that range.
For an expanded network, size counted only the prefixes falling inside the
block while used counted the subnets of every prefix the network owns,
including those outside it. The two answered different questions and were
presented as a ratio, so a network straddling a block boundary could report
more space used than it has.

used was also reset inside the loop over in-block prefixes, so a network with
no prefixes inside the block never had it reset and reported the whole
network's subnet total against a size of zero.

Initialise used once and count only subnets inside the block. Block and space
totals are unaffected, as they never consulted subnets.
The utilization tests asserted block and space arithmetic but nothing
verified a network's own size and used, nor that external networks reported
utilization at all. Both defects fixed in this branch would have passed CI.

Assert that an expanded network never reports more used than its size, that
an external network's used is the space assigned to its subnets, that an
external subnet's used is its endpoint count, and that a block counts an
external network once rather than its subnets again.
…ions

The suite used capitalized Function and Param, making it the only PowerShell
file in the repo not following the lowercase keyword style that
PSScriptAnalyzerSettings.psd1 describes and every other script applies.

It also carried the repo's only analyzer Error, from normalizing an Azure
access token to a SecureString. Suppress it inline with the same
justification already used in deploy.ps1, migrate.ps1 and update.ps1.

The suite now reports no analyzer findings at any severity.
version.ps1 declared ValueFromPipelineByPropertyName on all four parameters
without a process block, so piping several objects would have silently
processed only the last. Nothing pipes to the script, and no comparable
script in the repository declares pipeline binding -- update.ps1 has 81
parameters and migrate.ps1 41, both with none. deploy.ps1 is the only script
that does, and it implements begin and process blocks to match. Removing the
unused bindings leaves the parameter sets and every CI invocation unchanged.

version.ps1 also called get-date -format in lowercase, where every other
script uses Get-Date -Format, and carried a trailing space in its header
banner.

Separately, all five scripts declared $logPath and then referenced $logpath
on the very next line when creating the log directory. PowerShell variable
names are case-insensitive so the behaviour was correct, but the mismatch
had been copied into every script. Corrected in deploy.ps1, migrate.ps1,
update.ps1 and version.ps1; build.ps1 is handled separately.

No behavioural change.
The lint job covered the UI, engine, Bicep templates and Dockerfiles, but no
PowerShell file was ever analyzed, which is how the Pester suite drifted from
the repo conventions and accumulated the only analyzer finding in the tree.

The step is scoped to the four rules PSScriptAnalyzerSettings.psd1 configures.
Running PSUseCorrectCasing alongside the full default rule set intermittently
crashes the analyzer through a thread-safety defect in its command cache,
which would fail the step on a clean tree roughly a third of the time. That is
upstream PowerShell/PSScriptAnalyzer#1708, open since 2021 and present in both
1.24.0 and 1.25.0, so pinning an older version does not avoid it.

Verified in a clean Ubuntu container: the step passes on the current tree,
fails on reintroduced casing drift, and fails rather than passing vacuously
when no scripts are found.
…package

Add a wheel-tag layer to the native module gate in the build script. After pip
install, each *.dist-info/WHEEL is parsed and a distribution is rejected unless
one of its tags is `any` or a manylinux at or below the glibc ceiling derived
from PIP_PLATFORM. This catches locally compiled wheels (bare linux_x86_64) and
over-new ones (manylinux_2_28/2_34) by name, before the ELF symbol scan has to
find them. The gate is now three layers: ABI naming, wheel tag, GLIBC_ symbols.
The scan additionally records the highest glibc symbol required by any bundled
module and reports it on every build rather than only on failure.

Write build.json to the archive root, recording the build timestamp, app and
python versions, pip platform, glibc ceiling versus observed, native module
count and the resolved wheel tags per package. Air-gapped clouds cannot share
build logs, so the artifact has to be able to identify itself.

Resolve the engine's Python version from the archive being deployed rather than
the local checkout, and let configuration drift retarget LinuxFxVersion to
match. An archive built for one Python version could previously be deployed
onto a site configured for another with nothing detecting the mismatch; the
bundled wheels are ABI-specific, so every native module silently becomes
unimportable and the failure surfaces at startup rather than at deploy time.

Gate WEBSITES_INCLUDE_CLOUD_CERTS on the cloud rather than the deployment
shape, in the deployment modules, the staging slot modules and the update
script's target settings map. The trust store is a property of the cloud; the
run-from-package shape is not. The setting consequently now also reaches
container deployments in sovereign clouds, which the previous nesting excluded.

Add a hidden -RunFromPackage switch to the deployment and update scripts,
threaded through to a forceRunFromPackage template parameter, so the
internet-restricted deployment path can be exercised on clouds that would
otherwise build on the server. The switch is marked DontShow and is
intentionally undocumented.

Document deployments that report success without changing the running
application, covering how to confirm which package is actually mounted,
package retention, and the expected virtual environment warning.

Refs: GLIBC_2.33 import errors reported from IL6
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

4 participants