Conversation
…the same StorageType. Co-authored-by: https://github.com/xichen01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Co-authored-by: https://github.com/xichen01
What changes were proposed in this pull request?
On a cluster with mixed storage media, the
ContainerBalancercan currently move a replica from one storage tier to another. A replica placed on SSD can be relocated to DISK or ARCHIVE purely because the destination datanode has more free space. This silently defeats the storage policy the key was written with, and the balancer has no way to notice: it measures utilization across all of a datanode's volumes, so a node that is full on SSD but empty on ARCHIVE looks half-used.This PR makes the balancer tier-aware end to end, so a move keeps a replica on the storage type it is already on.
Protocol changes
Two optional fields:
An SCM without storage-type support sends neither, which leaves the target free to choose a volume, so mixed-version clusters behave as they do today.
What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16653
How was this patch tested?
New and extended unit tests, each verified to fail when its production change is reverted:
TestReplicateContainerCommand— new, 9 tests covering the proto round-trip with and without the storage typeTestDatanodeUsageInfo— per-tier utilization,getStorageTypes()reporting only tiers with capacity, and an absent tier reporting no usage rather than failingTestSCMContainerPlacementCapacity— ranking on the requested tier, and unchanged overall-usage ranking when no tier is given. Both count outcomes over 2000 draws, since chooseNode picks two random indices and returns without comparing when they matchTestContainerBalancerSelectionCriteria— exclusion of replicas not on the tier being balancedTestMoveManager,TestContainerImporter,TestSendContainerOutputStream— extended for the nullable storage type, including@EnumSource(StorageType.class)coverage and a null case