Repository navigation
Conversation
…n_batch knn_search (open3d.ml.torch.ops) documents that it returns fewer than k neighbors per query whenever there are fewer than k candidate points -- its neighbors_index output is then ragged (not a multiple of k in total length), and its neighbors_row_splits marks each query's sub-range. knn_batch ignored this and called ans.neighbors_index.reshape(-1, k) directly, which raises "shape '[-1, 16]' is invalid for input of size N" (or silently misaligns rows on a lucky divisor) whenever any query in a batch has fewer than k neighbors -- e.g. a training config with widely spaced points, or a point cloud smaller than k. FixedRadiusSearch's ragged output hits the same shape elsewhere in this repo (kpconv.py's batch_neighbors) and is already densified with open3d.ml.torch.ops.ragged_to_dense using neighbors_row_splits; apply the same op to knn_search's output here. knn_batch's result is used by queryandgroup to index directly into points/feat, so padding cannot use ragged_to_dense's out-of-range default_value as-is (points.shape[0] would raise IndexError instead). After densifying, replace padded slots with that query's own first neighbor (a valid, already in-range index), and fall back to point 0 only for the pathological case of a query with zero neighbors at all. Distances are padded the same way, defaulting to 0.0 in that same all-empty edge case so no inf reaches downstream softmax/attention weights. Fixes isl-org#683 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Training PointTransformer on a point cloud (or a batch region) with fewer candidate points than
kcrashes:Root cause
open3d.ml.torch.ops.knn_search's own docstring states it "supports returning less than k neighbors if there are less than k points" and that its output format is "compatible with the radius_search and fixed_radius_search ops" — i.e.neighbors_indexis ragged: its total length is not guaranteed to be a multiple ofk, andneighbors_row_splitsmarks each query's actual sub-range.knn_batchinml3d/torch/models/point_transformer.pyignored this and calledans.neighbors_index.reshape(-1, k)directly. When every query has exactlykneighbors this happens to work; the moment any query has fewer, the reshape either raises (this issue) or, if the ragged total happens to still divide evenly byk, silently misaligns which neighbors belong to which query.Fix
ml3d/torch/models/kpconv.py'sbatch_neighborsalready solves the identical ragged-output problem forFixedRadiusSearchusingopen3d.ml.torch.ops.ragged_to_dense(values, row_splits, out_col_size, default_value). I applied the same op here, usingneighbors_row_splitsto densify to(num_queries, k).One difference from
kpconv.py's usage:knn_batch's result is used byqueryandgroupto index directly intopoints/feat(points[idx.view(-1).long(), :]), so I can't reuseragged_to_dense's out-of-rangedefault_value(points.shape[0]) as the final padding — that would just trade the reshape crash for anIndexError. After densifying with that sentinel, padded slots are replaced with that query's own first (real) neighbor, which keeps every returned index valid and contributes a harmless, already-nearby duplicate to the fixed-size attention window instead of an arbitrary or out-of-range one. The one further edge case — a query with zero neighbors at all — falls back to point index0. Distances get the same treatment, defaulting to0.0in that same all-empty case so noinfreaches the downstream attention/softmax.Testing
test_pointtransformer_knn_batch_ragged_neighborsintests/test_models_torch.py, using a 5-point cloud withk=16(guaranteeing every query gets fewer thankneighbors, matching the issue's actual trigger condition) and asserting the returnedidx/disthave the correct(n_points, k)shape with every index in[0, n_points).ragged_to_dense(this environment has no compiled Open3D ML ops), covering: a full row (unaffected), a short row (padded with its own first neighbor), and a fully empty row (falls back to index 0) — all three produced correctly shaped, in-range output.knn_search's own docstring (extracted from theopen3d==0.20.0wheel) that its ragged output sharesneighbors_row_splitswithFixedRadiusSearch's result, validating that thekpconv.pydensification pattern applies unchanged.python3 -m py_compilepasses on both changed files.test_pointtransformer_torchhere — this environment doesn't have the compiledopen3d/open3d.ml.torchpackage with GPU ops. Please runpytest tests/test_models_torch.py -k pointtransformerin CI/a configured dev environment to confirm.Fixes #683
🤖 Generated with Claude Code