fix: a retrain dropped rows from probed vector searches - #96
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: pgEdge/coldfront/.coderabbit.yaml Review profile: CHILL Plan: Essentials Run ID: ⛔ Files ignored due to path filters (3)
📒 Files selected for processing (10)
💤 Files with no reviewable changes (1)
Included review availability: This review used your included allowance. 4 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 5 reviews per hour. 📝 WalkthroughWalkthroughVector training now handles centroid initialization, Lloyd iterations, and cold-row assignment. Retraining at the configured cluster count reuses live centroid identities; changed counts use k-means++ seeding. The shared nearest-centroid expression supports training assignment and scoring. Documentation and journey tests reflect these changes. The ChangesVector Training
Priority: ➖ Normal Estimated code review effort: 4 (Complex) | ~40 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The documented retraining workflow has no established blocker to merging after normal checks. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Up to standards ✅🟢 Issues
|
Clustering assigns every cold row to the cluster of its nearest centroid, and a probed search reads only the clusters nearest the query. A retrain computed new centroids and pointed searches at them but left every row assigned under the previous centroids, so a row's cluster number now named a different centroid, or none, and probed searches silently left rows out: a table clustered into two groups and retrained into one returned three of its eight rows to a search that reads every cluster.
vector_trainnow ends by assigning every cold row to its nearest centroid of the set it just wrote, in the same transaction and under the table's claim, so the centroids and the assignments change together; only a row whose cluster changed is rewritten, and a retrain at the samenliststarts from the current centroids so each keeps its identity and most rows stay put.vector_assignis removed, since training covers the rows that predate it. The claim is taken once per transaction, before the sample, and released at commit, as the compactor holds it, so the modelled protocol is unchanged.