Repository navigation
Semantic relations with a judged p; doc sections linked to their code by format and position - #2574
Merged
Merged
Conversation
The TF-IDF signal keyed each term by the token's POSITION in the function, so any two functions of similar length shared their first terms whatever their words: 0.20 of every SEMANTICALLY_RELATED score measured length, not vocabulary. Terms are now keyed by corpus token (cbm_sem_tfidf_terms: one term per distinct word, weight term frequency x idf, ascending token index, which the sparse cosine merges on). Measured on a blind judged sample: 222 function pairs from six held-out repositories, stratified by score band, read in three rounds. The share of admitted SEMANTICALLY_RELATED edges a reader judges related rises from 0.713 to 0.742, with about 19% fewer edges (fastapi 507 -> 426, django 64 -> 66). The same sample shows that no score threshold reaches 0.90 precision, so the documented "~95% at 0.75" did not hold before this change either. The harness that measured it ships with the fix, off by default: CBM_SEM_PAIR_SIGNALS=<path> writes one line per scored pair (both qualified names, the score, every signal value, whether the pair was admitted, both locations); CBM_SEM_PAIR_SIGNALS_FLOOR also records pairs below the threshold, never admitting them. cbm_sem_combined_score is now cbm_sem_signal_values + cbm_sem_combine (same arithmetic, same order). CBM_SEMANTIC_INDEX_VERSION 3 -> 4: indexes built before score pairs differently. Tests: sem_tfidf_terms_compare_vocabulary_not_positions, sem_combine_weighs_signals, pipeline_semantic_pair_signals_never_change_the_graph; each is RED with its behavior reverted. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
When a function has more candidate partners than its budget of ten SEMANTICALLY_RELATED edges, admission took the pairs in discovery order: the first ten won, whatever their scores. Admission is now best-first (cbm_sem_admit_best_first: descending score, ties in the canonical order, still a pure function of the sorted inputs). On a blind judged sample from held-out repositories, the pairs first-come admitted over budget were 0 of 17 related; the pairs best-first admits instead, 7 of 20. Where no budget binds nothing changes (django 66 edges identical; fastapi 426 -> 430 edges, 222 changed). Every SEMANTICALLY_RELATED edge now carries "p": the probability that a reader judges the pair related, by score band, from 120 blind-judged admitted pairs (score < 0.778: 0.56, < 0.806: 0.64, < 0.852: 0.73, else 0.86). p is computed from the score as stored (three decimals), so one shown score always has one p. Scores and the 0.75 threshold are unchanged; a scoring change (TF-IDF alone, or weights fitted on the first sample) was judged on the same sample and did not beat the shipped score by the pre-registered margin. Tests: sem_admit_best_first_keeps_the_best, sem_calibrated_p_bands, pipeline_semantic_edges_carry_p; each is RED with its behavior reverted. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The semantic pass read every string property of a function (body text, signature, return type, decorators, docstring) with a reader that stopped at the first '"' and kept escapes raw: a value holding an escaped quote lost everything after it, and "\nreturn" became the token "nreturn". cbm_sem_json_str honours the escapes: \" \\ \/ keep their character, \n \t \r \b \f and \uXXXX become a separator. The tokens change, so admissions change: on 13 repositories 1,340 -> 1,387 admitted SEMANTICALLY_RELATED pairs, 614 of them unchanged. Of the dropped pairs that earlier samples had judged, 52 of 89 were related (0.58); of the kept ones, 73 of 92 (0.79). A pre-registered blind sample of 100 newly admitted pairs from 11 held-out repositories (three reading rounds) judged 65 related (0.65, Wilson 0.55-0.74), so the admitted set is 0.72 precise against 0.69 before. The same index run twice gives identical pairs and scores. "p" is refit on the new admitted population (bands at the quartiles of the admitted scores, pooled where a higher band judged lower; added and kept pairs weighted by their share): score < 0.773: 0.63, else 0.76 (was 0.56 / 0.64 / 0.73 / 0.86 in four bands). Tests: sem_json_str_honours_escapes (RED with escapes ignored), sem_calibrated_p_bands. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A documentation section that explains a function rarely names it, so the deterministic doc links cannot connect them. The semantic pass now ranks, for every Markdown section that can be about code, the functions whose TF-IDF terms best match the section's heading and body, over the functions' own corpus (function-pair scores are untouched): - gate: no candidates for release-number or project-meta headings (changelog, licence, contributing, links, contents, install, ...), bodies under 20 words, or bodies that are mostly links; - targets: functions and methods, never tests (by path or name); - stored: a section's best five with tfidf >= 0.20, each with p, the probability that the section is about the function, from 274 blind-judged candidates in five held-out repositories (three reading rounds): 0.36 below 0.30, 0.42 below 0.40, else 0.55. Below 0.20 the judged share was 0.06, so nothing is stored there. These are candidates for an agent to verify, not graph edges: at best about half are right, far below the bar for a link. They live in a side table, doc_link_candidates (section and function qualified names, rank, score, p, evidence), created when a generation is published, so older indexes read as "no candidates". Full and legacy-partial reindexes replace a project's rows; the delta route keeps the previous rows (its proxy buffer has no section text or function tokens to recompute them), and a row whose node is gone is skipped when read. get_code_snippet shows them: on a Section, "possibly_about" (its candidate functions); on a function or method, "possibly_described_in" (the sections), at most five, highest p then score first, with a candidates_note saying what they are. Measured on fastapi (6,144 gated sections): 24,293 rows, identical to the generator's dump; index time within noise; the database grows by the table (37 MB there, with a 148-character project prefix on every qualified name). Tests: semantic:sem_doc_calibrated_p_bands, store_arch:doc_candidates_replace_and_get (RED when a replace keeps old rows), pipeline:pipeline_doc_candidates_published_and_kept_by_delta (RED with test functions admitted, the heading gate off, the publish write skipped, or the delta route replacing rows), mcp:snippet_doc_candidates_both_directions (RED with stale rows shown). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The candidate rows and their strings were allocated with calloc/strdup in the semantic pass and in cbm_store_doc_candidates_get, and released with free. They now go through the memory core in CBM_MEM_CLASS_STORE (cbm_calloc, cbm_mem_strdup, cbm_free) on both sides, so the rows the pass hands to the pipeline and the rows a read returns are accounted like every other store buffer. make lint-memory-core: no file grew. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
pipeline_doc_candidates_published_and_kept_by_delta appended its delta-route edit to "user_config.go" while the fixture writes "User_config.go". On a case-insensitive filesystem (macOS, Windows) that is the same file; on Linux it is a new file, closure repair declines new files (correctly), the run falls back to a full rebuild, the route assertion fails and the early return leaks the pipeline (LeakSanitizer on the arm64 GCC leg). The edit now names the fixture file exactly. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…d home folders Doc -> code candidates now carry a p judged for their kind, their document format and where the document sits. Format (sample 6, 234 blind-judged candidates): reStructuredText, AsciiDoc and PDF sections each get their own curve (reST stored from tfidf 0.30, AsciiDoc from 0.20, PDF not measured and not stored). A candidate that repeats an exact MENTIONS link of the same section is not stored. Position (sample 7, 490 blind-judged candidates from four monorepos): every document gets a home folder -- its folder; a docs folder documents its parent; a folder without code hands over to the nearest one with code; the repository root means the whole project. Markdown candidates outside the document's home were right 6 of 90 and are never stored; inside it they are stored from tfidf 0.30 (0.33 / 0.53); for documents of the whole project samples 4 and 7 pooled give 0.22 / 0.51 from 0.30. New kinds, all in doc_link_candidates (evidence gains "kind" and "position"; no new edge type, no schema change): - local: up to two home functions ranked below the top five (0.32); - file: a whole file, its vector the sum of its functions' terms and its definitions' names, config keys included, cut to 64 terms (0.30, from tfidf 0.40); - folder: a README or index document -> its home folder (0.87). get_code_snippet on a section lists every kind with its label; get_file_outline shows the sections about the file or a folder holding it (cbm_store_doc_candidates_for_path). cbm_sem_p_2dp rounded 0.325 and 0.525 down (their float values sit 1.2e-6 below the half); its epsilon is 1e-4 on the x100 scale. Replayed on the sample repositories, the build stores exactly the rows the judged curves select (950 rows, 0 differences). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…cabulary Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Brings the doc_links block test fix. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
DavidHLP
pushed a commit
to DavidHLP/codebase-memory-mcp
that referenced
this pull request
Oct 11, 2026
The doc-section candidate pass (DeusData#2574) made a Linux kernel index 2.8x slower: 207 s before DeusData#2574, 580 s after, on the same machine. Every candidate of every section got a dot product that merged the section's up to 1024 terms with the function's, and then every candidate was sorted for a top 10. - The section's vector is held densely per worker, indexed by term and stamped per section, so a candidate's dot product walks only the function's (or file's) own terms. The same products are added in the same ascending term order, so every score is the same float. - Only the candidates the results read are sorted: the top SEM_DOC_TOP_K (and the local ones down to CBM_SEM_DOC_MIN_SCORE), the top CBM_SEM_DOC_FILE_K files. cmp_doc_cand is a total order, so the sorted prefix is the full sort's prefix. Kernel (perf-bench/linux, M5 Pro, same plain build): 579.7 s -> 391.5 s, nodes and edges unchanged. Exactness, measured directly: a build that also ran the old merge and the old full sort found 0 score, 0 shared-term and 0 order differences over 5,003 django sections. Candidate generation itself (a section's key terms, df <= funcs/20) is unchanged here: narrowing it would change the stored rows. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Semantic relations you can trust a number on, and doc sections that point at their code
Seven commits on
fix/semantic-tfidf-vocabulary(plus merges that bringmainup to date), building on #2550 (doc-comment links) and #2551 (documents as link sources), which have both landed; each change was measured on real repositories and judged blind before it was kept.1. TF-IDF compares vocabulary, not token positions (
fix(semantic))The TF-IDF signal compared the i-th token of one function with the i-th token of the other. It now compares the shared vocabulary (
cbm_sem_tfidf_terms). On a pre-registered blind sample from six held-out repositories, admitted-pair precision went from 0.713 (main) to 0.742.CBM_SEMANTIC_INDEX_VERSIONgoes 6 -> 7 (after #2550 and #2551), so existing indexes rebuild once.2. Best-first admission; every edge carries
p(feat(semantic))When a function has more candidate partners than its budget of ten, admission now keeps the best-scoring pairs instead of the first ones found. Judged: the pairs first-come admitted over budget were 0 of 17 related; best-first admits 7 of 20 instead. Every
SEMANTICALLY_RELATEDedge carriesp, the judged probability that the pair is related.3. JSON strings read with their escapes;
precalibrated (fix(semantic))The pass read function properties with a reader that stopped at the first
"and kept\nraw ("\nreturn" -> token "nreturn"). Fixing it changes about half of the admitted pairs (13 repos: 1,340 -> 1,387, 614 unchanged). A pre-registered blind sample of 100 newly admitted pairs judged 0.65 related, so the admitted set goes from 0.69 to 0.72 precision.pis refit on the new population: score < 0.773 -> 0.63, else 0.76.4. Doc section -> code candidates (
feat(semantic)) + memory-core allocation (fix(semantic))A section that explains a function rarely names it. For every Markdown section that can be about code (gated: no changelog/licence/install/... headings, at least 20 words, not mostly links), the pass ranks the functions whose TF-IDF terms best match it. Tests are never targets. A section's best five with tfidf >= 0.20 are stored with
pfrom 274 blind-judged candidates (0.36 / 0.42 / 0.55 by band; sections 5 and 6 replace this curve by format and by position).These are candidates for an agent to verify, not edges: at best about half are right. They live in a side table,
doc_link_candidates, keyed by qualified name and created at publish (older indexes read as "none"). The delta route keeps the previous rows, and a stale row is skipped when read.get_code_snippetshows them aspossibly_abouton a Section andpossibly_described_inon a function or method, with acandidates_note.fastapi: 24,293 rows, identical to the generator's dump; index time within noise; the DB grows by the table (37 MB there, with a long project prefix on every QN).
5. One judged curve per document format; no candidate repeats an exact link (
feat(semantic))With the documents layer (#2551), reStructuredText, AsciiDoc and PDF sections are candidates too (on django, 20,770 reST-in-
.txtcandidates against 19 Markdown ones), so the Markdown curve alone would mislabel almost everything. A second pre-registered blind sample (234 candidates, three Sphinx projects and one AsciiDoc project, three reading rounds) gives each format its own curve:A candidate the section already links exactly (a MENTIONS edge) is not stored again (4% of django's reST candidates). Replayed on the sample's repositories, the code stores exactly the rows the pre-registered rule selects (618 rows, 0 differences).
6. Where a doc sits: position, local extras, whole files, home folders (
feat(semantic))Every doc gets a home folder: its own folder, except that a
docs/folder documents its parent, and a folder without code hands over to the nearest one with code. The repository root means "this doc is about the whole project". A third pre-registered blind sample (490 candidates from four monorepos with per-component docs, three reading rounds) measured what that position is worth, plus three new kinds of candidate:Position is the strongest signal measured so far: code outside a component doc's own folder was right 6 times in 90. Candidates now point at functions, whole files or folders, all in the same
doc_link_candidatestable withkindandpositionin the evidence (no new edge type, no schema change).get_code_snippeton a section lists every kind with its label;get_file_outlineshows the sections about the file or a folder holding it. Indexing time is unchanged on the five repositories measured (gritql 39.1 -> 39.5 s). Replay of the shipped curves against the frozen build's candidates: 950 rows, 0 differences. Fixed on the way:cbm_sem_p_2dprounded 0.325 and 0.525 down (float representation below the half; epsilon 1e-6 -> 1e-4).Tests (each RED with its behaviour reverted)
semantic: sem_admit_best_first_keeps_the_best, sem_calibrated_p_bands, sem_doc_calibrated_p_bands, sem_doc_format_by_extension, sem_p_two_decimals_half_up, sem_json_str_honours_escapes.store_arch: doc_candidates_replace_and_get, doc_candidates_for_path_file_and_folders.pipeline: pipeline_semantic_edges_carry_p, pipeline_doc_candidates_published_and_kept_by_delta, pipeline_doc_candidates_skip_exact_links_and_use_format_curves, pipeline_doc_candidates_home_position_and_kinds.mcp: snippet_doc_candidates_both_directions, tool_get_file_outline_shows_doc_candidates_of_file_and_folders.Local verification
On the tree stacked on #2551 and #2550, with
main72a2c0b merged in:make lint-ciclean.run.sh test): 8,965 passed, 0 failed, 10 skipped (173 suites).On this tip (
maincb8b641 merged in): it differs from that tree only by #2575's SHA-256 change and #2551's reST lint fix (cb24c41), and neither touches the semantic pass. macOS build plus thesemantic,doc_links_rstanddoc_mentionssuites: 123 passed. CI's cppcheck 2.20.0 is clean on all 35 production C files this branch changes. Oneclitest (cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothing, install code this PR does not touch) failed once in a multi-suite run and passed on both reruns of the suite; it is recorded as a sighting, not counted as fixed. The hosted 3-OS CI runs on this tip.Notes
CBM_SEMANTIC_INDEX_VERSION7 (fix(extract): C typedefs become nodes; macros and enumerators get their own QNs #2543 = 4, feat(doc-links): C# doc-comment references become MENTIONS edges #2550 = 5, feat(doc-links): Markdown, ADR, reST, AsciiDoc and PDF sections link to the code they name #2551 = 6 have landed): an index built before rebuilds once on upgrade.docs/docs/...takes the outerdocsfolder as its home (the rule cuts at the deepest docs folder); cutting at the outermost one is a follow-up.