Skip to content

Semantic relations with a judged p; doc sections linked to their code by format and position - #2574

Merged
DeusData merged 12 commits into
mainfrom
fix/semantic-tfidf-vocabulary
Oct 10, 2026
Merged

DeusData merged 12 commits into
mainfrom
fix/semantic-tfidf-vocabulary

Conversation

@DeusData

@DeusData DeusData commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

Semantic relations you can trust a number on, and doc sections that point at their code

Seven commits on fix/semantic-tfidf-vocabulary (plus merges that bring main up to date), building on #2550 (doc-comment links) and #2551 (documents as link sources), which have both landed; each change was measured on real repositories and judged blind before it was kept.

1. TF-IDF compares vocabulary, not token positions (fix(semantic))

The TF-IDF signal compared the i-th token of one function with the i-th token of the other. It now compares the shared vocabulary (cbm_sem_tfidf_terms). On a pre-registered blind sample from six held-out repositories, admitted-pair precision went from 0.713 (main) to 0.742. CBM_SEMANTIC_INDEX_VERSION goes 6 -> 7 (after #2550 and #2551), so existing indexes rebuild once.

2. Best-first admission; every edge carries p (feat(semantic))

When a function has more candidate partners than its budget of ten, admission now keeps the best-scoring pairs instead of the first ones found. Judged: the pairs first-come admitted over budget were 0 of 17 related; best-first admits 7 of 20 instead. Every SEMANTICALLY_RELATED edge carries p, the judged probability that the pair is related.

3. JSON strings read with their escapes; p recalibrated (fix(semantic))

The pass read function properties with a reader that stopped at the first " and kept \n raw ("\nreturn" -> token "nreturn"). Fixing it changes about half of the admitted pairs (13 repos: 1,340 -> 1,387, 614 unchanged). A pre-registered blind sample of 100 newly admitted pairs judged 0.65 related, so the admitted set goes from 0.69 to 0.72 precision. p is refit on the new population: score < 0.773 -> 0.63, else 0.76.

4. Doc section -> code candidates (feat(semantic)) + memory-core allocation (fix(semantic))

A section that explains a function rarely names it. For every Markdown section that can be about code (gated: no changelog/licence/install/... headings, at least 20 words, not mostly links), the pass ranks the functions whose TF-IDF terms best match it. Tests are never targets. A section's best five with tfidf >= 0.20 are stored with p from 274 blind-judged candidates (0.36 / 0.42 / 0.55 by band; sections 5 and 6 replace this curve by format and by position).

These are candidates for an agent to verify, not edges: at best about half are right. They live in a side table, doc_link_candidates, keyed by qualified name and created at publish (older indexes read as "none"). The delta route keeps the previous rows, and a stale row is skipped when read. get_code_snippet shows them as possibly_about on a Section and possibly_described_in on a function or method, with a candidates_note.

fastapi: 24,293 rows, identical to the generator's dump; index time within noise; the DB grows by the table (37 MB there, with a long project prefix on every QN).

5. One judged curve per document format; no candidate repeats an exact link (feat(semantic))

With the documents layer (#2551), reStructuredText, AsciiDoc and PDF sections are candidates too (on django, 20,770 reST-in-.txt candidates against 19 Markdown ones), so the Markdown curve alone would mislabel almost everything. A second pre-registered blind sample (234 candidates, three Sphinx projects and one AsciiDoc project, three reading rounds) gives each format its own curve:

format judged p by tfidf band 0.20 / 0.30 / 0.40+ stored
Markdown see section 6 by position
reST 120 (0.125) / 0.35 / 0.63 from 0.30
AsciiDoc 114 (one repository) 0.23 / 0.25 / 0.50 from 0.20
PDF not measured - nothing

A candidate the section already links exactly (a MENTIONS edge) is not stored again (4% of django's reST candidates). Replayed on the sample's repositories, the code stores exactly the rows the pre-registered rule selects (618 rows, 0 differences).

6. Where a doc sits: position, local extras, whole files, home folders (feat(semantic))

Every doc gets a home folder: its own folder, except that a docs/ folder documents its parent, and a folder without code hands over to the nearest one with code. The repository root means "this doc is about the whole project". A third pre-registered blind sample (490 candidates from four monorepos with per-component docs, three reading rounds) measured what that position is worth, plus three new kinds of candidate:

Markdown candidate judged p by tfidf band 0.20 / 0.30 / 0.40+ stored
function inside the doc's home 120 (0.18) / 0.33 / 0.53 from 0.30
function outside the doc's home 90 (0.00 / 0.07 / 0.13) never
function, doc of the whole project 158 (samples 4 and 7 pooled) (0.17) / 0.22 / 0.51 from 0.30
home function ranked below the top 5 (at most 2) 60 0.32 at every band from 0.20
whole file (functions + definition names, config keys included) 90 (0.12 / 0.12) / 0.30 from 0.40
a README/index doc -> its home folder 30 0.87 yes
any other doc -> its home folder 30 (0.17) no

Position is the strongest signal measured so far: code outside a component doc's own folder was right 6 times in 90. Candidates now point at functions, whole files or folders, all in the same doc_link_candidates table with kind and position in the evidence (no new edge type, no schema change). get_code_snippet on a section lists every kind with its label; get_file_outline shows the sections about the file or a folder holding it. Indexing time is unchanged on the five repositories measured (gritql 39.1 -> 39.5 s). Replay of the shipped curves against the frozen build's candidates: 950 rows, 0 differences. Fixed on the way: cbm_sem_p_2dp rounded 0.325 and 0.525 down (float representation below the half; epsilon 1e-6 -> 1e-4).

Tests (each RED with its behaviour reverted)

semantic: sem_admit_best_first_keeps_the_best, sem_calibrated_p_bands, sem_doc_calibrated_p_bands, sem_doc_format_by_extension, sem_p_two_decimals_half_up, sem_json_str_honours_escapes. store_arch: doc_candidates_replace_and_get, doc_candidates_for_path_file_and_folders. pipeline: pipeline_semantic_edges_carry_p, pipeline_doc_candidates_published_and_kept_by_delta, pipeline_doc_candidates_skip_exact_links_and_use_format_curves, pipeline_doc_candidates_home_position_and_kinds. mcp: snippet_doc_candidates_both_directions, tool_get_file_outline_shows_doc_candidates_of_file_and_folders.

Local verification

On the tree stacked on #2551 and #2550, with main 72a2c0b merged in:

  • macOS: 9,127 passed, 0 failed, 11 skipped (177 suites); make lint-ci clean.
  • Linux (arm64 GCC container, run.sh test): 8,965 passed, 0 failed, 10 skipped (173 suites).
  • Windows (arm64 VM, all 177 suites): 8,955 passed, 0 failed, 90 skipped.
  • Replay of the judged curves on the sample-7 repositories: 950 rows, 0 differences.

On this tip (main cb8b641 merged in): it differs from that tree only by #2575's SHA-256 change and #2551's reST lint fix (cb24c41), and neither touches the semantic pass. macOS build plus the semantic, doc_links_rst and doc_mentions suites: 123 passed. CI's cppcheck 2.20.0 is clean on all 35 production C files this branch changes. One cli test (cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothing, install code this PR does not touch) failed once in a multi-suite run and passed on both reruns of the suite; it is recorded as a sighting, not counted as fixed. The hosted 3-OS CI runs on this tip.

Notes

The TF-IDF signal keyed each term by the token's POSITION in the
function, so any two functions of similar length shared their first
terms whatever their words: 0.20 of every SEMANTICALLY_RELATED score
measured length, not vocabulary. Terms are now keyed by corpus token
(cbm_sem_tfidf_terms: one term per distinct word, weight term
frequency x idf, ascending token index, which the sparse cosine
merges on).

Measured on a blind judged sample: 222 function pairs from six
held-out repositories, stratified by score band, read in three rounds.
The share of admitted SEMANTICALLY_RELATED edges a reader judges
related rises from 0.713 to 0.742, with about 19% fewer edges (fastapi
507 -> 426, django 64 -> 66). The same sample shows that no score
threshold reaches 0.90 precision, so the documented "~95% at 0.75" did
not hold before this change either.

The harness that measured it ships with the fix, off by default:
CBM_SEM_PAIR_SIGNALS=<path> writes one line per scored pair (both
qualified names, the score, every signal value, whether the pair was
admitted, both locations); CBM_SEM_PAIR_SIGNALS_FLOOR also records
pairs below the threshold, never admitting them. cbm_sem_combined_score
is now cbm_sem_signal_values + cbm_sem_combine (same arithmetic, same
order). CBM_SEMANTIC_INDEX_VERSION 3 -> 4: indexes built before score
pairs differently.

Tests: sem_tfidf_terms_compare_vocabulary_not_positions,
sem_combine_weighs_signals,
pipeline_semantic_pair_signals_never_change_the_graph; each is RED with
its behavior reverted.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
When a function has more candidate partners than its budget of ten
SEMANTICALLY_RELATED edges, admission took the pairs in discovery order:
the first ten won, whatever their scores. Admission is now best-first
(cbm_sem_admit_best_first: descending score, ties in the canonical
order, still a pure function of the sorted inputs). On a blind judged
sample from held-out repositories, the pairs first-come admitted over
budget were 0 of 17 related; the pairs best-first admits instead, 7 of
20. Where no budget binds nothing changes (django 66 edges identical;
fastapi 426 -> 430 edges, 222 changed).

Every SEMANTICALLY_RELATED edge now carries "p": the probability that a
reader judges the pair related, by score band, from 120 blind-judged
admitted pairs (score < 0.778: 0.56, < 0.806: 0.64, < 0.852: 0.73,
else 0.86). p is computed from the score as stored (three decimals),
so one shown score always has one p. Scores and the 0.75 threshold are
unchanged; a scoring change (TF-IDF alone, or weights fitted on the
first sample) was judged on the same sample and did not beat the
shipped score by the pre-registered margin.

Tests: sem_admit_best_first_keeps_the_best, sem_calibrated_p_bands,
pipeline_semantic_edges_carry_p; each is RED with its behavior
reverted.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The semantic pass read every string property of a function (body text,
signature, return type, decorators, docstring) with a reader that
stopped at the first '"' and kept escapes raw: a value holding an
escaped quote lost everything after it, and "\nreturn" became the token
"nreturn". cbm_sem_json_str honours the escapes: \" \\ \/ keep their
character, \n \t \r \b \f and \uXXXX become a separator.

The tokens change, so admissions change: on 13 repositories 1,340 ->
1,387 admitted SEMANTICALLY_RELATED pairs, 614 of them unchanged. Of
the dropped pairs that earlier samples had judged, 52 of 89 were
related (0.58); of the kept ones, 73 of 92 (0.79). A pre-registered
blind sample of 100 newly admitted pairs from 11 held-out repositories
(three reading rounds) judged 65 related (0.65, Wilson 0.55-0.74), so
the admitted set is 0.72 precise against 0.69 before. The same index
run twice gives identical pairs and scores.

"p" is refit on the new admitted population (bands at the quartiles of
the admitted scores, pooled where a higher band judged lower; added and
kept pairs weighted by their share): score < 0.773: 0.63, else 0.76
(was 0.56 / 0.64 / 0.73 / 0.86 in four bands).

Tests: sem_json_str_honours_escapes (RED with escapes ignored),
sem_calibrated_p_bands.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A documentation section that explains a function rarely names it, so
the deterministic doc links cannot connect them. The semantic pass now
ranks, for every Markdown section that can be about code, the functions
whose TF-IDF terms best match the section's heading and body, over the
functions' own corpus (function-pair scores are untouched):

- gate: no candidates for release-number or project-meta headings
  (changelog, licence, contributing, links, contents, install, ...),
  bodies under 20 words, or bodies that are mostly links;
- targets: functions and methods, never tests (by path or name);
- stored: a section's best five with tfidf >= 0.20, each with p, the
  probability that the section is about the function, from 274
  blind-judged candidates in five held-out repositories (three reading
  rounds): 0.36 below 0.30, 0.42 below 0.40, else 0.55. Below 0.20 the
  judged share was 0.06, so nothing is stored there.

These are candidates for an agent to verify, not graph edges: at best
about half are right, far below the bar for a link. They live in a side
table, doc_link_candidates (section and function qualified names, rank,
score, p, evidence), created when a generation is published, so older
indexes read as "no candidates". Full and legacy-partial reindexes
replace a project's rows; the delta route keeps the previous rows (its
proxy buffer has no section text or function tokens to recompute
them), and a row whose node is gone is skipped when read.

get_code_snippet shows them: on a Section, "possibly_about" (its
candidate functions); on a function or method, "possibly_described_in"
(the sections), at most five, highest p then score first, with a
candidates_note saying what they are.

Measured on fastapi (6,144 gated sections): 24,293 rows, identical to
the generator's dump; index time within noise; the database grows by
the table (37 MB there, with a 148-character project prefix on every
qualified name).

Tests: semantic:sem_doc_calibrated_p_bands,
store_arch:doc_candidates_replace_and_get (RED when a replace keeps old
rows), pipeline:pipeline_doc_candidates_published_and_kept_by_delta
(RED with test functions admitted, the heading gate off, the publish
write skipped, or the delta route replacing rows),
mcp:snippet_doc_candidates_both_directions (RED with stale rows shown).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The candidate rows and their strings were allocated with calloc/strdup
in the semantic pass and in cbm_store_doc_candidates_get, and released
with free. They now go through the memory core in CBM_MEM_CLASS_STORE
(cbm_calloc, cbm_mem_strdup, cbm_free) on both sides, so the rows the
pass hands to the pipeline and the rows a read returns are accounted
like every other store buffer. make lint-memory-core: no file grew.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
pipeline_doc_candidates_published_and_kept_by_delta appended its
delta-route edit to "user_config.go" while the fixture writes
"User_config.go". On a case-insensitive filesystem (macOS, Windows)
that is the same file; on Linux it is a new file, closure repair
declines new files (correctly), the run falls back to a full rebuild,
the route assertion fails and the early return leaks the pipeline
(LeakSanitizer on the arm64 GCC leg). The edit now names the fixture
file exactly.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…d home folders

Doc -> code candidates now carry a p judged for their kind, their document
format and where the document sits.

Format (sample 6, 234 blind-judged candidates): reStructuredText, AsciiDoc and
PDF sections each get their own curve (reST stored from tfidf 0.30, AsciiDoc
from 0.20, PDF not measured and not stored). A candidate that repeats an
exact MENTIONS link of the same section is not stored.

Position (sample 7, 490 blind-judged candidates from four monorepos): every
document gets a home folder -- its folder; a docs folder documents its
parent; a folder without code hands over to the nearest one with code; the
repository root means the whole project. Markdown candidates outside the
document's home were right 6 of 90 and are never stored; inside it they are
stored from tfidf 0.30 (0.33 / 0.53); for documents of the whole project
samples 4 and 7 pooled give 0.22 / 0.51 from 0.30.

New kinds, all in doc_link_candidates (evidence gains "kind" and
"position"; no new edge type, no schema change):
- local: up to two home functions ranked below the top five (0.32);
- file: a whole file, its vector the sum of its functions' terms and its
  definitions' names, config keys included, cut to 64 terms (0.30, from
  tfidf 0.40);
- folder: a README or index document -> its home folder (0.87).

get_code_snippet on a section lists every kind with its label;
get_file_outline shows the sections about the file or a folder holding it
(cbm_store_doc_candidates_for_path).

cbm_sem_p_2dp rounded 0.325 and 0.525 down (their float values sit 1.2e-6
below the half); its epsilon is 1e-4 on the x100 scale.

Replayed on the sample repositories, the build stores exactly the rows the
judged curves select (950 rows, 0 differences).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…cabulary

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Brings the doc_links block test fix.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData changed the base branch from feat/doc-links-documents to main October 10, 2026 07:35
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData merged commit 5f906a2 into main Oct 10, 2026
27 checks passed
DavidHLP pushed a commit to DavidHLP/codebase-memory-mcp that referenced this pull request Oct 11, 2026
The doc-section candidate pass (DeusData#2574) made a Linux kernel index 2.8x
slower: 207 s before DeusData#2574, 580 s after, on the same machine. Every
candidate of every section got a dot product that merged the section's up
to 1024 terms with the function's, and then every candidate was sorted for
a top 10.

- The section's vector is held densely per worker, indexed by term and
  stamped per section, so a candidate's dot product walks only the
  function's (or file's) own terms. The same products are added in the
  same ascending term order, so every score is the same float.
- Only the candidates the results read are sorted: the top
  SEM_DOC_TOP_K (and the local ones down to CBM_SEM_DOC_MIN_SCORE), the
  top CBM_SEM_DOC_FILE_K files. cmp_doc_cand is a total order, so the
  sorted prefix is the full sort's prefix.

Kernel (perf-bench/linux, M5 Pro, same plain build): 579.7 s -> 391.5 s,
nodes and edges unchanged. Exactness, measured directly: a build that also
ran the old merge and the old full sort found 0 score, 0 shared-term and
0 order differences over 5,003 django sections.

Candidate generation itself (a section's key terms, df <= funcs/20) is
unchanged here: narrowing it would change the stored rows.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant