Skip to content

perf(semantic): exact, cheaper doc-section candidate scoring - #2582

Merged
DeusData merged 1 commit into
mainfrom
fix/doc-section-exact-topk
Oct 10, 2026
Merged

DeusData merged 1 commit into
mainfrom
fix/doc-section-exact-topk

Conversation

@DeusData

Copy link
Copy Markdown
Owner

The doc-section candidate pass from #2574 made indexing the Linux kernel about 2.8x slower: 207 s before #2574 and 580 s after it, same machine, same build. A teammate's call trees put about 85 % of that time in the per-candidate dot products and about 14 % in sorting.

What changes (output unchanged)

  • Dot product. Each section's term vector is kept as a dense array per worker, so a candidate's dot product reads only the candidate's own 20–60 terms. Before, it merged them with the section's up to 1,024 terms. The products and the order of addition are the same, so every score is the same float.
  • Sorting. Only the candidates the results actually read are sorted: the top 10, plus every candidate scoring 0.20 or more (the "local" hits) for functions, and the top 3 for files. The comparator cmp_doc_cand is a total order, so sorting this subset gives exactly the same first entries as the full sort did.

Measured

  • Linux kernel (perf-bench/linux, M5 Pro, same plain build, quiet machine): 579.7 s → 391.5 s. Nodes (8,540,291) and edges (16,765,478) are identical.
  • Exactness was checked directly. A scratch build ran the old merge dot product and the old full sort next to the new code on django: 0 score differences (bitwise), 0 shared-term differences and 0 order differences over 5,003 sections.
  • Suites: semantic, pipeline, pipeline_semantic_manifest_repro, store_arch, mcp all pass. cppcheck 2.20, clang-format and the memory-core lint are clean.

Not in this PR

  • Remaining cost: candidate generation is unchanged. A term counts as a key term when at most one in 20 functions uses it, so on the kernel that's tens of thousands of functions per term. Narrowing that would change the stored rows, so it needs its own decision.
  • Run-to-run differences in doc_link_candidates: the kernel A/B also showed these, but they come from main itself (two runs of the same build differ). The cause is in the reST section text (rst_body); a separate fix follows.

The doc-section candidate pass (#2574) made a Linux kernel index 2.8x
slower: 207 s before #2574, 580 s after, on the same machine. Every
candidate of every section got a dot product that merged the section's up
to 1024 terms with the function's, and then every candidate was sorted for
a top 10.

- The section's vector is held densely per worker, indexed by term and
  stamped per section, so a candidate's dot product walks only the
  function's (or file's) own terms. The same products are added in the
  same ascending term order, so every score is the same float.
- Only the candidates the results read are sorted: the top
  SEM_DOC_TOP_K (and the local ones down to CBM_SEM_DOC_MIN_SCORE), the
  top CBM_SEM_DOC_FILE_K files. cmp_doc_cand is a total order, so the
  sorted prefix is the full sort's prefix.

Kernel (perf-bench/linux, M5 Pro, same plain build): 579.7 s -> 391.5 s,
nodes and edges unchanged. Exactness, measured directly: a build that also
ran the old merge and the old full sort found 0 score, 0 shared-term and
0 order differences over 5,003 django sections.

Candidate generation itself (a section's key terms, df <= funcs/20) is
unchanged here: narrowing it would change the stored rows.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData merged commit 76a2485 into main Oct 10, 2026
28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant