Skip to content

Improve search: independent keywords, tags, and fuzzy matching - #1

Merged
FloLey merged 2 commits into
mainfrom
claude/search-improvements-cohru5
Jul 7, 2026
Merged

FloLey merged 2 commits into
mainfrom
claude/search-improvements-cohru5

Conversation

@FloLey

@FloLey FloLey commented Jul 6, 2026

Copy link
Copy Markdown
Owner

Why

The search was weak. search_wiki did a single case-insensitive substring scan, line by line, which meant:

  • Phrase-only matching — a multi-word query like florent soeur only hit when those exact words appeared adjacent on one line. Word order and spacing mattered.
  • Tags weren't searchable as tags — frontmatter tags: [...] was just text buried in a fenced block, not treated as first-class content.
  • Zero typo tolerance — one wrong letter (identty) returned nothing.

What changed

Reworked search_wiki (src/wiki_server/query.py) to:

  • Independent keywords — the query is split on whitespace and each keyword is matched on its own, so order and phrasing no longer matter.
  • Tag search — frontmatter tags are parsed and searched alongside page text, surfaced in results as line 0: path:0: [tags: ...].
  • Fuzzy matching — typos and near-misses still hit, via difflib (stdlib only — keeps the project's no-dependency design). Fuzzy only applies to tokens ≥4 chars at a 0.8 ratio, so short words and unrelated terms don't produce noise.
  • Relevance ranking — pages covering more of the query come first, with exact hits ranked above fuzzy ones.

Reported line numbers now skip the frontmatter block, so they still point at the real line in the file.

The MCP search tool docstring (src/wiki_server/server.py) is updated to describe the new behavior.

Tests

Added cases to tests/test_query.py covering keyword independence and order-insensitivity, coverage-based ranking, tag matches, fuzzy typo tolerance, frontmatter-aware line numbers, and the empty/no-match messages. Full suite: 81 passed.

Manual check against seed/: identty → matches identity, self → hits the [tags: self] line, zzznope → no false fuzzy matches.

🤖 Generated with Claude Code


Generated by Claude Code

The old search was a single case-insensitive substring scan, so a
multi-word query only matched when the exact phrase appeared on one line,
frontmatter tags were not searched as such, and any typo missed entirely.

Rework search_wiki to:
- split the query into independent keywords, matched order-insensitively;
- search frontmatter `tags` as first-class content (surfaced as line 0);
- tolerate typos via difflib fuzzy matching (stdlib, no new deps);
- rank pages that cover more of the query first, exact hits over fuzzy.

Line numbers now account for the frontmatter block so they still point at
the real file. Update the MCP tool docstring and add tests covering
keyword independence, ranking, tag hits, fuzzy typos, and line numbering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S9dBrktZvypAK5R6SXLXD9

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances the wiki search functionality by transitioning from a simple substring match to an order-insensitive keyword search that supports frontmatter tag matching, fuzzy matching for typo tolerance, and relevance-based ranking. Feedback on these changes focuses on performance and logic optimizations: specifically, optimizing the fuzzy matching loop using fast length-based pre-filters and quick ratios to reduce search latency, deduplicating query keywords to avoid redundant matching overhead, and refining the truncation logic to ensure the limit message is only displayed when actual truncation occurs.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread src/wiki_server/query.py
Comment on lines +184 to +188
for tok in tokens:
if len(tok) >= _FUZZY_MIN_LEN and \
difflib.SequenceMatcher(None, kw, tok).ratio() >= _FUZZY_THRESHOLD:
hits[kw] = _FUZZY
break

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Using difflib.SequenceMatcher on every token of every line when there is no exact match is highly inefficient and can cause significant search latency. We can optimize this by:

  1. Applying a fast length-based pre-filter (since a ratio of >= 0.8 mathematically requires the token length to be within a specific range relative to the keyword length).
  2. Leveraging SequenceMatcher.real_quick_ratio() and SequenceMatcher.quick_ratio() as fast upper-bound checks before computing the full, expensive ratio().
        for tok in tokens:
            if len(tok) >= _FUZZY_MIN_LEN:
                if 2 * min(len(kw), len(tok)) < _FUZZY_THRESHOLD * (len(kw) + len(tok)):
                    continue
                matcher = difflib.SequenceMatcher(None, kw, tok)
                if (matcher.real_quick_ratio() >= _FUZZY_THRESHOLD and
                        matcher.quick_ratio() >= _FUZZY_THRESHOLD and
                        matcher.ratio() >= _FUZZY_THRESHOLD):
                    hits[kw] = _FUZZY
                    break

Comment thread src/wiki_server/query.py Outdated
Comment on lines +158 to +160
def _keywords(query: str) -> list[str]:
"""Split a query into independent, lowercased keywords."""
return [k for k in query.lower().split() if k]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If the search query contains duplicate keywords (e.g., "florent florent"), the current implementation will perform redundant matching and ranking calculations for the duplicate terms on every line of every file. Deduplicating the keywords while preserving their order avoids this unnecessary overhead.

Suggested change
def _keywords(query: str) -> list[str]:
"""Split a query into independent, lowercased keywords."""
return [k for k in query.lower().split() if k]
def _keywords(query: str) -> list[str]:
"""Split a query into independent, lowercased keywords."""
return list(dict.fromkeys(k for k in query.lower().split() if k))

Comment thread src/wiki_server/query.py
Comment on lines +259 to +266
results: list[str] = []
for _coverage, _exact, rel, out_lines in pages:
for lineno, line in out_lines:
results.append(f"{rel}:{lineno}: {line}")
if len(results) >= max_results:
results.append(f"... (stopped at {max_results} matches)")
return "\n".join(results)
return "\n".join(results)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current truncation logic appends the "... (stopped at {max_results} matches)" message even when there are exactly max_results matches in total (meaning no actual truncation occurred). We can track whether there are additional matches beyond max_results to ensure the message is only appended when actual truncation happens.

Suggested change
results: list[str] = []
for _coverage, _exact, rel, out_lines in pages:
for lineno, line in out_lines:
results.append(f"{rel}:{lineno}: {line}")
if len(results) >= max_results:
results.append(f"... (stopped at {max_results} matches)")
return "\n".join(results)
return "\n".join(results)
results: list[str] = []
truncated = False
for _coverage, _exact, rel, out_lines in pages:
for lineno, line in out_lines:
if len(results) < max_results:
results.append(f"{rel}:{lineno}: {line}")
else:
truncated = True
break
if truncated:
break
if truncated:
results.append(f"... (stopped at {max_results} matches)")
return "\n".join(results)

From the PR review:
- Fuzzy matching now applies a length-based pre-filter and the cheaper
  real_quick_ratio/quick_ratio upper bounds before the full ratio(), and
  reuses one SequenceMatcher per keyword, avoiding a full comparison
  against every token of every line.
- De-duplicate query keywords so a repeated word is not matched and
  ranked twice.
- Only append the "stopped at N matches" line when matches actually
  exceed the cap, not when the total equals it exactly.

Add tests for the exact-limit (no truncation message), over-limit
(truncation message), and duplicate-keyword cases.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S9dBrktZvypAK5R6SXLXD9
@FloLey
FloLey merged commit 8315481 into main Jul 7, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants