Improve search: independent keywords, tags, and fuzzy matching - #1
Conversation
The old search was a single case-insensitive substring scan, so a multi-word query only matched when the exact phrase appeared on one line, frontmatter tags were not searched as such, and any typo missed entirely. Rework search_wiki to: - split the query into independent keywords, matched order-insensitively; - search frontmatter `tags` as first-class content (surfaced as line 0); - tolerate typos via difflib fuzzy matching (stdlib, no new deps); - rank pages that cover more of the query first, exact hits over fuzzy. Line numbers now account for the frontmatter block so they still point at the real file. Update the MCP tool docstring and add tests covering keyword independence, ranking, tag hits, fuzzy typos, and line numbering. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S9dBrktZvypAK5R6SXLXD9
There was a problem hiding this comment.
Code Review
This pull request enhances the wiki search functionality by transitioning from a simple substring match to an order-insensitive keyword search that supports frontmatter tag matching, fuzzy matching for typo tolerance, and relevance-based ranking. Feedback on these changes focuses on performance and logic optimizations: specifically, optimizing the fuzzy matching loop using fast length-based pre-filters and quick ratios to reduce search latency, deduplicating query keywords to avoid redundant matching overhead, and refining the truncation logic to ensure the limit message is only displayed when actual truncation occurs.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| for tok in tokens: | ||
| if len(tok) >= _FUZZY_MIN_LEN and \ | ||
| difflib.SequenceMatcher(None, kw, tok).ratio() >= _FUZZY_THRESHOLD: | ||
| hits[kw] = _FUZZY | ||
| break |
There was a problem hiding this comment.
Using difflib.SequenceMatcher on every token of every line when there is no exact match is highly inefficient and can cause significant search latency. We can optimize this by:
- Applying a fast length-based pre-filter (since a ratio of >= 0.8 mathematically requires the token length to be within a specific range relative to the keyword length).
- Leveraging
SequenceMatcher.real_quick_ratio()andSequenceMatcher.quick_ratio()as fast upper-bound checks before computing the full, expensiveratio().
for tok in tokens:
if len(tok) >= _FUZZY_MIN_LEN:
if 2 * min(len(kw), len(tok)) < _FUZZY_THRESHOLD * (len(kw) + len(tok)):
continue
matcher = difflib.SequenceMatcher(None, kw, tok)
if (matcher.real_quick_ratio() >= _FUZZY_THRESHOLD and
matcher.quick_ratio() >= _FUZZY_THRESHOLD and
matcher.ratio() >= _FUZZY_THRESHOLD):
hits[kw] = _FUZZY
break| def _keywords(query: str) -> list[str]: | ||
| """Split a query into independent, lowercased keywords.""" | ||
| return [k for k in query.lower().split() if k] |
There was a problem hiding this comment.
If the search query contains duplicate keywords (e.g., "florent florent"), the current implementation will perform redundant matching and ranking calculations for the duplicate terms on every line of every file. Deduplicating the keywords while preserving their order avoids this unnecessary overhead.
| def _keywords(query: str) -> list[str]: | |
| """Split a query into independent, lowercased keywords.""" | |
| return [k for k in query.lower().split() if k] | |
| def _keywords(query: str) -> list[str]: | |
| """Split a query into independent, lowercased keywords.""" | |
| return list(dict.fromkeys(k for k in query.lower().split() if k)) |
| results: list[str] = [] | ||
| for _coverage, _exact, rel, out_lines in pages: | ||
| for lineno, line in out_lines: | ||
| results.append(f"{rel}:{lineno}: {line}") | ||
| if len(results) >= max_results: | ||
| results.append(f"... (stopped at {max_results} matches)") | ||
| return "\n".join(results) | ||
| return "\n".join(results) |
There was a problem hiding this comment.
The current truncation logic appends the "... (stopped at {max_results} matches)" message even when there are exactly max_results matches in total (meaning no actual truncation occurred). We can track whether there are additional matches beyond max_results to ensure the message is only appended when actual truncation happens.
| results: list[str] = [] | |
| for _coverage, _exact, rel, out_lines in pages: | |
| for lineno, line in out_lines: | |
| results.append(f"{rel}:{lineno}: {line}") | |
| if len(results) >= max_results: | |
| results.append(f"... (stopped at {max_results} matches)") | |
| return "\n".join(results) | |
| return "\n".join(results) | |
| results: list[str] = [] | |
| truncated = False | |
| for _coverage, _exact, rel, out_lines in pages: | |
| for lineno, line in out_lines: | |
| if len(results) < max_results: | |
| results.append(f"{rel}:{lineno}: {line}") | |
| else: | |
| truncated = True | |
| break | |
| if truncated: | |
| break | |
| if truncated: | |
| results.append(f"... (stopped at {max_results} matches)") | |
| return "\n".join(results) |
From the PR review: - Fuzzy matching now applies a length-based pre-filter and the cheaper real_quick_ratio/quick_ratio upper bounds before the full ratio(), and reuses one SequenceMatcher per keyword, avoiding a full comparison against every token of every line. - De-duplicate query keywords so a repeated word is not matched and ranked twice. - Only append the "stopped at N matches" line when matches actually exceed the cap, not when the total equals it exactly. Add tests for the exact-limit (no truncation message), over-limit (truncation message), and duplicate-keyword cases. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S9dBrktZvypAK5R6SXLXD9
Why
The search was weak.
search_wikidid a single case-insensitive substring scan, line by line, which meant:florent soeuronly hit when those exact words appeared adjacent on one line. Word order and spacing mattered.tags: [...]was just text buried in a fenced block, not treated as first-class content.identty) returned nothing.What changed
Reworked
search_wiki(src/wiki_server/query.py) to:tagsare parsed and searched alongside page text, surfaced in results as line0:path:0: [tags: ...].difflib(stdlib only — keeps the project's no-dependency design). Fuzzy only applies to tokens ≥4 chars at a 0.8 ratio, so short words and unrelated terms don't produce noise.Reported line numbers now skip the frontmatter block, so they still point at the real line in the file.
The MCP
searchtool docstring (src/wiki_server/server.py) is updated to describe the new behavior.Tests
Added cases to
tests/test_query.pycovering keyword independence and order-insensitivity, coverage-based ranking, tag matches, fuzzy typo tolerance, frontmatter-aware line numbers, and the empty/no-match messages. Full suite: 81 passed.Manual check against
seed/:identty→ matchesidentity,self→ hits the[tags: self]line,zzznope→ no false fuzzy matches.🤖 Generated with Claude Code
Generated by Claude Code