Skip to content

Swift: overloaded methods collapse into one node, producing fake callers in trace_path #2061

Description

@jaganio

Version

0.10.8

Platform

macOS (Apple Silicon), Darwin 25.6.0

Summary

Two Swift methods on the same type that differ only in their parameter labels end up as a single node in the graph. The node that survives keeps one overload's line range, but it ends up holding the CALLS edges from both of them.

The practical fallout is that trace_path reports callers that don't exist. A function that only ever calls one overload shows up as a caller of whatever the other overload calls.

Overloading by argument label is pretty common in Swift (work(flag:) vs work(name:)), so this comes up a lot in real code.

Minimal reproduction

Three files, no dependencies, nothing to build.

Sources/Sink.swift

class Sink {
	func target() {}
}

Sources/Service.swift

class Service {
	let sink = Sink()

	// Overload A: DOES call target()
	func work(flag: Bool) {
		self.sink.target()
	}

	// Overload B: does NOT call target()
	func work(name: String) {
		print(name)
	}
}

Sources/Caller.swift

class Caller {
	let service = Service()

	// Calls ONLY overload B, which never reaches target()
	func onlyCallsOverloadB() {
		self.service.work(name: "x")
	}
}
codebase-memory-mcp cli index_repository --repo_path <repro> --name overload-repro --mode fast
codebase-memory-mcp cli search_graph --project overload-repro --name_pattern '^work$'
codebase-memory-mcp cli trace_path  --project overload-repro --function_name target --direction inbound --depth 3

Actual

search_graph finds one node where the source has two methods:

total: 1
  work Method 10-12 1 2

Lines 10-12 are overload B, work(name:). Overload A at lines 5-7 gets no node at all. But the surviving node reports out 2, and the body of work(name:) is just print(name), so it has picked up the call to target() from the overload that no longer exists.

trace_path then reports a caller that isn't real:

function: target
direction: inbound
callers_total: 2
  onlyCallsOverloadB 2      <- should not be here
  work 1

onlyCallsOverloadB calls work(name:), and that body is a single print, so it never reaches target() at any depth.

Expected

Two separate nodes, one per overload, with target() reachable only from work(flag:).

Notes

parse_partial_count is 0 for this repo, so nothing is failing to parse and this isn't a grammar issue. It looks like the qualified name doesn't include the parameter labels that tell Swift overloads apart, so the second declaration just overwrites the first.

This might belong to the same family as the fabricated CALLS edges filed for Rust (#2053) and Go (#1906, #1909, #1927), but the mechanism looks different. There's no stdlib receiver involved here, and both symbols are project-local methods on the same type.

Activity

  1. added
    bugSomething isn't working
    priority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.
    on Sep 5, 2026
  2. DeusData commented on Sep 25, 2026

    @DeusData
    Owner

    Thank you @jaganio again for the clean, dependency-free reproduction, and @DavidHLP for the exceptionally thorough investigation and candidate branch. The hardened-container reproduction, the repeat runs across worker counts and the RED/GREEN pipeline test are exactly the evidence we needed.

    We've settled the direction. While checking it we found the collapse isn't Swift-specific: Java, C#, C++ and Kotlin overloads merge into one node the same way today (on Alamofire, the 12 Session.upload overloads become a single node carrying callees from all of them). Because a qualified name is persisted graph identity, we're adopting one signature-qualified identity for every overloading language (Java, C#, C++, Kotlin, Scala, Objective-C, Swift, …): base_qn[<tparams>](params)[cvref], applied to every callable, not only when a sibling exists, so adding an overload never renames an existing node. The name column stays bare, so searching work still finds every overload, and each language ships behind a one-time index-format rebuild. When a call still matches several overloads after label/arity selection, we'll emit an edge to each compatible candidate, tagged with the candidate count, rather than guess or drop it.

    Your label model becomes the Swift instance of that scheme, with two extensions we found while measuring Alamofire:

    • Types alongside labels. upload(_ data: Data, to:…), upload(_ fileURL: URL, to:…) and upload(_ stream: InputStream, to:…) share the selector upload(_:to:…), so labels alone still merge them. The Swift form becomes work(flag:Bool) / work(name:String), and work() for zero-argument callables so the rule is uniform.
    • Label-compatible call resolution instead of exact selector equality. Default arguments and trailing closures need to match too, so call-site labels feed an overload selector rather than an exact-name lookup.

    The shared plumbing (the signature builder, base-QN-aware lookups and splitters, no language enabled yet, byte-identical graphs) is up as #2342. Languages will then be enabled one PR at a time. We'd love for you to carry the Swift PR on top of that, with your helper and regression test as its core. We'll ping you here as soon as the plumbing is on main. Thank you again for such careful work!

  3. added a commit that references this issue on Oct 4, 2026
  4. DavidHLP commented on Oct 5, 2026

    @DavidHLP
    Contributor

    I removed the original investigation while consolidating the progress comments. I'm restoring it here in full.

    This was first posted on 2026-09-14 (original comment 5662955218). The findings and results below belong to that checkpoint. The current Swift follow-up and review are in PR #2436; the original snapshot remains available.


    Exploration update (PR intentionally deferred)

    I am posting the investigation and reproducible evidence here first. I am not opening an upstream pull request yet; the candidate branch on my fork is only investigation material at this point.

    1. Scope and environment

    • Repository: DeusData/codebase-memory-mcp
    • Issue: Swift: overloaded methods collapse into one node, producing fake callers in trace_path #2061
    • Baseline tested: main at 0b51555402aa6aea9e5fd93d9ed53136dec4c6e8
    • Host: Arch Linux
    • Container: the repository test image built from test-infrastructure/Dockerfile (Ubuntu 24.04/Noble)
    • Container hardening used for the CLI reproduction: read-only root filesystem, --security-opt no-new-privileges, --cap-drop ALL, PID limit, and an isolated writable tmpfs for build/cache data

    This issue is reported on macOS Apple Silicon, but this is not a claim of native macOS validation. Docker shares the host kernel and cannot emulate macOS. The defect is in Swift symbol identity and graph resolution, so the deterministic semantic reproduction is still useful on Linux; native macOS CI or real macOS hardware remains the platform-specific follow-up.

    2. Minimal fixture

    The reproduction uses the same three dependency-free Swift files described in this issue:

    Sources/Sink.swift

    class Sink {
        func target() {}
    }

    Sources/Service.swift

    class Service {
        let sink = Sink()
    
        // Overload A: DOES call target()
        func work(flag: Bool) {
            self.sink.target()
        }
    
        // Overload B: does NOT call target()
        func work(name: String) {
            print(name)
        }
    }

    Sources/Caller.swift

    class Caller {
        let service = Service()
    
        // Calls ONLY overload B, which never reaches target()
        func onlyCallsOverloadB() {
            self.service.work(name: "x")
        }
    }

    The important invariant is that work(flag:) calls target(), while work(name:) does not, and onlyCallsOverloadB() calls only the latter.

    3. Reproduction procedure

    Inside the isolated container I built the production binary from the baseline commit and indexed the fixture in fast mode. The equivalent requests were:

    index_repository {"repo_path":"/repro","name":"overload-repro","mode":"fast"}
    search_graph {"project":"overload-repro","name_pattern":"^work$","format":"json","limit":20}
    trace_path {"project":"overload-repro","function_name":"target","direction":"inbound","depth":3,"format":"json"}
    

    The index status reported parse_partial_count=0, so this is not a grammar parse failure.

    4. Baseline result: the issue reproduces

    The unmodified baseline produced the same semantic failure reported here:

    nodes=17
    edges=25
    search_graph total=1
      work | Method | lines 10-12 | in=1 | out=2
    trace_path callers_total=2
      work | hop 1
      onlyCallsOverloadB | hop 2   <- false caller
    

    The only returned work node retains the line range of work(name:), but it has the outgoing CALLS edges from the other overload as well. Because onlyCallsOverloadB() calls that merged node, the traversal reaches target() through an edge that cannot exist in the source program.

    This gives a compact red assertion: one work node, the wrong merged degree/line identity, and a fabricated inbound caller.

    5. Root-cause analysis

    The failure is caused by losing Swift argument-label identity at more than one extraction boundary:

    1. Definition extraction generated the qualified name from the base callable name (work). Both declarations on Service therefore addressed the same graph key.
    2. Call extraction and unified call-scope construction also treated the callee as the base name, so the call resolver could not distinguish work(flag:) from work(name:).
    3. The graph merge was not independently inventing the bad edge; it was receiving two logically different definitions under one identity and accumulating their edges. The surviving node kept one declaration range while retaining relationships from both bodies.
    4. trace_path then correctly traversed the already-corrupted graph, which made the symptom look like a fabricated caller.

    The shared semantic rule needed by the graph is therefore:

    declaration: work(flag:) / work(name:)
    call site:   work(flag:) / work(name:)
    unlabeled argument: _
    no-argument callable: bare name
    

    6. Candidate solution explored locally

    The candidate change is on my fork at commit 0bf62ed, branch fix/issue-2061. It is not an upstream PR yet.

    The change applies one canonical Swift callable-name rule across the affected layers:

    • Adds the shared cbm_swift_callable_name helper in internal/cbm/helpers.c / helpers.h.
    • Uses declaration argument labels (external_name, falling back to the parameter name) when creating Swift definition qualified names.
    • Uses call-site argument labels when creating the resolution name; unlabeled arguments are represented as _.
    • Keeps the raw CBMCall.callee_name for source-facing data and adds a canonical resolution_name for lookup.
    • Applies canonical lookup in both sequential and parallel call-resolution passes.
    • Preserves resolution_name when synthetic calls are copied across the LSP arena.
    • Adds a pipeline regression test that asserts exact node names, line ranges, degrees, CALLS edges, and inbound trace rows.

    This is intended as a root-cause fix rather than a trace_path-specific filter: definitions, calls, resolution, and the regression contract all use the same overload identity.

    7. Candidate green observations already obtained

    Before pausing this work, the candidate was checked with:

    • Production Docker build using the Makefile's -Werror flags: passed.
    • Strict CLI assertions: 6/6 passed, with three independent runs under CBM_WORKERS=1 and three under CBM_WORKERS=4.
    • Fixed result in those runs: nodes=18, edges=27; search_graph returned exactly work(flag:) and work(name:); only work(flag:) had the edge to target(); inbound trace_path returned exactly work(flag:).
    • Focused extraction plus pipeline suites: 621 passed, including the new regression test.

    The remaining pre-submission gates are deliberately not being presented as green here: the full scripts/test.sh run and scripts/lint.sh --ci still need to be completed, and native macOS validation has not been performed.

    8. Request for maintainer review

    Please confirm whether this identity model matches the project's intended Swift qualified-name semantics, especially the representation of unlabeled arguments and any edge cases involving default argument labels. Once that direction is confirmed, I can resume the full gate run and open a focused PR with the complete red/green evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingparsing/qualityGraph extraction bugs, false positives, missing edgespriority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions