Skip to content

Diff attribution: establish dependency coverage before excluding a standing finding #557

Description

@pengfei-threemoonslab

Whole-function source comparison — 2026-09-09

Delivered in #606, merged 7a5abcdd897b71205d6d53e111b98b6b3e4e96ac. The additional sdk_boolean_function/v1 profile inspects the entire finite Boolean function, the selected pure imported predicate and one literal source Agent binding. It distinguishes return values, true-return input sets, source membership and the bound true-return relation. A narrowed first guard can coexist with a wider whole-function true domain; unchanged returns can coexist with added membership. Unknown operations/configuration, unsupported or unreachable statements, ambiguous identities and redaction retain unresolved evidence.

The model's complete Boolean-domain table is a source comparison, not action effect, human approval or deployed authority. The outer capability dependency coverage remains incomplete, every finding-exclusion flag remains false, and the acceptance checklist below is deliberately unchanged. Newly isolated sub-issue #607 owns the relationship between a modeled operation and an existing finding's policy predicates; it is deferred from this implementation pass. The supported operation/effect/dependency boundary, #515's TypeScript cal-1 case and #563/#312's reviewed historical evidence remain outstanding.

Validation: final full suite 9,347 passed, 5 skipped, 113 focused checks, exact final-head CI including aggregate coverage and current-main verification passed. Review 5149940982 found multiplicative helper evaluation; the regression reproduced 202,496 helper invocations and the fix reduces it to exactly two for that Boolean domain. The independent reviewer rechecked all 10 final published blobs; address 5596207129 records the fix and clean confirmation. #606 has one review/address loop, bringing this issue to three across #599/#606 in this implementation pass. No release-qualification claim is made.

First bounded implementation — 2026-09-09

Delivered in #599, merged 5c65f19c62a8ba64ce5d98d5be541b13bb762a39. The OpenAI Agents SDK reader now records and compares imported Boolean guard evidence for canonical tool observations, including source/definition locations, dependency digests and an explicit Boolean domain. Real paired repositories distinguish unchanged, wider, narrower and otherwise changed predicates. Missing, dynamic, ambiguous and redacted evidence remains unresolved. Actual dependency bytes and named present/absent/unconfirmable import candidates are bound to verifier identity and checked for current-control currency.

The first acceptance item is delivered for this documented profile. The remaining items retain broader obligations: every record still says dependency_coverage: incomplete and finding_exclusion_eligible: false; no standing finding is excluded, no default scope or release decision changes, and #312's historical safety measurement has not been satisfied. Full binding/configuration dependency closure and #515 remain open.

Validation: two independent GitHub coding-agent review/address rounds; a looping-symlink crash found in round 1 was fixed and independently rechecked in round 2. Full local suite 9,193 passed, 5 skipped; final-head CI, aggregate coverage and current-base verification passed. Deferred follow-ups are #596, #597 and #598. These checks do not establish release qualification.

Problem

Implementation of #515 found a narrower prerequisite than a CLI scope flag. ToolSurfaceDiffReference reduces base findings to identity rows; _finding_deltas joins those rows on fingerprint. Finding.support deliberately lives outside that fingerprint. A guard/policy predicate can therefore change its observed value or blocking eligibility while the finding remains in unchanged_findings. The current tool/action facts also do not assert a complete shared-helper/import/configuration dependency closure.

The #515 implementation pass preserves base finding evidence and exposes these comparison limits. It cannot establish causality merely by comparing that evidence. An unchanged support object may mean the reader never modeled the changed dependency, rather than that authority stayed unchanged.

Required implementation

Start with one supported source reader and a real paired repository fixture. Preserve the source/dependency evidence actually read for each affected capability, including shared guard/helper, binding and configuration inputs. Compare complete base/head snapshots. Track incomplete, dynamic, missing-base and ambiguous-subject cases explicitly; no missing edge may be interpreted as absence of influence.

Only after that proof may #515 consume attribution in the release decision and introduce the requested scope in report/receipt identity. Fingerprint equality, changed-files-only scans and a supplied support hash are not substitutes. Keep whole-tree findings available for audit and preserve the approved safety catch bars.

Acceptance

  • An unchanged tool declaration whose imported guard is weakened identifies the affected capability and base/head guard evidence (bounded SDK Boolean profile in Compare imported SDK guard predicates with bound input evidence #599).
  • A standing weakness, strict improvement and newly widened capability are distinguished without relabeling unresolved dependency coverage as safe.
  • Missing base inputs, dynamic dependencies, configuration indirection and ambiguous shared subjects cannot cause a finding to be excluded.
  • The evidence used to claim complete comparison is bound to the same verifier input identity; a stale receipt cannot validate a different scope.
  • benchmark: re-run the labeled corpus on the v0.18 candidate and compare with W27 #312's fixed-history measurement demonstrates the safety impact before the default changes.

The first bounded evidence profile is delivered; the complete dependency claim remains unresolved. #515 stays open; this is its dependency-proof prerequisite, not a second scope implementation. Refs #518, #520, #312.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Next after P0; blocks other work or ships a misleading resultarea:identityVerification identity, receipts, reproducibilityenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions