Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
106 changes: 106 additions & 0 deletions corpus/verdicts/can4hou6joeng4-boss-agent-cli.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
{
"repository": "C:\\Users\\Elfrost\\AppData\\Local\\Temp\\ai-patchlab-clone-10x7xz2a\\repo",
"generated_at": "2026-09-23T13:18:38.997055+00:00",
"total_dismissed": 27,
"by_reason": {
"ignore-pattern": 2,
"below-min-severity": 1,
"sql-identifier-fp": 8,
"by-design": 12,
"not-reachable": 2,
"product-surface": 1,
"confirmed-real": 1
},
"records": [
{
"source": "scanner",
"reason_code": "ignore-pattern",
"tool": "semgrep",
"rule": "html.security.audit.missing-integrity.missing-integrity",
"count": 2,
"verdict": "",
"detail": "Suppressed by an ignore pattern."
},
{
"source": "scanner",
"reason_code": "below-min-severity",
"tool": "trivy",
"rule": "No HEALTHCHECK defined",
"count": 1,
"verdict": "",
"detail": "Below the --min-severity floor (medium)."
},
{
"source": "curation",
"reason_code": "sql-identifier-fp",
"tool": "semgrep",
"rule": "python.sqlalchemy.security.sqlalchemy-execute-raw-query / formatted-sql-query",
"count": 8,
"verdict": "false-positive",
"detail": "5 high + 3 medium. Every f-string interpolates only a table or column NAME, all hardcoded literals at the call sites (stats.py _safe_count/_count_since called with 'greet_records'/'apply_records'/'greeted_at'; clean.py _clean_table with literal table names; cache/store.py joins in-code \"col = ?\" strings). Every VALUE is bound as ? — since is a ? param, the crawl_runs UPDATE binds all values. No user input reaches any SQL string. #1 recurring identifier FP."
},
{
"source": "curation",
"reason_code": "by-design",
"tool": "semgrep",
"rule": "typescript.react.security.react-insecure-request",
"count": 1,
"verdict": "by-design",
"detail": "extension/background.js:145 is fetch(http://127.0.0.1:19826/ping) — a loopback health check to the local bridge daemon, not a react render sink. http to loopback is intended."
},
{
"source": "curation",
"reason_code": "by-design",
"tool": "semgrep",
"rule": "python.lang.compatibility.python37.python37-compatibility-importlib2",
"count": 1,
"verdict": "not-applicable",
"detail": "endpoints_loader.py:2 importlib.resources — a Python 3.7 compatibility rule, not a security finding. Project targets 3.11+."
},
{
"source": "curation",
"reason_code": "by-design",
"tool": "semgrep",
"rule": "github-actions.security.github-actions-mutable-action-tag",
"count": 8,
"verdict": "hardening",
"detail": "8 of 16 mediums, one rule across pages.yml/release.yml/star-history.yml. Unpinned action tags — supply-chain hardening worth a SHA pin, not a vulnerability. Same-rule flood inflates the band."
},
{
"source": "curation",
"reason_code": "not-reachable",
"tool": "semgrep",
"rule": "python.lang.security.audit.dynamic-urllib-use-detected",
"count": 2,
"verdict": "by-design",
"detail": "browser_client.py:688 urlopen(cdp_http_url + '/json') is the local Chrome DevTools endpoint the user launched; scripts/update_contributors.py:27 fetches api.github.com in a build script. Neither is an attacker-controlled URL at runtime."
},
{
"source": "curation",
"reason_code": "by-design",
"tool": "semgrep",
"rule": "python.lang.security.insecure-hash-algorithm-sha1",
"count": 2,
"verdict": "by-design",
"detail": "zhilian_browser.py:305 and events.py:27 use sha1()[:16] to derive a dedup/reference id from element text; non-security digest, not a password or signature."
},
{
"source": "curation",
"reason_code": "product-surface",
"tool": "semgrep",
"rule": "python.lang.security.audit.non-literal-import",
"count": 1,
"verdict": "by-design",
"detail": "__init__.py:94 dynamic import is the plugin loader — the extensibility architecture, not an injection sink (import name is not user-supplied)."
},
{
"source": "curation",
"reason_code": "confirmed-real",
"tool": "manual-review",
"rule": "bridge-daemon-loopback-no-origin-host-guard",
"count": 1,
"verdict": "real",
"detail": "Browser Bridge daemon (127.0.0.1:19826, aiohttp, no middleware) exposes POST /command and the /ext WebSocket with NO Origin and NO Host check and no shared secret. Verified server-side against the shipped daemon: foreign Host served (200), text/plain JSON body parsed and forwarded (CORS-simple request, no preflight), hostile Origin not rejected. The extension executes action=exec via chrome.debugger Runtime.evaluate and action=cookies via chrome.cookies.getAll (returns httpOnly) on the *.zhipin.com session; workspace='boss' auto-targets the logged-in tab. Any web page visited while bridge mode is active can blind-CSRF arbitrary JS into the authenticated recruitment tab, and read httpOnly session cookies via DNS rebinding. Desktop-loopback-inversion class. Reported privately (strict-norm: SECURITY.md forbids public issues; PVR enabled)."
}
]
}
7 changes: 4 additions & 3 deletions docs/index.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
layout: default
title: AI PatchLab Scans
description: "109 curated security scans of open-source AI agents, MCP servers and LLM apps - 25 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit."
description: "110 curated security scans of open-source AI agents, MCP servers and LLM apps - 25 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit."
---

# AI PatchLab Scans
Expand All @@ -20,7 +20,7 @@ remediation and confidence rules to normalize the findings.

> **Want this run privately against your own codebase?** I do independent
> security review of AI agents, MCP servers, and LLM apps —
> [**work with me →**]({{ '/work-with-me' | relative_url }}). 109 scans, 25 confirmed fixes, methodology in the open.
> [**work with me →**]({{ '/work-with-me' | relative_url }}). 110 scans, 25 confirmed fixes, methodology in the open.

> **OpenAI just launched [Daybreak](https://openai.com/index/daybreak-securing-the-world/) and Patch the Planet.**
> Same remediation loop, opposite trade-off: their path is a cloud frontier model;
Expand Down Expand Up @@ -103,7 +103,7 @@ login and static assets. Fifty-two flagged, none reported.

## All scans

109 scans, newest first. **Findings** is the raw count the tools produced;
110 scans, newest first. **Findings** is the raw count the tools produced;
**Real** is what survived curation. The gap between those two columns is the
entire job.

Expand All @@ -118,6 +118,7 @@ filed, which is the usual outcome of a clean scan.

| Date | Repository | Findings | Real | Outcome |
| --- | --- | ---: | --- | --- |
| 2026-09-23 | [can4hou6joeng4/boss-agent-cli](scans/can4hou6joeng4-boss-agent-cli.html) | 26 | 1 real — withheld | private |
| 2026-09-22 | [overwirehq/claude-code-telegram](scans/overwirehq-claude-code-telegram.html) | 54 | 0 first-party — dependency | — |
| 2026-09-21 | [HarnessRouter/harnessrouter](scans/harnessrouter-harnessrouter.html) | 117 | 1 real — dependency | — |
| 2026-09-20 | [TencentCloud/Octop](scans/tencentcloud-octop.html) | 300 | 1 real — withheld | private |
Expand Down
3 changes: 2 additions & 1 deletion docs/scan-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,9 @@ description: "The complete AI PatchLab scan log: every public repository scanned

# Full scan log

Every scan in the series, newest first, with the summary written on the day of the scan. 109 scans. For the compact index, see the [scan log home]({{ '/' | relative_url }}).
Every scan in the series, newest first, with the summary written on the day of the scan. 110 scans. For the compact index, see the [scan log home]({{ '/' | relative_url }}).

- **2026-09-23** — [can4hou6joeng4/boss-agent-cli](scans/can4hou6joeng4-boss-agent-cli.html) — 26 findings at `medium+` (7 high, 16 medium, 3 meta), **1 real — withheld, filed privately** — a **CLI for the BOSS Zhipin recruitment platform** (2k★, MIT; a human drives it from a terminal wizard, an agent drives the same workflows over a JSON-envelope API, an ~77-tool MCP server, or a Python SDK) that authenticates by reusing the operator's own logged-in browser session. Responsive maintainer — 67 merged PRs from 7 human authors in 60 days — which is why it cleared the pre-check. **Zero of the 26 scanner findings is the vulnerability**, which is the whole point of the entry: the real item is the *absence* of a check on an otherwise-working local channel, the shape no rule can match, and it is the [desktop-loopback-inversion](scans/liaohch3-claude-tap.html) class — a `127.0.0.1` control surface that binds neither *who* may call it nor which host the request claims, driving privileged browser actions against a site where the operator is authenticated. Confirmed with a **server-side exploit primitive against the shipped daemon** (a foreign `Host` is served, a cross-origin "simple" request reaches the command channel with no preflight, a hostile `Origin` is not rejected) plus a read of the extension source for the browser half — the [run-the-primitive discipline](scans/observal-observal.html), because reading alone could not separate "the channel is drivable" from "the channel is drivable *and* blind does not blunt it." Graded **High not Critical** on an honest precondition (the bridge component must be actively running — a bounded window, not an always-on service) and the sensitive scope is the one site the tool is built around; **the fix was written before it was named**, narrow and verified not to break the tool's own two legitimate clients ([write the patch first](scans/roflcoopter-viseron.html)). The 26: eight SQL hits (five high) are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) — every interpolation is a hardcoded table/column name with values bound as `?`, settled by reading the call sites; eight of sixteen mediums are one `github-actions-mutable-action-tag` flood; the rest are two `sha1()[:16]` dedup-id digests, two `urllib` calls to a local DevTools endpoint and a build script, a React rule on a loopback health-check `fetch`, and a Python-3.7 compatibility rule that is not a security check. Coverage `partial` and stated so — Semgrep did not cover every file and the dependency scan left `uv.lock` unaudited (Trivy read it, the [luck that hides the bug](scans/roflcoopter-viseron.html)), neither touching the finding. **Strict-norm** (`SECURITY.md` forbids public vulnerability issues, gives an email and a PVR link) · PVR-enabled, filed **privately, accepted into triage first try** ([GHSA-xp54-2hxc-w78q](https://github.com/can4hou6joeng4/boss-agent-cli/security/advisories/GHSA-xp54-2hxc-w78q)) **with the `vulnerabilities` array** — the [viseron rule](scans/roflcoopter-viseron.html) holds again · post-only, finding withheld at class level under embargo
- **2026-09-22** — [overwirehq/claude-code-telegram](scans/overwirehq-claude-code-telegram.html) — 54 findings at `medium+` (1 critical, 21 high, 30 medium), **0 first-party defects — dependency currency, post-only** — a **Telegram bot that gives authorised users remote access to Claude Code**, i.e. it runs file and Bash tools on a host machine by design, so the whole security surface *is* the boundary around that execution (2k★, MIT, org-backed; active — 15 merged PRs from 7 authors and 9 closed issues in 60 days; strict-norm: real `SECURITY.md`, PVR enabled). Picked with the manual queue at zero. **The story is what precise security documentation looks like.** The real boundary is the `can_use_tool` callback that stops Claude — when steered off-course by content it reads mid-task — from acting outside the approved directory, and the maintainers know and *state* exactly what it covers: path validation is scoped in the docs to precisely six tools (`Read`, `Write`, `Edit`, `MultiEdit`, `NotebookEdit`, `NotebookRead`), while `Grep`/`Glob`/`LS` are deliberately left to the OS sandbox. Reading that as "Read is guarded but Grep is not, so the boundary is incomplete" is the plausible-but-wrong finding this site exists to resist — the boundary is [advertised, not accidental](scans/tracecathq-tracecat.html), and `ROADMAP-v2.md` even records the SDK's own `CanUseToolShadowedWarning` naming every tool the callback will never see, filed as tracking item #221. Two more choices earn credit: guarded tools are deliberately *stripped from* the SDK `allowed_tools` list (a pre-approved tool never produces a `can_use_tool` request, so leaving them in would render the checks silently inert — [issue #219](https://github.com/overwirehq/claude-code-telegram/issues/219)), and `autoAllowBashIfSandboxed` is disabled whenever the boundary checks must run, closing a second bypass. That is a maintainer who traced how the framework resolves permissions rather than trusting that a config allow-list is an enforcement boundary — it isn't, and the code says so. **When the machine is built and documented this carefully, the finding moves to the dependency manifest.** The 30 Trivy advisories in `poetry.lock` split cleanly by reachability: **17 are opt-in-only and not reachable on a default install** — `starlette` (6) and `python-multipart` (3) need the FastAPI webhook server (`ENABLE_API_SERVER=false` default); `pyjwt` (5) needs token auth (`ENABLE_TOKEN_AUTH=false` default, and the SECURITY.md documents token auth as non-functional, [#58](https://github.com/overwirehq/claude-code-telegram/issues/58)); the MCP SDK highs (3) need MCP enabled — while **13 are unconditionally installed** (`anyio` incl. the critical IDNA/TLS advisory, the `cryptography`/OpenSSL cluster, `urllib3`, `idna`, `requests`, `pydantic-settings`, `python-dotenv`) and are the genuine currency gap. The direct deps are current; the drift is transitive, which the repo's existing `.github/dependabot.yml` (configured for *version* updates, not *security* updates) will not raise on its own. **No advisory filed** — there is no first-party vulnerability. Of the other 24: 16 `github-actions-mutable-action-tag` (SHA-pin hardening); one `pull_request_target` checkout wrapped in a 40-line threat-model header (read-only tool allowlist, secrets scrubbed, "worst case is a prompt-injected review comment" — [the trigger decides severity](scans/lightseekorg-tokenspeed.html)); two `sqlalchemy-execute-raw-query` [identifier FPs](scans/aurelio-labs-semantic-router.html) (only a generated run of `?` placeholders is interpolated, values bound via `execute(query, params)`); and three gitleaks hits that are a `your-api-secret` placeholder plus two fake tokens in a test asserting the project's own `_redact_secrets()` helper scrubs them ([credited defence](scans/realiti4-claude-swap.html)). Coverage `partial` and honestly so: pip-audit resolved **no dependencies** from a `pyproject.toml` declaring `dynamic = ["dependencies"]` — reading identically to "clean" — and Trivy's `poetry.lock` read carried the run, the [same dependency blind spot](scans/harnessrouter-harnessrouter.html) as yesterday reached by a different manifest shape.

- **2026-09-21** — [HarnessRouter/harnessrouter](scans/harnessrouter-harnessrouter.html) — 117 findings (2 critical, 42 high, 70 medium), **1 real — dependency currency, post-only** — a **self-hosted control plane that turns agent CLIs (Codex, Claude Code, Hermes, and a dozen more) into a single OpenAI-compatible API** (1.7k★, Apache-2.0, org-backed with a hosted "Cloud" edition; very active — 91 merged PRs from 5 authors in 60 days). Picked with the manual queue at zero: strict-norm (real `SECURITY.md`, PVR enabled, commercial backing), which is a fair target when the backlog is clear. **The story here is a well-built authorization layer that survived the sweep the series usually breaks projects on, so the one finding moved to the dependency manifest.** The route inventory — 113 gateway routes — comes back with exactly five answering without a credential once the project's own identity idioms (`_owned_session`, `_pub_org_member`, `_principal`) are resolved: `/v1/uhp`, `/healthz`, `/readyz`, `/version`, and `/share/{token}` where the unguessable token *is* the credential. That is the [name-matched sweep](scans/mai-with-u-maibot.html) returning empty for the right reason. Two design choices earn explicit credit: the self-hosted BFF stamps its internal trust key onto a gateway call **only for a request carrying a valid session cookie** — an earlier build stamped it unconditionally, and the code comments document catching and fixing exactly that [conditional-verification](scans/sentelabsai-openexecutive.html) class before I arrived — and the gateway **binds loopback with only the UI port published**, so the [DNS-rebinding shape](scans/liaohch3-claude-tap.html) that has caught several desktop apps here does not apply (the published surface is the session-gated Next.js app with `X-Frame-Options: DENY`). **The actionable finding is dependency currency, and it only surfaced because the two dependency tools disagreed about whether there was anything to scan.** pip-audit reported **no manifest** — it scans the repo root, and the requirements live in `gateway/requirements.txt` and `runner/requirements.txt` one level down — which on a monorepo reads identically to "clean". Trivy's whole-tree walk read both and found the pinned `next` **15.5.23** is one patch behind **15.5.24**, which fixes two criticals: [CVE-2026-75604](https://github.com/advisories) (Windows-hosted RCE — **dropped, the image is Linux**) and [GHSA-2xp9-vwfh-vxw4](https://github.com/advisories) (AVIF image-optimizer RCE — the optimizer runs at its default-enabled setting and the middleware matcher **excludes `_next/image`**, so the surface is reachable unauthenticated; default-empty `remotePatterns` constrains full exploitation, and with no Docker Linux engine available I could not run the primitive, so I claim the reachable surface and the currency gap, not a demonstrated RCE). `gateway/requirements.txt` also carries `PyJWT` 2.10.1 (CVE-2026-48526, auth-bypass) and `cryptography`/`aiohttp`/`python-multipart` CVEs — low reachability self-hosted (gateway loopback, `HR_IDENTITY_MODE=off`) but the hosted build shares the code. **No advisory filed**: the actionable item is a published upstream CVE with a one-line fix (`next >=15.5.24` + a `dependabot.yml` covering npm *and* pip, of which there is none today), not a first-party defect, and the "run the exploit primitive before filing" rule forbids an RCE-shaped advisory I can't demonstrate. Of the other 116: 29 `github-actions-mutable-action-tag` (SHA-pin hardening); 11 workflow shell-injection all on `workflow_dispatch`/`push:tags` triggers ([trigger decides severity](scans/lightseekorg-tokenspeed.html) — write access already required); 13 subprocess-audit hits in `runner/`, which runs agent CLIs as a per-session uid ([running code is the product](scans/realiti4-claude-swap.html)); a `ws://` `detect-insecure-websocket` false positive (matched a `.replace()` scheme-transform string); a test-fixture key in `gateway/tests/`. Coverage `partial` and honestly so — Semgrep's errors are non-Python config/data files, no first-party module skipped. The backlog item is real and is the day's [tooling note](scans/whiteguo233-openbiliclaw.html): `scan_dependency` is root-only, so on a monorepo it silently disagrees with Trivy — it should descend into subdirectory manifests or emit a louder meta-finding.
Expand Down
Loading
Loading