diff --git a/docs/fixed.md b/docs/fixed.md index fcb5a7b..97e2aa6 100644 --- a/docs/fixed.md +++ b/docs/fixed.md @@ -1,12 +1,12 @@ --- layout: default title: Findings a maintainer fixed -description: "The 25 scans in the AI PatchLab series where the maintainer shipped a fix — what was reported, and what landed upstream." +description: "The 26 scans in the AI PatchLab series where the maintainer shipped a fix — what was reported, and what landed upstream." --- # Findings a maintainer fixed -Of 109 scans, **25** ended with a maintainer shipping a fix. This page is the +Of 111 scans, **26** ended with a maintainer shipping a fix. This page is the short version of the argument: a report is only worth writing if someone can act on it. The fastest turnaround in the series was about six hours from filing to a merged pull @@ -20,6 +20,7 @@ release notes. | 2026-09-14 | [superlinked/sie](scans/superlinked-sie.html) | 156 | 1 real | | 2026-09-03 | [samuelgursky/davinci-resolve-mcp](scans/samuelgursky-davinci-resolve-mcp.html) | 97 | 1 real | | 2026-08-31 | [shy3130/tick-stock-panel](scans/shy3130-tick-stock-panel.html) | 77 | 1 real | +| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | | 2026-08-27 | [Zleap-AI/SAG](scans/zleap-ai-sag.html) | 60 | 1 real | | 2026-08-20 | [whiteguo233/OpenBiliClaw](scans/whiteguo233-openbiliclaw.html) | 373 | 1 real | | 2026-08-17 | [zilliztech/memsearch](scans/zilliztech-memsearch.html) | 90 | 1 real | diff --git a/docs/index.md b/docs/index.md index 9c2ec4a..65febd8 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,7 +1,7 @@ --- layout: default title: AI PatchLab Scans -description: "111 curated security scans of open-source AI agents, MCP servers and LLM apps - 25 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit." +description: "111 curated security scans of open-source AI agents, MCP servers and LLM apps - 26 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit." --- # AI PatchLab Scans @@ -20,7 +20,7 @@ remediation and confidence rules to normalize the findings. > **Want this run privately against your own codebase?** I do independent > security review of AI agents, MCP servers, and LLM apps — -> [**work with me →**]({{ '/work-with-me' | relative_url }}). 111 scans, 25 confirmed fixes, methodology in the open. +> [**work with me →**]({{ '/work-with-me' | relative_url }}). 111 scans, 26 confirmed fixes, methodology in the open. > **OpenAI just launched [Daybreak](https://openai.com/index/daybreak-securing-the-world/) and Patch the Planet.** > Same remediation loop, opposite trade-off: their path is a cloud frontier model; @@ -35,7 +35,7 @@ remediation and confidence rules to normalize the findings. ## Two shorter ways in -- [**Findings a maintainer fixed**]({{ '/fixed' | relative_url }}) — the 25 that resolved +- [**Findings a maintainer fixed**]({{ '/fixed' | relative_url }}) — the 26 that resolved upstream. The shortest version of the argument: a report is only worth writing if someone can act on it. - [**Scans that found nothing**]({{ '/clean' | relative_url }}) — 40 of them, published as @@ -141,7 +141,7 @@ filed, which is the usual outcome of a clean scan. | 2026-09-01 | [future-agi/future-agi](scans/future-agi-future-agi.html) | 1227 | 1 real — withheld | private | | 2026-08-31 | [shy3130/tick-stock-panel](scans/shy3130-tick-stock-panel.html) | 77 | 1 real | **fixed** | | 2026-08-30 | [SenteLabsAI/OpenExecutive](scans/sentelabsai-openexecutive.html) | 124 | 1 real — withheld | private | -| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | open | +| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | **fixed** | | 2026-08-28 | [Ontos-AI/knowhere](scans/ontos-ai-knowhere.html) | 129 | 1 real — withheld | private | | 2026-08-27 | [Zleap-AI/SAG](scans/zleap-ai-sag.html) | 60 | 1 real | **fixed** | | 2026-08-26 | [ascending-llc/jarvis-registry](scans/ascending-llc-jarvis-registry.html) | 235 | 1 real — withheld | private | diff --git a/docs/scan-log.md b/docs/scan-log.md index 108d98b..f0e1638 100644 --- a/docs/scan-log.md +++ b/docs/scan-log.md @@ -8,7 +8,7 @@ description: "The complete AI PatchLab scan log: every public repository scanned Every scan in the series, newest first, with the summary written on the day of the scan. 111 scans. For the compact index, see the [scan log home]({{ '/' | relative_url }}). -- **2026-09-24** — [superdesigndev/treg](scans/superdesigndev-treg.html) — 96 findings at `medium+` (1 critical, 23 high, 69 medium, 3 meta), **1 real — withheld, reported privately by email** — **"OpenRouter for agent tools"** (3k★, source-available; hosted at treg.to and self-hostable), a multi-tenant server holding a team's provider credentials Fernet-encrypted and proxying calls to 3,000+ upstream endpoints with the credential injected server-side so it never reaches the caller. Its `SECURITY.md` makes two load-bearing promises: the proxy never hands a key to the caller, and outbound requests are *"re-resolved at call time and internal/metadata addresses are refused."* **Zero of the 96 scanner findings is the vulnerability** — the real item came from tabulating the server's outbound request paths and asking which the call-time host guard actually covers. It does have a proper guard, `host_is_public` in `infra/upstream/ssrf.py`, and its address classification is genuinely thorough (every loopback/link-local/CGNAT/NAT64 and decimal/hex/octal/short-form IPv4 literal I tried was refused). The finding is that the guard is wired into the main proxy path and **not** into a sibling outbound path a plain member can drive; the registration-time check that does run there is, by its own docstring, a static check that allows any DNS name, so a name resolving to an internal address passes it and nothing re-resolves before the request goes out — a member-authenticated, semi-blind [SSRF on the unguarded sibling of a guarded transport](scans/liaohch3-claude-tap.html). Confirmed with the [run-the-primitive](scans/observal-observal.html) discipline: treg's own two guard functions run side by side on a host that resolves internally return "allowed" (registration) and "refused" (call-time), and the sibling path calls only the first — one differential, using the project's code not my reasoning. Graded **Medium**: the injected credential on the affected path is the member's own, so it is internal + cloud-metadata reach, not cross-tenant theft; the org boundary itself held everywhere ([the probe could return "no" on isolation and did](scans/tracecathq-tracecat.html)). **The fix is one guard applied at the outbound path** and changes nothing for a legitimate public target. The 96: seventeen `sqlalchemy-text` are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) (Alembic `SET`/`CREATE INDEX CONCURRENTLY` from module constants + a test-reset `TRUNCATE` over dialect-quoted table names, no value interpolated); eighteen `v-html` render the project's own static tutorial/snippet HTML with an `esc(&<>)` helper on every interpolated field; four `insecure-file-permissions` are the project [tightening to `0o700`](scans/realiti4-claude-swap.html) (the rule's suggestion would loosen them); one Trivy critical is `anyio` 4.14.1→4.14.2 IDNA hostname-confusion, a worthwhile lockfile bump reachable only by a network-positioned attacker against a non-ASCII upstream host. Coverage `partial` and stated so — Semgrep hit a Windows per-file temp error on five large modules (recovered by a targeted re-run of just those five), and pip-audit left `uv.lock` unaudited while Trivy read it; neither touches the finding. **Strict-norm** (`SECURITY.md` forbids public vuln issues, PVR disabled — checked `{"enabled":false}` — names an email) · reported **privately by email to the maintainer**; post-only here, finding withheld at class level, **the send is the operator's step** +- **2026-09-24** — [superdesigndev/treg](scans/superdesigndev-treg.html) — 96 findings at `medium+` (1 critical, 23 high, 69 medium, 3 meta), **1 real — withheld, reported privately by email** — **"OpenRouter for agent tools"** (3k★, source-available; hosted at treg.to and self-hostable), a multi-tenant server holding a team's provider credentials Fernet-encrypted and proxying calls to 3,000+ upstream endpoints with the credential injected server-side so it never reaches the caller. Its `SECURITY.md` makes two load-bearing promises: the proxy never hands a key to the caller, and outbound requests are *"re-resolved at call time and internal/metadata addresses are refused."* **Zero of the 96 scanner findings is the vulnerability** — the real item came from tabulating the server's outbound request paths and asking which the call-time host guard actually covers. It does have a proper guard, `host_is_public` in `infra/upstream/ssrf.py`, and its address classification is genuinely thorough (every loopback/link-local/CGNAT/NAT64 and decimal/hex/octal/short-form IPv4 literal I tried was refused). The finding is that the guard is wired into the main proxy path and **not** into a sibling outbound path a plain member can drive; the registration-time check that does run there is, by its own docstring, a static check that allows any DNS name, so a name resolving to an internal address passes it and nothing re-resolves before the request goes out — a member-authenticated, semi-blind [SSRF on the unguarded sibling of a guarded transport](scans/liaohch3-claude-tap.html). Confirmed with the [run-the-primitive](scans/observal-observal.html) discipline: treg's own two guard functions run side by side on a host that resolves internally return "allowed" (registration) and "refused" (call-time), and the sibling path calls only the first — one differential, using the project's code not my reasoning. Graded **Medium**: the injected credential on the affected path is the member's own, so it is internal + cloud-metadata reach, not cross-tenant theft; the org boundary itself held everywhere ([the probe could return "no" on isolation and did](scans/tracecathq-tracecat.html)). **The fix is one guard applied at the outbound path** and changes nothing for a legitimate public target. The 96: seventeen `sqlalchemy-text` are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) (Alembic `SET`/`CREATE INDEX CONCURRENTLY` from module constants + a test-reset `TRUNCATE` over dialect-quoted table names, no value interpolated); eighteen `v-html` render the project's own static tutorial/snippet HTML with an `esc(&<>)` helper on every interpolated field; four `insecure-file-permissions` are the project [tightening to `0o700`](scans/realiti4-claude-swap.html) (the rule's suggestion would loosen them); one Trivy critical is `anyio` 4.14.1→4.14.2 IDNA hostname-confusion, a worthwhile lockfile bump reachable only by a network-positioned attacker against a non-ASCII upstream host. Coverage `partial` and stated so — Semgrep hit a Windows per-file temp error on five large modules (recovered by a targeted re-run of just those five), and pip-audit left `uv.lock` unaudited while Trivy read it; neither touches the finding. **Strict-norm** (`SECURITY.md` forbids public vuln issues, PVR disabled — checked `{"enabled":false}` — names an email) · reported **privately by email to the maintainer**; post-only here, finding withheld at class level · sent by email 2026-09-24, ten minutes after this entry was published - **2026-09-23** — [can4hou6joeng4/boss-agent-cli](scans/can4hou6joeng4-boss-agent-cli.html) — 26 findings at `medium+` (7 high, 16 medium, 3 meta), **1 real — withheld, filed privately** — a **CLI for the BOSS Zhipin recruitment platform** (2k★, MIT; a human drives it from a terminal wizard, an agent drives the same workflows over a JSON-envelope API, an ~77-tool MCP server, or a Python SDK) that authenticates by reusing the operator's own logged-in browser session. Responsive maintainer — 67 merged PRs from 7 human authors in 60 days — which is why it cleared the pre-check. **Zero of the 26 scanner findings is the vulnerability**, which is the whole point of the entry: the real item is the *absence* of a check on an otherwise-working local channel, the shape no rule can match, and it is the [desktop-loopback-inversion](scans/liaohch3-claude-tap.html) class — a `127.0.0.1` control surface that binds neither *who* may call it nor which host the request claims, driving privileged browser actions against a site where the operator is authenticated. Confirmed with a **server-side exploit primitive against the shipped daemon** (a foreign `Host` is served, a cross-origin "simple" request reaches the command channel with no preflight, a hostile `Origin` is not rejected) plus a read of the extension source for the browser half — the [run-the-primitive discipline](scans/observal-observal.html), because reading alone could not separate "the channel is drivable" from "the channel is drivable *and* blind does not blunt it." Graded **High not Critical** on an honest precondition (the bridge component must be running, and the user has to start it deliberately; once started it stays up while the browser side is connected — *corrected 2026-09-24: first published as "a bounded window", the one claim reasoned about rather than checked, and wrong both ways; the private report was corrected the same day*) and the sensitive scope is the one site the tool is built around; **the fix was written before it was named**, narrow and verified not to break the tool's own two legitimate clients ([write the patch first](scans/roflcoopter-viseron.html)). The 26: eight SQL hits (five high) are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) — every interpolation is a hardcoded table/column name with values bound as `?`, settled by reading the call sites; eight of sixteen mediums are one `github-actions-mutable-action-tag` flood; the rest are two `sha1()[:16]` dedup-id digests, two `urllib` calls to a local DevTools endpoint and a build script, a React rule on a loopback health-check `fetch`, and a Python-3.7 compatibility rule that is not a security check. Coverage `partial` and stated so — Semgrep did not cover every file and the dependency scan left `uv.lock` unaudited (Trivy read it, the [luck that hides the bug](scans/roflcoopter-viseron.html)), neither touching the finding. **Strict-norm** (`SECURITY.md` forbids public vulnerability issues, gives an email and a PVR link) · PVR-enabled, filed **privately, accepted into triage first try** ([GHSA-xp54-2hxc-w78q](https://github.com/can4hou6joeng4/boss-agent-cli/security/advisories/GHSA-xp54-2hxc-w78q)) **with the `vulnerabilities` array** — the [viseron rule](scans/roflcoopter-viseron.html) holds again · post-only, finding withheld at class level under embargo - **2026-09-22** — [overwirehq/claude-code-telegram](scans/overwirehq-claude-code-telegram.html) — 54 findings at `medium+` (1 critical, 21 high, 30 medium), **0 first-party defects — dependency currency, post-only** — a **Telegram bot that gives authorised users remote access to Claude Code**, i.e. it runs file and Bash tools on a host machine by design, so the whole security surface *is* the boundary around that execution (2k★, MIT, org-backed; active — 15 merged PRs from 7 authors and 9 closed issues in 60 days; strict-norm: real `SECURITY.md`, PVR enabled). Picked with the manual queue at zero. **The story is what precise security documentation looks like.** The real boundary is the `can_use_tool` callback that stops Claude — when steered off-course by content it reads mid-task — from acting outside the approved directory, and the maintainers know and *state* exactly what it covers: path validation is scoped in the docs to precisely six tools (`Read`, `Write`, `Edit`, `MultiEdit`, `NotebookEdit`, `NotebookRead`), while `Grep`/`Glob`/`LS` are deliberately left to the OS sandbox. Reading that as "Read is guarded but Grep is not, so the boundary is incomplete" is the plausible-but-wrong finding this site exists to resist — the boundary is [advertised, not accidental](scans/tracecathq-tracecat.html), and `ROADMAP-v2.md` even records the SDK's own `CanUseToolShadowedWarning` naming every tool the callback will never see, filed as tracking item #221. Two more choices earn credit: guarded tools are deliberately *stripped from* the SDK `allowed_tools` list (a pre-approved tool never produces a `can_use_tool` request, so leaving them in would render the checks silently inert — [issue #219](https://github.com/overwirehq/claude-code-telegram/issues/219)), and `autoAllowBashIfSandboxed` is disabled whenever the boundary checks must run, closing a second bypass. That is a maintainer who traced how the framework resolves permissions rather than trusting that a config allow-list is an enforcement boundary — it isn't, and the code says so. **When the machine is built and documented this carefully, the finding moves to the dependency manifest.** The 30 Trivy advisories in `poetry.lock` split cleanly by reachability: **17 are opt-in-only and not reachable on a default install** — `starlette` (6) and `python-multipart` (3) need the FastAPI webhook server (`ENABLE_API_SERVER=false` default); `pyjwt` (5) needs token auth (`ENABLE_TOKEN_AUTH=false` default, and the SECURITY.md documents token auth as non-functional, [#58](https://github.com/overwirehq/claude-code-telegram/issues/58)); the MCP SDK highs (3) need MCP enabled — while **13 are unconditionally installed** (`anyio` incl. the critical IDNA/TLS advisory, the `cryptography`/OpenSSL cluster, `urllib3`, `idna`, `requests`, `pydantic-settings`, `python-dotenv`) and are the genuine currency gap. The direct deps are current; the drift is transitive, which the repo's existing `.github/dependabot.yml` (configured for *version* updates, not *security* updates) will not raise on its own. **No advisory filed** — there is no first-party vulnerability. Of the other 24: 16 `github-actions-mutable-action-tag` (SHA-pin hardening); one `pull_request_target` checkout wrapped in a 40-line threat-model header (read-only tool allowlist, secrets scrubbed, "worst case is a prompt-injected review comment" — [the trigger decides severity](scans/lightseekorg-tokenspeed.html)); two `sqlalchemy-execute-raw-query` [identifier FPs](scans/aurelio-labs-semantic-router.html) (only a generated run of `?` placeholders is interpolated, values bound via `execute(query, params)`); and three gitleaks hits that are a `your-api-secret` placeholder plus two fake tokens in a test asserting the project's own `_redact_secrets()` helper scrubs them ([credited defence](scans/realiti4-claude-swap.html)). Coverage `partial` and honestly so: pip-audit resolved **no dependencies** from a `pyproject.toml` declaring `dynamic = ["dependencies"]` — reading identically to "clean" — and Trivy's `poetry.lock` read carried the run, the [same dependency blind spot](scans/harnessrouter-harnessrouter.html) as yesterday reached by a different manifest shape. @@ -34,7 +34,7 @@ Every scan in the series, newest first, with the summary written on the day of t - **2026-08-28** — [Ontos-AI/knowhere](scans/ontos-ai-knowhere.html) — 129 findings, **1 real — withheld** — a document-ingestion pipeline sold as the memory layer between dirty documents and AI agents (2.7k★): one missing-authentication gap on an internal storage-event webhook whose verifiers are stubs, verified by execution (wrong token rejected, absent header accepted), and a 68-CVE dependency table that all resolves to a single lockfile · drafted for private delivery, post-only - **2026-08-27** — [Zleap-AI/SAG](scans/zleap-ai-sag.html) — 60 findings, **1 real** — an agent framework whose login did not verify the password: any name signed in as the owner account · ✅ **Resolved 2026-08-30 in 3 days (19th fix in the series)** — posing the product question beat proposing a patch - **2026-09-01** — [future-agi/future-agi](scans/future-agi-future-agi.html) — 1227 findings (1224 above the medium floor), **1 real — withheld** — an **end-to-end LLM/agent observability and evaluation platform** (1.9k★, Apache-2.0 + EE): OTLP trace ingestion into a ClickHouse store of prompts/completions/tokens/costs, an evaluation engine, a simulation runner, a Go **AI gateway** (`agentcc-gateway`, an OpenAI/Anthropic/Gemini-compatible proxy with per-org keys, rotation and RBAC) and a Django backend, with a real `SECURITY.md` (named email, SLA table, safe harbor, in-scope list, public vuln issues forbidden). **Largest raw count in the series — and one real finding no rule represented.** The class: **one service in a multi-service deployment is published on all interfaces while every service beside it, including every datastore, is correctly loopback-bound, and the production overlay inherits the exposure unchanged** — the exposed service is the most dangerous one in the file to leave open, and reaching it hands an unauthenticated caller a path to the trace store. Three reportability properties, each a series pattern: **the maintainer's own compose comment is the oracle** (*"bound to 127.0.0.1 so only the host can reach them"*, applied to Postgres/ClickHouse/Redis/MinIO/Temporal/collector — the finding is the outlier), the [contract-vs-artifact](scans/vexa-ai-vexa.html) footing in a deployment file; **N deployment paths, and the question is whether they agree** — the production overlay re-binds every secret with `${VAR:?must be set for production}` (genuinely good) but names only `environment:`/`restart:`, not `ports:`, so the documented production command inherits the exposure; and **I proved it with the vendor's own merge engine** — rendered the exact `deploy/README.md` command through `docker compose config` and read the published bindings out of the result (exposed service on `0.0.0.0`, every datastore on `127.0.0.1`), [running the primitive](scans/nottelabs-notte.html) rather than reasoning about override semantics. **Scope discipline kept it to one**: two other `0.0.0.0` services are the intended public surface (frontend + reverse-proxied API), and two management UIs published without a loopback bind sit behind Compose `profiles:` — opt-in, absent from the default and documented-production deployments, and the rendered config confirmed they don't appear. **Two auth hypotheses returned NO and that is part of the record**: the gateway mounts a large admin plane *outside* its `/v1/` auth prefix (mint keys, set org configs, rotate creds) — the [claude-tap inversion](scans/liaohch3-claude-tap.html) shape — but all **24** admin handlers across four files call an in-handler `requireAdmin`/`checkAdminAuth` with a constant-time compare that **fails closed on an unset token**; and the `/v1/` middleware short-circuits on a valid license token (a [caller-influenced auth branch](scans/liaohch3-claude-tap.html)) but the verifier pins RS256, resolves `kid` from a **local** key map, and validates the full claim set, so branch-then-serve is safe. **Noise**: 100 mutable-action-tag GHA rows, 94+44 raw/formatted-SQL that resolve to the [#1 identifier FP](scans/mnemosyne-oss-mnemosyne.html) (ClickHouse builders interpolate regex-validated keys and an allowlisted bucket function, bind every value via `%(param)s` — the four taint-tracked hits were audit-log f-strings and the allowlisted bucket fn), 3 gitleaks hits all covered by the repo's own `.gitleaks.toml`. **The 26 criticals are a coverage story**: `scan_dependency` found *no root manifest* because this is a **monorepo** (manifests under `futureagi/`), so the ChromaDB/LiteLLM/Authlib/Django/langchain criticals came from Trivy's lockfile parse and want the *version-match → reachable → mitigated* gate — several are Proxy-only/transport-only in libs used as clients. **Monorepo root-only coverage gap, re-confirmed as the top backlog item.** Strict-norm · reported **privately** to `security@futureagi.com` with a full dossier + `docker compose config` reproduction · post-only, finding **withheld** (class only) until remediated · the email send is the operator's manual step -- **2026-08-29** — [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) — 372 findings above the medium floor, **1 real** — a **financial-market agent platform** (1.7k★, Apache-2.0, daily merges, FastAPI + React + a Daytona/Docker code sandbox), and one of the most carefully defended codebases in the series. Semgrep's top finding — `pull_request_target` checking out untrusted fork code — **retired on evidence outside the repository tree**: the GitHub environments API shows a `fork-ci` environment with required reviewers and `prevent_self_review`, and the workflow pins the immutable head SHA, sets `persist-credentials: false`, and drops the token to `contents: read`. The author even documented that `uv sync` runs fork build hooks before pytest — a better threat model than the rule that flagged it. All 30 gitleaks hits are fixtures in the project's **own** secret-redaction and leak-detection tests; the 102-finding SQL cluster binds every value with `%s` and interpolates only a constant column list. What survived is a **composite**: `GET /api/v1/preview/{workspace_id}/{port}` is unauthenticated and its docstring promises it "does NOT start stopped sandboxes (to prevent denial-of-wallet)" — but it enforces that with a DB `status == "running"` check and then calls the *acquisition* path, which reaches `PTCSandbox.reconnect`, documented as "it **starts a stopped sandbox**." **Four sibling unauthenticated routes in the same codebase refuse exactly that call**, citing "a stale 'running' DB row with no warm session," and use the no-wake `get_session_if_ready` accessor whose fence parameter was made mandatory "so a new caller cannot omit the fence." The majority sibling is the contract; the fifth route never inherited it. The lead came from the tool's **coverage warning**, not its findings: Semgrep could not parse `claude.yml`, `release.yml`, or `sandbox-integration.yml`, which were precisely the workflows worth reading. +- **2026-08-29** — [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) — 372 findings above the medium floor, **1 real** — a **financial-market agent platform** (1.7k★, Apache-2.0, daily merges, FastAPI + React + a Daytona/Docker code sandbox), and one of the most carefully defended codebases in the series. Semgrep's top finding — `pull_request_target` checking out untrusted fork code — **retired on evidence outside the repository tree**: the GitHub environments API shows a `fork-ci` environment with required reviewers and `prevent_self_review`, and the workflow pins the immutable head SHA, sets `persist-credentials: false`, and drops the token to `contents: read`. The author even documented that `uv sync` runs fork build hooks before pytest — a better threat model than the rule that flagged it. All 30 gitleaks hits are fixtures in the project's **own** secret-redaction and leak-detection tests; the 102-finding SQL cluster binds every value with `%s` and interpolates only a constant column list. What survived is a **composite**: `GET /api/v1/preview/{workspace_id}/{port}` is unauthenticated and its docstring promises it "does NOT start stopped sandboxes (to prevent denial-of-wallet)" — but it enforces that with a DB `status == "running"` check and then calls the *acquisition* path, which reaches `PTCSandbox.reconnect`, documented as "it **starts a stopped sandbox**." **Four sibling unauthenticated routes in the same codebase refuse exactly that call**, citing "a stale 'running' DB row with no warm session," and use the no-wake `get_session_if_ready` accessor whose fence parameter was made mandatory "so a new caller cannot omit the fence." The majority sibling is the contract; the fifth route never inherited it. The lead came from the tool's **coverage warning**, not its findings: Semgrep could not parse `claude.yml`, `release.yml`, or `sandbox-integration.yml`, which were precisely the workflows worth reading. · one focused public issue filed ([#378](https://github.com/ginlix-ai/LangAlpha/issues/378)), PVR off and no `SECURITY.md` anywhere · ✅ **Resolved 2026-09-25 in 27 days (26th fix in the series)** — [#425](https://github.com/ginlix-ai/LangAlpha/pull/425) did not take the proposed accessor swap; it removed the premise. The old preview route now reads the database and redirects to the app's owner-only link, so it touches no session and can wake nothing, and the workspace UUID stopped being a credential altogether: report frames load under an HMAC-signed grant that expires. A regression test names #378. **Re-verified at the fix commit**: the differential that found it went from one waking unauthenticated route out of five to none, and the new redirect was probed as a new check — the link code it hands a UUID holder opens nothing, because the schema itself forbids sharing an app link (a `CHECK` constraint requires `shared_at IS NULL`). - **2026-08-26** — [ascending-llc/jarvis-registry](scans/ascending-llc-jarvis-registry.html) — 235 findings above the medium floor, **1 real — withheld** — an **enterprise MCP/A2A gateway** that brokers per-user OAuth credentials to downstream tool servers (2.8k★, Apache-2.0, commercial backing, **7 distinct human PR authors in 60 days**), and the cleanest demonstration yet of why the mitigation gate is not optional. All four criticals retired: the **LiteLLM CVEs** (SQL injection, Host-header auth bypass, MCP command execution, key-gen privilege escalation) are Proxy-server-only, and this repo declares `litellm>=1.50.0` then **never imports it** — one grep outside the lockfile returns the dependency declaration and nothing else. The largest family, **21× GitHub Actions shell injection**, retired on a single fact: there is no `pull_request_target` anywhere in the repository, so every flagged workflow runs either from a fork with a read-only token and no secrets, or from a context that already requires write access. **16 gitleaks "secrets"** are docs example output and MongoDB seed fixtures, against a `tartufo.toml` that carries a written reason on every exclusion; **12× `logger-credential-leak`** logs usernames and source ids, never a credential; the lone **`unverified-jwt-decode`** is by design and its two callers both verify-then-branch, over a token layer that pins RS256, pins the `kid`, enforces issuer and audience, and blocks token-class confusion with positive equality checks. **The near-miss is the story**: the registry allows CORS origins by a regex whose `.*\.compute.*\.amazonaws\.com` arm admits the public DNS name of *every EC2 instance on the internet*, with `allow_credentials=True` — confirmed allowed against real Starlette 1.6.0, with the controls correctly denied. It is still **inert**, because every cookie the app sets is `SameSite=Lax` and `amazonaws.com` compute domains are on the Public Suffix List, so the attacker's `fetch()` carries no session at all. Four verification steps said "real"; the fifth — grep for the mitigation before flagging — retired it, and this series has filed that same class as a genuine finding against five other projects. **The one real defect is object-level authorization no scanner can see**: two sibling handlers in one router return per-user state from the same deterministic identifier, one enforces ownership against the caller and the other takes no caller identity at all, behind a permission every role holds. The intra-repo differential is the proof — the guarded sibling *is* the project's own contract. Routed privately to `security@ascendingdc.com` per the project's SECURITY.md, with a patch; public detail withheld. - **2026-08-25** — [langflow-ai/openrag](scans/langflow-ai-openrag.html) — 213 findings above the medium floor, **1 real — withheld** — a **single-package Retrieval-Augmented Generation platform** built on OpenSearch + Langflow (4.5k★, Apache-2.0, commercial backing, **47 merged PRs from 13 authors in 60 days**), and a textbook "the scanner count is the least informative thing about this repo" target. 213 medium-plus findings resolve to one recurring best-practice class (**121× GitHub Actions mutable action tags**) and a long tail of well-understood false positives: **15 gitleaks "secrets"** all in `.secrets.baseline` / empty `.env.example` slots / `Makefile` shell variables / docs / a test fixture (credit the `.secrets.baseline` and they collapse to zero), a **"critical" Kubernetes RBAC** that is the OpenRAG operator's *own* secret-management role doing its job, and **10× `logger-credential-leak`** on metadata-only log lines — the series' most reliable false positive, still awaiting its confidence downgrade. **The one finding that matters is a composite no scanner can see**: a hard-coded/shipped-default cryptographic secret on a token-signing boundary (CWE-321/798), reachable in the default self-hosted deployment, where the hardening applied to one token path was not applied to its sibling. It was found by diffing `securityconfig/` against `cloud_securityconfig/` and tracing a token from mint to validation — neither of which any rule performs — verified offline against OpenRAG's own token-validation logic, and reported privately via GitHub PVR (**GHSA-xv8v-6c28-v78p**, `triage`) with a concrete fix written against the repo's own key-generation architecture. Detail withheld per OpenRAG's SECURITY.md; this entry carries the class only. diff --git a/docs/scans/ginlix-ai-langalpha.md b/docs/scans/ginlix-ai-langalpha.md index ad5dd2d..4674946 100644 --- a/docs/scans/ginlix-ai-langalpha.md +++ b/docs/scans/ginlix-ai-langalpha.md @@ -9,12 +9,16 @@ date: 2026-08-29 **Repository:** [ginlix-ai/LangAlpha](https://github.com/ginlix-ai/LangAlpha) **Commit scanned:** `f5232fa2` (main at scan time) **Scan date:** 2026-08-29 -**Disclosure status:** disclosed — one real finding filed as a single focused -public issue. No `SECURITY.md` at the repository root, in `.github/`, in -`docs/`, or at the organisation level, and private vulnerability reporting is -disabled (confirmed with an empty-payload control request, which returned -`403 Repository does not have private vulnerability reporting enabled` — it -files nothing). A public issue is the only channel the project offers. +**Disclosure status:** ✅ **resolved** — one real finding, filed as a single +focused public issue on the day of the scan and fixed by the maintainer 27 days +later in [ginlix-ai/LangAlpha#425](https://github.com/ginlix-ai/LangAlpha/pull/425), +which closed the issue as completed. The fix went further than the one proposed +here: it retired the design the finding took as given. No `SECURITY.md` at the +repository root, in `.github/`, in `docs/`, or at the organisation level, and +private vulnerability reporting is disabled (confirmed with an empty-payload +control request, which returned `403 Repository does not have private +vulnerability reporting enabled` — it files nothing). A public issue is the +only channel the project offers. ## Summary @@ -181,6 +185,35 @@ high-severity findings automatically instead of by hand. [issue #378](https://github.com/ginlix-ai/LangAlpha/issues/378). One finding, not a grouped review. No PR opened: the fix is one file, but choosing 404 vs 503 for a sleeping sandbox is a product decision the maintainer should make. +- **2026-09-25** — **fixed** in [#425](https://github.com/ginlix-ai/LangAlpha/pull/425) + (`2eab033f`) and the issue closed as completed, 27 days after filing. The + maintainer did not adopt the accessor swap proposed above; the fix removes + the premise instead. The old `/api/v1/preview/{workspace_id}/{port}` route no + longer resolves a signed URL or touches a session at all: it reads the + registered preview command from the database, gets or creates the app's one + link, and redirects to it. That link opens for the signed-in workspace owner + alone, and the owner's path is now the only one that resolves a preview — in + the new docstring's words, "a stopped sandbox is started for the owner and for + nobody else." This page described the workspace UUID as the route's bearer + credential, by design. #425 retired that design: report frames now load under + an HMAC-signed grant that expires, and running apps open through the owner's + private link. A regression test names #378 — with the workspace manager and + the sandbox lookup both patched to raise, the redirect still answers `302` + and the manager is never called. +- **2026-09-26** — re-verified at `2eab033f`, starting with the differential + that found the bug. At the scanned commit, five unauthenticated routes touched + a workspace session and one called the waking accessor; at the fix, none do, + and every remaining caller of `get_session_for_workspace` sits behind an + authenticated owner. The redirect was then probed as a new check, because it + still answers anyone who holds a UUID. What it hands back is an app link + code, and an app link cannot be shared — the migration's `share_links_app_shape` + constraint requires `shared_at IS NULL` on every app row — so for anyone but + the owner the code opens nothing and returns the same `404` as an unknown one. + The public page that does resolve a preview reaches the waking path only past + two independent ownership checks (`resolve_link`, then `require_workspace_owner` + inside `_get_sandbox`). Nor can the redirect be used to rewrite the owner's + link: its insert is `ON CONFLICT DO NOTHING`, and an old URL's page rides along + as `?path=` rather than being stored. ## Reproduce diff --git a/docs/scans/superdesigndev-treg.md b/docs/scans/superdesigndev-treg.md index 2c007c4..427236e 100644 --- a/docs/scans/superdesigndev-treg.md +++ b/docs/scans/superdesigndev-treg.md @@ -14,8 +14,9 @@ date: 2026-09-24 project's `SECURITY.md` (which asks that vulnerabilities not be opened as public issues; GitHub private vulnerability reporting is disabled on the repo, and the policy names an email address, so email is the channel). Detail below is kept at -class level until the maintainer has had a chance to respond. **The private send -is the operator's step and had not gone out when this page was published.** +class level until the maintainer has had a chance to respond. The private send +was the operator's step and had not gone out when this page was published; it +went out ten minutes later (see the timeline below). ## Summary @@ -161,6 +162,17 @@ curated by hand; scanner output alone is not a vulnerability report. This page w updated with full technical detail once the maintainer has responded and a fix has shipped.* +## Disclosure timeline + +- 2026-09-24 — scan run at `240c595`; finding verified by running treg's own two host + guards side by side +- 2026-09-24 — channel probed: `SECURITY.md` asks that vulnerabilities not be opened as + public issues and names an email address; GitHub private vulnerability reporting is + disabled (`{"enabled":false}`) +- 2026-09-24 13:31 UTC — this page published, finding withheld at class level +- 2026-09-24 13:41 UTC — reported privately by email to the address in `SECURITY.md` +- 2026-09-26 — no reply yet, no bounce; none of the reported files has changed upstream + ## Reproduce ```bash diff --git a/docs/work-with-me.md b/docs/work-with-me.md index 53a9a29..090810e 100644 --- a/docs/work-with-me.md +++ b/docs/work-with-me.md @@ -1,7 +1,7 @@ --- layout: default title: Work with me — AI agent & LLM security review -description: "Independent security review of AI agents, MCP servers and LLM applications. 109 public scans and 25 confirmed fixes as the proof; first-look engagements from $750." +description: "Independent security review of AI agents, MCP servers and LLM applications. 111 public scans and 26 confirmed fixes as the proof; first-look engagements from $750." --- # Work with me @@ -12,7 +12,7 @@ If you're shipping an agent framework, an MCP server, a RAG pipeline, or anythin ## The proof is public -Everything I'd do for you, I've done in the open — [109 curated scans]({{ '/' | relative_url }}) of well-known AI/agent projects, with **25 confirmed fixes** where maintainers acted on the findings, and a documented methodology behind each one. +Everything I'd do for you, I've done in the open — [111 curated scans]({{ '/' | relative_url }}) of well-known AI/agent projects, with **26 confirmed fixes** where maintainers acted on the findings, and a documented methodology behind each one. - **Real findings, not scanner noise.** A typical scan starts at hundreds of raw findings and ends at a handful that matter — because the rest are false positives, by-design patterns, or vendored third-party code. The curation is the value. - **Reachability over rule-count.** A `subprocess(shell=True)` that's safe under operator trust can become a multi-tenant sandbox escape once you trace who reaches it — and a scary-looking SQL string can be fully gated by a five-character allowlist. The verdict goes wherever the *code* goes, which means reading the code, not the rule.