From 4322651a5b4c9db60c4c4983b7782be8a665de7e Mon Sep 17 00:00:00 2001 From: elfrost <5491654+elfrost@users.noreply.github.com> Date: Sat, 26 Sep 2026 09:09:26 -0400 Subject: [PATCH] docs: LangAlpha #378 fixed in #425 (26th fix); treg report delivered LangAlpha: the maintainer closed #378 as completed on 2026-09-25, 27 days after filing, with #425 (2eab033f). The fix is not the accessor swap the issue proposed. The old /api/v1/preview/{workspace}/{port} route no longer resolves a signed URL or touches a session: it reads the registered preview command, gets or creates the app's one link, and redirects to it. That link opens for the signed-in owner alone, and the workspace UUID stopped being a credential altogether (report frames load under an expiring HMAC grant). A regression test names #378. Re-verified at 2eab033f, differential first: five unauthenticated routes touched a workspace session at the scanned commit and one called the waking accessor; at the fix, none do, and every remaining get_session_for_workspace caller is behind an authenticated owner. The redirect was probed as a new check, since it still answers any UUID holder: the code it returns is an app link, and the share_links_app_shape CHECK constraint makes an app link unshareable, so the code opens nothing for anyone but the owner. The public metadata route reaches the waking path only past two independent ownership checks. The redirect cannot rewrite the owner's link (ON CONFLICT DO NOTHING; the old path rides along as ?path=). - scan post: status resolved, two timeline entries (fix, re-verification) - index: outcome open -> fixed; confirmed-fix counters 25 -> 26 - fixed.md: LangAlpha row added; counters had drifted at 109 scans - scan log: resolution suffix on the 2026-08-29 entry - work-with-me: counters had drifted at 109 scans; now 111 / 26 treg: the post said the private send "had not gone out when this page was published". True at 13:31 UTC; Gmail Sent shows it went to the SECURITY.md address at 13:41 UTC the same day. Added a disclosure timeline and reworded the status line and the scan-log entry so the page no longer reads as if the maintainer was never told. The outcome is unchanged: no reply yet, no bounce, and none of the reported files has changed upstream. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/fixed.md | 5 ++-- docs/index.md | 8 +++--- docs/scan-log.md | 4 +-- docs/scans/ginlix-ai-langalpha.md | 45 ++++++++++++++++++++++++++----- docs/scans/superdesigndev-treg.md | 16 +++++++++-- docs/work-with-me.md | 4 +-- 6 files changed, 64 insertions(+), 18 deletions(-) diff --git a/docs/fixed.md b/docs/fixed.md index fcb5a7b..97e2aa6 100644 --- a/docs/fixed.md +++ b/docs/fixed.md @@ -1,12 +1,12 @@ --- layout: default title: Findings a maintainer fixed -description: "The 25 scans in the AI PatchLab series where the maintainer shipped a fix — what was reported, and what landed upstream." +description: "The 26 scans in the AI PatchLab series where the maintainer shipped a fix — what was reported, and what landed upstream." --- # Findings a maintainer fixed -Of 109 scans, **25** ended with a maintainer shipping a fix. This page is the +Of 111 scans, **26** ended with a maintainer shipping a fix. This page is the short version of the argument: a report is only worth writing if someone can act on it. The fastest turnaround in the series was about six hours from filing to a merged pull @@ -20,6 +20,7 @@ release notes. | 2026-09-14 | [superlinked/sie](scans/superlinked-sie.html) | 156 | 1 real | | 2026-09-03 | [samuelgursky/davinci-resolve-mcp](scans/samuelgursky-davinci-resolve-mcp.html) | 97 | 1 real | | 2026-08-31 | [shy3130/tick-stock-panel](scans/shy3130-tick-stock-panel.html) | 77 | 1 real | +| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | | 2026-08-27 | [Zleap-AI/SAG](scans/zleap-ai-sag.html) | 60 | 1 real | | 2026-08-20 | [whiteguo233/OpenBiliClaw](scans/whiteguo233-openbiliclaw.html) | 373 | 1 real | | 2026-08-17 | [zilliztech/memsearch](scans/zilliztech-memsearch.html) | 90 | 1 real | diff --git a/docs/index.md b/docs/index.md index 9c2ec4a..65febd8 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,7 +1,7 @@ --- layout: default title: AI PatchLab Scans -description: "111 curated security scans of open-source AI agents, MCP servers and LLM apps - 25 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit." +description: "111 curated security scans of open-source AI agents, MCP servers and LLM apps - 26 confirmed fixes, run local-first with Semgrep, Gitleaks, Trivy and pip-audit." --- # AI PatchLab Scans @@ -20,7 +20,7 @@ remediation and confidence rules to normalize the findings. > **Want this run privately against your own codebase?** I do independent > security review of AI agents, MCP servers, and LLM apps — -> [**work with me →**]({{ '/work-with-me' | relative_url }}). 111 scans, 25 confirmed fixes, methodology in the open. +> [**work with me →**]({{ '/work-with-me' | relative_url }}). 111 scans, 26 confirmed fixes, methodology in the open. > **OpenAI just launched [Daybreak](https://openai.com/index/daybreak-securing-the-world/) and Patch the Planet.** > Same remediation loop, opposite trade-off: their path is a cloud frontier model; @@ -35,7 +35,7 @@ remediation and confidence rules to normalize the findings. ## Two shorter ways in -- [**Findings a maintainer fixed**]({{ '/fixed' | relative_url }}) — the 25 that resolved +- [**Findings a maintainer fixed**]({{ '/fixed' | relative_url }}) — the 26 that resolved upstream. The shortest version of the argument: a report is only worth writing if someone can act on it. - [**Scans that found nothing**]({{ '/clean' | relative_url }}) — 40 of them, published as @@ -141,7 +141,7 @@ filed, which is the usual outcome of a clean scan. | 2026-09-01 | [future-agi/future-agi](scans/future-agi-future-agi.html) | 1227 | 1 real — withheld | private | | 2026-08-31 | [shy3130/tick-stock-panel](scans/shy3130-tick-stock-panel.html) | 77 | 1 real | **fixed** | | 2026-08-30 | [SenteLabsAI/OpenExecutive](scans/sentelabsai-openexecutive.html) | 124 | 1 real — withheld | private | -| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | open | +| 2026-08-29 | [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) | 372 | 1 real | **fixed** | | 2026-08-28 | [Ontos-AI/knowhere](scans/ontos-ai-knowhere.html) | 129 | 1 real — withheld | private | | 2026-08-27 | [Zleap-AI/SAG](scans/zleap-ai-sag.html) | 60 | 1 real | **fixed** | | 2026-08-26 | [ascending-llc/jarvis-registry](scans/ascending-llc-jarvis-registry.html) | 235 | 1 real — withheld | private | diff --git a/docs/scan-log.md b/docs/scan-log.md index 108d98b..f0e1638 100644 --- a/docs/scan-log.md +++ b/docs/scan-log.md @@ -8,7 +8,7 @@ description: "The complete AI PatchLab scan log: every public repository scanned Every scan in the series, newest first, with the summary written on the day of the scan. 111 scans. For the compact index, see the [scan log home]({{ '/' | relative_url }}). -- **2026-09-24** — [superdesigndev/treg](scans/superdesigndev-treg.html) — 96 findings at `medium+` (1 critical, 23 high, 69 medium, 3 meta), **1 real — withheld, reported privately by email** — **"OpenRouter for agent tools"** (3k★, source-available; hosted at treg.to and self-hostable), a multi-tenant server holding a team's provider credentials Fernet-encrypted and proxying calls to 3,000+ upstream endpoints with the credential injected server-side so it never reaches the caller. Its `SECURITY.md` makes two load-bearing promises: the proxy never hands a key to the caller, and outbound requests are *"re-resolved at call time and internal/metadata addresses are refused."* **Zero of the 96 scanner findings is the vulnerability** — the real item came from tabulating the server's outbound request paths and asking which the call-time host guard actually covers. It does have a proper guard, `host_is_public` in `infra/upstream/ssrf.py`, and its address classification is genuinely thorough (every loopback/link-local/CGNAT/NAT64 and decimal/hex/octal/short-form IPv4 literal I tried was refused). The finding is that the guard is wired into the main proxy path and **not** into a sibling outbound path a plain member can drive; the registration-time check that does run there is, by its own docstring, a static check that allows any DNS name, so a name resolving to an internal address passes it and nothing re-resolves before the request goes out — a member-authenticated, semi-blind [SSRF on the unguarded sibling of a guarded transport](scans/liaohch3-claude-tap.html). Confirmed with the [run-the-primitive](scans/observal-observal.html) discipline: treg's own two guard functions run side by side on a host that resolves internally return "allowed" (registration) and "refused" (call-time), and the sibling path calls only the first — one differential, using the project's code not my reasoning. Graded **Medium**: the injected credential on the affected path is the member's own, so it is internal + cloud-metadata reach, not cross-tenant theft; the org boundary itself held everywhere ([the probe could return "no" on isolation and did](scans/tracecathq-tracecat.html)). **The fix is one guard applied at the outbound path** and changes nothing for a legitimate public target. The 96: seventeen `sqlalchemy-text` are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) (Alembic `SET`/`CREATE INDEX CONCURRENTLY` from module constants + a test-reset `TRUNCATE` over dialect-quoted table names, no value interpolated); eighteen `v-html` render the project's own static tutorial/snippet HTML with an `esc(&<>)` helper on every interpolated field; four `insecure-file-permissions` are the project [tightening to `0o700`](scans/realiti4-claude-swap.html) (the rule's suggestion would loosen them); one Trivy critical is `anyio` 4.14.1→4.14.2 IDNA hostname-confusion, a worthwhile lockfile bump reachable only by a network-positioned attacker against a non-ASCII upstream host. Coverage `partial` and stated so — Semgrep hit a Windows per-file temp error on five large modules (recovered by a targeted re-run of just those five), and pip-audit left `uv.lock` unaudited while Trivy read it; neither touches the finding. **Strict-norm** (`SECURITY.md` forbids public vuln issues, PVR disabled — checked `{"enabled":false}` — names an email) · reported **privately by email to the maintainer**; post-only here, finding withheld at class level, **the send is the operator's step** +- **2026-09-24** — [superdesigndev/treg](scans/superdesigndev-treg.html) — 96 findings at `medium+` (1 critical, 23 high, 69 medium, 3 meta), **1 real — withheld, reported privately by email** — **"OpenRouter for agent tools"** (3k★, source-available; hosted at treg.to and self-hostable), a multi-tenant server holding a team's provider credentials Fernet-encrypted and proxying calls to 3,000+ upstream endpoints with the credential injected server-side so it never reaches the caller. Its `SECURITY.md` makes two load-bearing promises: the proxy never hands a key to the caller, and outbound requests are *"re-resolved at call time and internal/metadata addresses are refused."* **Zero of the 96 scanner findings is the vulnerability** — the real item came from tabulating the server's outbound request paths and asking which the call-time host guard actually covers. It does have a proper guard, `host_is_public` in `infra/upstream/ssrf.py`, and its address classification is genuinely thorough (every loopback/link-local/CGNAT/NAT64 and decimal/hex/octal/short-form IPv4 literal I tried was refused). The finding is that the guard is wired into the main proxy path and **not** into a sibling outbound path a plain member can drive; the registration-time check that does run there is, by its own docstring, a static check that allows any DNS name, so a name resolving to an internal address passes it and nothing re-resolves before the request goes out — a member-authenticated, semi-blind [SSRF on the unguarded sibling of a guarded transport](scans/liaohch3-claude-tap.html). Confirmed with the [run-the-primitive](scans/observal-observal.html) discipline: treg's own two guard functions run side by side on a host that resolves internally return "allowed" (registration) and "refused" (call-time), and the sibling path calls only the first — one differential, using the project's code not my reasoning. Graded **Medium**: the injected credential on the affected path is the member's own, so it is internal + cloud-metadata reach, not cross-tenant theft; the org boundary itself held everywhere ([the probe could return "no" on isolation and did](scans/tracecathq-tracecat.html)). **The fix is one guard applied at the outbound path** and changes nothing for a legitimate public target. The 96: seventeen `sqlalchemy-text` are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) (Alembic `SET`/`CREATE INDEX CONCURRENTLY` from module constants + a test-reset `TRUNCATE` over dialect-quoted table names, no value interpolated); eighteen `v-html` render the project's own static tutorial/snippet HTML with an `esc(&<>)` helper on every interpolated field; four `insecure-file-permissions` are the project [tightening to `0o700`](scans/realiti4-claude-swap.html) (the rule's suggestion would loosen them); one Trivy critical is `anyio` 4.14.1→4.14.2 IDNA hostname-confusion, a worthwhile lockfile bump reachable only by a network-positioned attacker against a non-ASCII upstream host. Coverage `partial` and stated so — Semgrep hit a Windows per-file temp error on five large modules (recovered by a targeted re-run of just those five), and pip-audit left `uv.lock` unaudited while Trivy read it; neither touches the finding. **Strict-norm** (`SECURITY.md` forbids public vuln issues, PVR disabled — checked `{"enabled":false}` — names an email) · reported **privately by email to the maintainer**; post-only here, finding withheld at class level · sent by email 2026-09-24, ten minutes after this entry was published - **2026-09-23** — [can4hou6joeng4/boss-agent-cli](scans/can4hou6joeng4-boss-agent-cli.html) — 26 findings at `medium+` (7 high, 16 medium, 3 meta), **1 real — withheld, filed privately** — a **CLI for the BOSS Zhipin recruitment platform** (2k★, MIT; a human drives it from a terminal wizard, an agent drives the same workflows over a JSON-envelope API, an ~77-tool MCP server, or a Python SDK) that authenticates by reusing the operator's own logged-in browser session. Responsive maintainer — 67 merged PRs from 7 human authors in 60 days — which is why it cleared the pre-check. **Zero of the 26 scanner findings is the vulnerability**, which is the whole point of the entry: the real item is the *absence* of a check on an otherwise-working local channel, the shape no rule can match, and it is the [desktop-loopback-inversion](scans/liaohch3-claude-tap.html) class — a `127.0.0.1` control surface that binds neither *who* may call it nor which host the request claims, driving privileged browser actions against a site where the operator is authenticated. Confirmed with a **server-side exploit primitive against the shipped daemon** (a foreign `Host` is served, a cross-origin "simple" request reaches the command channel with no preflight, a hostile `Origin` is not rejected) plus a read of the extension source for the browser half — the [run-the-primitive discipline](scans/observal-observal.html), because reading alone could not separate "the channel is drivable" from "the channel is drivable *and* blind does not blunt it." Graded **High not Critical** on an honest precondition (the bridge component must be running, and the user has to start it deliberately; once started it stays up while the browser side is connected — *corrected 2026-09-24: first published as "a bounded window", the one claim reasoned about rather than checked, and wrong both ways; the private report was corrected the same day*) and the sensitive scope is the one site the tool is built around; **the fix was written before it was named**, narrow and verified not to break the tool's own two legitimate clients ([write the patch first](scans/roflcoopter-viseron.html)). The 26: eight SQL hits (five high) are the [#1 identifier FP](scans/aurelio-labs-semantic-router.html) — every interpolation is a hardcoded table/column name with values bound as `?`, settled by reading the call sites; eight of sixteen mediums are one `github-actions-mutable-action-tag` flood; the rest are two `sha1()[:16]` dedup-id digests, two `urllib` calls to a local DevTools endpoint and a build script, a React rule on a loopback health-check `fetch`, and a Python-3.7 compatibility rule that is not a security check. Coverage `partial` and stated so — Semgrep did not cover every file and the dependency scan left `uv.lock` unaudited (Trivy read it, the [luck that hides the bug](scans/roflcoopter-viseron.html)), neither touching the finding. **Strict-norm** (`SECURITY.md` forbids public vulnerability issues, gives an email and a PVR link) · PVR-enabled, filed **privately, accepted into triage first try** ([GHSA-xp54-2hxc-w78q](https://github.com/can4hou6joeng4/boss-agent-cli/security/advisories/GHSA-xp54-2hxc-w78q)) **with the `vulnerabilities` array** — the [viseron rule](scans/roflcoopter-viseron.html) holds again · post-only, finding withheld at class level under embargo - **2026-09-22** — [overwirehq/claude-code-telegram](scans/overwirehq-claude-code-telegram.html) — 54 findings at `medium+` (1 critical, 21 high, 30 medium), **0 first-party defects — dependency currency, post-only** — a **Telegram bot that gives authorised users remote access to Claude Code**, i.e. it runs file and Bash tools on a host machine by design, so the whole security surface *is* the boundary around that execution (2k★, MIT, org-backed; active — 15 merged PRs from 7 authors and 9 closed issues in 60 days; strict-norm: real `SECURITY.md`, PVR enabled). Picked with the manual queue at zero. **The story is what precise security documentation looks like.** The real boundary is the `can_use_tool` callback that stops Claude — when steered off-course by content it reads mid-task — from acting outside the approved directory, and the maintainers know and *state* exactly what it covers: path validation is scoped in the docs to precisely six tools (`Read`, `Write`, `Edit`, `MultiEdit`, `NotebookEdit`, `NotebookRead`), while `Grep`/`Glob`/`LS` are deliberately left to the OS sandbox. Reading that as "Read is guarded but Grep is not, so the boundary is incomplete" is the plausible-but-wrong finding this site exists to resist — the boundary is [advertised, not accidental](scans/tracecathq-tracecat.html), and `ROADMAP-v2.md` even records the SDK's own `CanUseToolShadowedWarning` naming every tool the callback will never see, filed as tracking item #221. Two more choices earn credit: guarded tools are deliberately *stripped from* the SDK `allowed_tools` list (a pre-approved tool never produces a `can_use_tool` request, so leaving them in would render the checks silently inert — [issue #219](https://github.com/overwirehq/claude-code-telegram/issues/219)), and `autoAllowBashIfSandboxed` is disabled whenever the boundary checks must run, closing a second bypass. That is a maintainer who traced how the framework resolves permissions rather than trusting that a config allow-list is an enforcement boundary — it isn't, and the code says so. **When the machine is built and documented this carefully, the finding moves to the dependency manifest.** The 30 Trivy advisories in `poetry.lock` split cleanly by reachability: **17 are opt-in-only and not reachable on a default install** — `starlette` (6) and `python-multipart` (3) need the FastAPI webhook server (`ENABLE_API_SERVER=false` default); `pyjwt` (5) needs token auth (`ENABLE_TOKEN_AUTH=false` default, and the SECURITY.md documents token auth as non-functional, [#58](https://github.com/overwirehq/claude-code-telegram/issues/58)); the MCP SDK highs (3) need MCP enabled — while **13 are unconditionally installed** (`anyio` incl. the critical IDNA/TLS advisory, the `cryptography`/OpenSSL cluster, `urllib3`, `idna`, `requests`, `pydantic-settings`, `python-dotenv`) and are the genuine currency gap. The direct deps are current; the drift is transitive, which the repo's existing `.github/dependabot.yml` (configured for *version* updates, not *security* updates) will not raise on its own. **No advisory filed** — there is no first-party vulnerability. Of the other 24: 16 `github-actions-mutable-action-tag` (SHA-pin hardening); one `pull_request_target` checkout wrapped in a 40-line threat-model header (read-only tool allowlist, secrets scrubbed, "worst case is a prompt-injected review comment" — [the trigger decides severity](scans/lightseekorg-tokenspeed.html)); two `sqlalchemy-execute-raw-query` [identifier FPs](scans/aurelio-labs-semantic-router.html) (only a generated run of `?` placeholders is interpolated, values bound via `execute(query, params)`); and three gitleaks hits that are a `your-api-secret` placeholder plus two fake tokens in a test asserting the project's own `_redact_secrets()` helper scrubs them ([credited defence](scans/realiti4-claude-swap.html)). Coverage `partial` and honestly so: pip-audit resolved **no dependencies** from a `pyproject.toml` declaring `dynamic = ["dependencies"]` — reading identically to "clean" — and Trivy's `poetry.lock` read carried the run, the [same dependency blind spot](scans/harnessrouter-harnessrouter.html) as yesterday reached by a different manifest shape. @@ -34,7 +34,7 @@ Every scan in the series, newest first, with the summary written on the day of t - **2026-08-28** — [Ontos-AI/knowhere](scans/ontos-ai-knowhere.html) — 129 findings, **1 real — withheld** — a document-ingestion pipeline sold as the memory layer between dirty documents and AI agents (2.7k★): one missing-authentication gap on an internal storage-event webhook whose verifiers are stubs, verified by execution (wrong token rejected, absent header accepted), and a 68-CVE dependency table that all resolves to a single lockfile · drafted for private delivery, post-only - **2026-08-27** — [Zleap-AI/SAG](scans/zleap-ai-sag.html) — 60 findings, **1 real** — an agent framework whose login did not verify the password: any name signed in as the owner account · ✅ **Resolved 2026-08-30 in 3 days (19th fix in the series)** — posing the product question beat proposing a patch - **2026-09-01** — [future-agi/future-agi](scans/future-agi-future-agi.html) — 1227 findings (1224 above the medium floor), **1 real — withheld** — an **end-to-end LLM/agent observability and evaluation platform** (1.9k★, Apache-2.0 + EE): OTLP trace ingestion into a ClickHouse store of prompts/completions/tokens/costs, an evaluation engine, a simulation runner, a Go **AI gateway** (`agentcc-gateway`, an OpenAI/Anthropic/Gemini-compatible proxy with per-org keys, rotation and RBAC) and a Django backend, with a real `SECURITY.md` (named email, SLA table, safe harbor, in-scope list, public vuln issues forbidden). **Largest raw count in the series — and one real finding no rule represented.** The class: **one service in a multi-service deployment is published on all interfaces while every service beside it, including every datastore, is correctly loopback-bound, and the production overlay inherits the exposure unchanged** — the exposed service is the most dangerous one in the file to leave open, and reaching it hands an unauthenticated caller a path to the trace store. Three reportability properties, each a series pattern: **the maintainer's own compose comment is the oracle** (*"bound to 127.0.0.1 so only the host can reach them"*, applied to Postgres/ClickHouse/Redis/MinIO/Temporal/collector — the finding is the outlier), the [contract-vs-artifact](scans/vexa-ai-vexa.html) footing in a deployment file; **N deployment paths, and the question is whether they agree** — the production overlay re-binds every secret with `${VAR:?must be set for production}` (genuinely good) but names only `environment:`/`restart:`, not `ports:`, so the documented production command inherits the exposure; and **I proved it with the vendor's own merge engine** — rendered the exact `deploy/README.md` command through `docker compose config` and read the published bindings out of the result (exposed service on `0.0.0.0`, every datastore on `127.0.0.1`), [running the primitive](scans/nottelabs-notte.html) rather than reasoning about override semantics. **Scope discipline kept it to one**: two other `0.0.0.0` services are the intended public surface (frontend + reverse-proxied API), and two management UIs published without a loopback bind sit behind Compose `profiles:` — opt-in, absent from the default and documented-production deployments, and the rendered config confirmed they don't appear. **Two auth hypotheses returned NO and that is part of the record**: the gateway mounts a large admin plane *outside* its `/v1/` auth prefix (mint keys, set org configs, rotate creds) — the [claude-tap inversion](scans/liaohch3-claude-tap.html) shape — but all **24** admin handlers across four files call an in-handler `requireAdmin`/`checkAdminAuth` with a constant-time compare that **fails closed on an unset token**; and the `/v1/` middleware short-circuits on a valid license token (a [caller-influenced auth branch](scans/liaohch3-claude-tap.html)) but the verifier pins RS256, resolves `kid` from a **local** key map, and validates the full claim set, so branch-then-serve is safe. **Noise**: 100 mutable-action-tag GHA rows, 94+44 raw/formatted-SQL that resolve to the [#1 identifier FP](scans/mnemosyne-oss-mnemosyne.html) (ClickHouse builders interpolate regex-validated keys and an allowlisted bucket function, bind every value via `%(param)s` — the four taint-tracked hits were audit-log f-strings and the allowlisted bucket fn), 3 gitleaks hits all covered by the repo's own `.gitleaks.toml`. **The 26 criticals are a coverage story**: `scan_dependency` found *no root manifest* because this is a **monorepo** (manifests under `futureagi/`), so the ChromaDB/LiteLLM/Authlib/Django/langchain criticals came from Trivy's lockfile parse and want the *version-match → reachable → mitigated* gate — several are Proxy-only/transport-only in libs used as clients. **Monorepo root-only coverage gap, re-confirmed as the top backlog item.** Strict-norm · reported **privately** to `security@futureagi.com` with a full dossier + `docker compose config` reproduction · post-only, finding **withheld** (class only) until remediated · the email send is the operator's manual step -- **2026-08-29** — [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) — 372 findings above the medium floor, **1 real** — a **financial-market agent platform** (1.7k★, Apache-2.0, daily merges, FastAPI + React + a Daytona/Docker code sandbox), and one of the most carefully defended codebases in the series. Semgrep's top finding — `pull_request_target` checking out untrusted fork code — **retired on evidence outside the repository tree**: the GitHub environments API shows a `fork-ci` environment with required reviewers and `prevent_self_review`, and the workflow pins the immutable head SHA, sets `persist-credentials: false`, and drops the token to `contents: read`. The author even documented that `uv sync` runs fork build hooks before pytest — a better threat model than the rule that flagged it. All 30 gitleaks hits are fixtures in the project's **own** secret-redaction and leak-detection tests; the 102-finding SQL cluster binds every value with `%s` and interpolates only a constant column list. What survived is a **composite**: `GET /api/v1/preview/{workspace_id}/{port}` is unauthenticated and its docstring promises it "does NOT start stopped sandboxes (to prevent denial-of-wallet)" — but it enforces that with a DB `status == "running"` check and then calls the *acquisition* path, which reaches `PTCSandbox.reconnect`, documented as "it **starts a stopped sandbox**." **Four sibling unauthenticated routes in the same codebase refuse exactly that call**, citing "a stale 'running' DB row with no warm session," and use the no-wake `get_session_if_ready` accessor whose fence parameter was made mandatory "so a new caller cannot omit the fence." The majority sibling is the contract; the fifth route never inherited it. The lead came from the tool's **coverage warning**, not its findings: Semgrep could not parse `claude.yml`, `release.yml`, or `sandbox-integration.yml`, which were precisely the workflows worth reading. +- **2026-08-29** — [ginlix-ai/LangAlpha](scans/ginlix-ai-langalpha.html) — 372 findings above the medium floor, **1 real** — a **financial-market agent platform** (1.7k★, Apache-2.0, daily merges, FastAPI + React + a Daytona/Docker code sandbox), and one of the most carefully defended codebases in the series. Semgrep's top finding — `pull_request_target` checking out untrusted fork code — **retired on evidence outside the repository tree**: the GitHub environments API shows a `fork-ci` environment with required reviewers and `prevent_self_review`, and the workflow pins the immutable head SHA, sets `persist-credentials: false`, and drops the token to `contents: read`. The author even documented that `uv sync` runs fork build hooks before pytest — a better threat model than the rule that flagged it. All 30 gitleaks hits are fixtures in the project's **own** secret-redaction and leak-detection tests; the 102-finding SQL cluster binds every value with `%s` and interpolates only a constant column list. What survived is a **composite**: `GET /api/v1/preview/{workspace_id}/{port}` is unauthenticated and its docstring promises it "does NOT start stopped sandboxes (to prevent denial-of-wallet)" — but it enforces that with a DB `status == "running"` check and then calls the *acquisition* path, which reaches `PTCSandbox.reconnect`, documented as "it **starts a stopped sandbox**." **Four sibling unauthenticated routes in the same codebase refuse exactly that call**, citing "a stale 'running' DB row with no warm session," and use the no-wake `get_session_if_ready` accessor whose fence parameter was made mandatory "so a new caller cannot omit the fence." The majority sibling is the contract; the fifth route never inherited it. The lead came from the tool's **coverage warning**, not its findings: Semgrep could not parse `claude.yml`, `release.yml`, or `sandbox-integration.yml`, which were precisely the workflows worth reading. · one focused public issue filed ([#378](https://github.com/ginlix-ai/LangAlpha/issues/378)), PVR off and no `SECURITY.md` anywhere · ✅ **Resolved 2026-09-25 in 27 days (26th fix in the series)** — [#425](https://github.com/ginlix-ai/LangAlpha/pull/425) did not take the proposed accessor swap; it removed the premise. The old preview route now reads the database and redirects to the app's owner-only link, so it touches no session and can wake nothing, and the workspace UUID stopped being a credential altogether: report frames load under an HMAC-signed grant that expires. A regression test names #378. **Re-verified at the fix commit**: the differential that found it went from one waking unauthenticated route out of five to none, and the new redirect was probed as a new check — the link code it hands a UUID holder opens nothing, because the schema itself forbids sharing an app link (a `CHECK` constraint requires `shared_at IS NULL`). - **2026-08-26** — [ascending-llc/jarvis-registry](scans/ascending-llc-jarvis-registry.html) — 235 findings above the medium floor, **1 real — withheld** — an **enterprise MCP/A2A gateway** that brokers per-user OAuth credentials to downstream tool servers (2.8k★, Apache-2.0, commercial backing, **7 distinct human PR authors in 60 days**), and the cleanest demonstration yet of why the mitigation gate is not optional. All four criticals retired: the **LiteLLM CVEs** (SQL injection, Host-header auth bypass, MCP command execution, key-gen privilege escalation) are Proxy-server-only, and this repo declares `litellm>=1.50.0` then **never imports it** — one grep outside the lockfile returns the dependency declaration and nothing else. The largest family, **21× GitHub Actions shell injection**, retired on a single fact: there is no `pull_request_target` anywhere in the repository, so every flagged workflow runs either from a fork with a read-only token and no secrets, or from a context that already requires write access. **16 gitleaks "secrets"** are docs example output and MongoDB seed fixtures, against a `tartufo.toml` that carries a written reason on every exclusion; **12× `logger-credential-leak`** logs usernames and source ids, never a credential; the lone **`unverified-jwt-decode`** is by design and its two callers both verify-then-branch, over a token layer that pins RS256, pins the `kid`, enforces issuer and audience, and blocks token-class confusion with positive equality checks. **The near-miss is the story**: the registry allows CORS origins by a regex whose `.*\.compute.*\.amazonaws\.com` arm admits the public DNS name of *every EC2 instance on the internet*, with `allow_credentials=True` — confirmed allowed against real Starlette 1.6.0, with the controls correctly denied. It is still **inert**, because every cookie the app sets is `SameSite=Lax` and `amazonaws.com` compute domains are on the Public Suffix List, so the attacker's `fetch()` carries no session at all. Four verification steps said "real"; the fifth — grep for the mitigation before flagging — retired it, and this series has filed that same class as a genuine finding against five other projects. **The one real defect is object-level authorization no scanner can see**: two sibling handlers in one router return per-user state from the same deterministic identifier, one enforces ownership against the caller and the other takes no caller identity at all, behind a permission every role holds. The intra-repo differential is the proof — the guarded sibling *is* the project's own contract. Routed privately to `security@ascendingdc.com` per the project's SECURITY.md, with a patch; public detail withheld. - **2026-08-25** — [langflow-ai/openrag](scans/langflow-ai-openrag.html) — 213 findings above the medium floor, **1 real — withheld** — a **single-package Retrieval-Augmented Generation platform** built on OpenSearch + Langflow (4.5k★, Apache-2.0, commercial backing, **47 merged PRs from 13 authors in 60 days**), and a textbook "the scanner count is the least informative thing about this repo" target. 213 medium-plus findings resolve to one recurring best-practice class (**121× GitHub Actions mutable action tags**) and a long tail of well-understood false positives: **15 gitleaks "secrets"** all in `.secrets.baseline` / empty `.env.example` slots / `Makefile` shell variables / docs / a test fixture (credit the `.secrets.baseline` and they collapse to zero), a **"critical" Kubernetes RBAC** that is the OpenRAG operator's *own* secret-management role doing its job, and **10× `logger-credential-leak`** on metadata-only log lines — the series' most reliable false positive, still awaiting its confidence downgrade. **The one finding that matters is a composite no scanner can see**: a hard-coded/shipped-default cryptographic secret on a token-signing boundary (CWE-321/798), reachable in the default self-hosted deployment, where the hardening applied to one token path was not applied to its sibling. It was found by diffing `securityconfig/` against `cloud_securityconfig/` and tracing a token from mint to validation — neither of which any rule performs — verified offline against OpenRAG's own token-validation logic, and reported privately via GitHub PVR (**GHSA-xv8v-6c28-v78p**, `triage`) with a concrete fix written against the repo's own key-generation architecture. Detail withheld per OpenRAG's SECURITY.md; this entry carries the class only. diff --git a/docs/scans/ginlix-ai-langalpha.md b/docs/scans/ginlix-ai-langalpha.md index ad5dd2d..4674946 100644 --- a/docs/scans/ginlix-ai-langalpha.md +++ b/docs/scans/ginlix-ai-langalpha.md @@ -9,12 +9,16 @@ date: 2026-08-29 **Repository:** [ginlix-ai/LangAlpha](https://github.com/ginlix-ai/LangAlpha) **Commit scanned:** `f5232fa2` (main at scan time) **Scan date:** 2026-08-29 -**Disclosure status:** disclosed — one real finding filed as a single focused -public issue. No `SECURITY.md` at the repository root, in `.github/`, in -`docs/`, or at the organisation level, and private vulnerability reporting is -disabled (confirmed with an empty-payload control request, which returned -`403 Repository does not have private vulnerability reporting enabled` — it -files nothing). A public issue is the only channel the project offers. +**Disclosure status:** ✅ **resolved** — one real finding, filed as a single +focused public issue on the day of the scan and fixed by the maintainer 27 days +later in [ginlix-ai/LangAlpha#425](https://github.com/ginlix-ai/LangAlpha/pull/425), +which closed the issue as completed. The fix went further than the one proposed +here: it retired the design the finding took as given. No `SECURITY.md` at the +repository root, in `.github/`, in `docs/`, or at the organisation level, and +private vulnerability reporting is disabled (confirmed with an empty-payload +control request, which returned `403 Repository does not have private +vulnerability reporting enabled` — it files nothing). A public issue is the +only channel the project offers. ## Summary @@ -181,6 +185,35 @@ high-severity findings automatically instead of by hand. [issue #378](https://github.com/ginlix-ai/LangAlpha/issues/378). One finding, not a grouped review. No PR opened: the fix is one file, but choosing 404 vs 503 for a sleeping sandbox is a product decision the maintainer should make. +- **2026-09-25** — **fixed** in [#425](https://github.com/ginlix-ai/LangAlpha/pull/425) + (`2eab033f`) and the issue closed as completed, 27 days after filing. The + maintainer did not adopt the accessor swap proposed above; the fix removes + the premise instead. The old `/api/v1/preview/{workspace_id}/{port}` route no + longer resolves a signed URL or touches a session at all: it reads the + registered preview command from the database, gets or creates the app's one + link, and redirects to it. That link opens for the signed-in workspace owner + alone, and the owner's path is now the only one that resolves a preview — in + the new docstring's words, "a stopped sandbox is started for the owner and for + nobody else." This page described the workspace UUID as the route's bearer + credential, by design. #425 retired that design: report frames now load under + an HMAC-signed grant that expires, and running apps open through the owner's + private link. A regression test names #378 — with the workspace manager and + the sandbox lookup both patched to raise, the redirect still answers `302` + and the manager is never called. +- **2026-09-26** — re-verified at `2eab033f`, starting with the differential + that found the bug. At the scanned commit, five unauthenticated routes touched + a workspace session and one called the waking accessor; at the fix, none do, + and every remaining caller of `get_session_for_workspace` sits behind an + authenticated owner. The redirect was then probed as a new check, because it + still answers anyone who holds a UUID. What it hands back is an app link + code, and an app link cannot be shared — the migration's `share_links_app_shape` + constraint requires `shared_at IS NULL` on every app row — so for anyone but + the owner the code opens nothing and returns the same `404` as an unknown one. + The public page that does resolve a preview reaches the waking path only past + two independent ownership checks (`resolve_link`, then `require_workspace_owner` + inside `_get_sandbox`). Nor can the redirect be used to rewrite the owner's + link: its insert is `ON CONFLICT DO NOTHING`, and an old URL's page rides along + as `?path=` rather than being stored. ## Reproduce diff --git a/docs/scans/superdesigndev-treg.md b/docs/scans/superdesigndev-treg.md index 2c007c4..427236e 100644 --- a/docs/scans/superdesigndev-treg.md +++ b/docs/scans/superdesigndev-treg.md @@ -14,8 +14,9 @@ date: 2026-09-24 project's `SECURITY.md` (which asks that vulnerabilities not be opened as public issues; GitHub private vulnerability reporting is disabled on the repo, and the policy names an email address, so email is the channel). Detail below is kept at -class level until the maintainer has had a chance to respond. **The private send -is the operator's step and had not gone out when this page was published.** +class level until the maintainer has had a chance to respond. The private send +was the operator's step and had not gone out when this page was published; it +went out ten minutes later (see the timeline below). ## Summary @@ -161,6 +162,17 @@ curated by hand; scanner output alone is not a vulnerability report. This page w updated with full technical detail once the maintainer has responded and a fix has shipped.* +## Disclosure timeline + +- 2026-09-24 — scan run at `240c595`; finding verified by running treg's own two host + guards side by side +- 2026-09-24 — channel probed: `SECURITY.md` asks that vulnerabilities not be opened as + public issues and names an email address; GitHub private vulnerability reporting is + disabled (`{"enabled":false}`) +- 2026-09-24 13:31 UTC — this page published, finding withheld at class level +- 2026-09-24 13:41 UTC — reported privately by email to the address in `SECURITY.md` +- 2026-09-26 — no reply yet, no bounce; none of the reported files has changed upstream + ## Reproduce ```bash diff --git a/docs/work-with-me.md b/docs/work-with-me.md index 53a9a29..090810e 100644 --- a/docs/work-with-me.md +++ b/docs/work-with-me.md @@ -1,7 +1,7 @@ --- layout: default title: Work with me — AI agent & LLM security review -description: "Independent security review of AI agents, MCP servers and LLM applications. 109 public scans and 25 confirmed fixes as the proof; first-look engagements from $750." +description: "Independent security review of AI agents, MCP servers and LLM applications. 111 public scans and 26 confirmed fixes as the proof; first-look engagements from $750." --- # Work with me @@ -12,7 +12,7 @@ If you're shipping an agent framework, an MCP server, a RAG pipeline, or anythin ## The proof is public -Everything I'd do for you, I've done in the open — [109 curated scans]({{ '/' | relative_url }}) of well-known AI/agent projects, with **25 confirmed fixes** where maintainers acted on the findings, and a documented methodology behind each one. +Everything I'd do for you, I've done in the open — [111 curated scans]({{ '/' | relative_url }}) of well-known AI/agent projects, with **26 confirmed fixes** where maintainers acted on the findings, and a documented methodology behind each one. - **Real findings, not scanner noise.** A typical scan starts at hundreds of raw findings and ends at a handful that matter — because the rest are false positives, by-design patterns, or vendored third-party code. The curation is the value. - **Reachability over rule-count.** A `subprocess(shell=True)` that's safe under operator trust can become a multi-tenant sandbox escape once you trace who reaches it — and a scary-looking SQL string can be fully gated by a five-character allowlist. The verdict goes wherever the *code* goes, which means reading the code, not the rule.