One review, end to end, on a sample committed to this repository.
By the end you should be able to say four things about the change under review, from the artifacts alone:
- what capability it added, widened or removed;
- why the top result matters;
- what the run did not establish; and
- what happens next, and who is allowed to do it.
The target is about 10 minutes. That is a target used to size this guide, not
a measured result — observed times are recorded in
design-partner-pilot-results.md.
Nothing here asks you to author a policy first. If you are evaluating a coding-agent host boundary rather than building a tool surface, you never need a manifest at all — see Route H.
Read this before you install. Agents Shipgate is published through more than
one channel, and they do not all implement the same runtime contract. The
newest published release is v0.15.0, which implements runtime contract
10; the agent-control envelope that later sections of the in-repo
documentation describe landed after that tag.
| Channel | How you get it | Runtime contract | Has control.* / current-control.json |
Accepts check --format agent-boundary-json |
Qualification |
|---|---|---|---|---|---|
Published release v0.15.0 |
pipx install agents-shipgate |
10 | No. agent-handoff.json is shipgate.agent_handoff/v1; read its gate block instead |
No. That build accepts only --format codex-boundary-json |
Qualified release |
| Unqualified preview | gh release download preview-<version> --repo ThreeMoonsLab/agents-shipgate --pattern '*.whl', then pip install ./<wheel> |
that of the source commit it was cut from | Yes | Yes | None, by construction — no adjudicated corpus, no qualification artifact, nothing signed. See release-evidence-policy-decision.md § Amendment 2 |
| Source checkout | git clone, then ./shipgate … from the checkout |
that of the checkout | Yes | Yes | Not a distributed build |
Pick one channel and stay on it for the whole walkthrough. Every command in
One review, end to end runs on the published
release, and reaches the same verdict there — blocked, can merge without human: false, exit 0. What differs is how much the artifacts say. The
excerpts below are from a source checkout; where the released build renders
something else, the step names v0.15.0 and shows what it prints instead —
a test fails if a step quotes output only one channel produces without saying
which.
Whatever you install, ask it what it is before you trust a field name:
agents-shipgate --version
agents-shipgate contract --jsoncontract --json prints the contract_version the installed build actually
implements. Do not assume the floor this documentation was written against.
pipx install agents-shipgatepipx upgrade agents-shipgate refreshes a stale copy — a plain install is a
no-op when an older build is already there. Alternatives:
python -m pip install agents-shipgate # global pip
uv tool install agents-shipgate # via uv
uvx agents-shipgate --help # one-shot via uv, no permanent install
python -m agents_shipgate --help # run from a pip install without PATHThe CLI binary is agents-shipgate; a short alias shipgate is also
installed. Agents Shipgate requires Python 3.12 or newer. If your project uses
an older runtime, install the CLI with pipx or uv against a 3.12+
interpreter rather than into the project environment — your agent project does
not need Python 3.12 itself.
One fetch, no install, stdlib only:
curl -sSL https://raw.githubusercontent.com/ThreeMoonsLab/agents-shipgate/main/tools/shipgate-detect.py \
| python3 - --workspace . --jsonContinue when is_agent_project: true, or when suggested_sources or
codex_plugin_candidates is non-empty, or the workspace already has a
shipgate.yaml.
If is_agent_project: false, suggested_sources: [],
codex_plugin_candidates: [], and python_parse_truncated: false, this is
not the right tool for this repository. python_parse_truncated: true means the
Python parse stopped at its cap, so that negative describes the files that were
read rather than the repository — re-run with
--max-python-files <workspace_signals.python_file_total> before concluding
anything. zero-install.md covers the uvx and GitHub
Action variants that also avoid a local install.
samples/ai_generated_refund_pr is a
committed two-commit history. The base is a support agent that can only search
a knowledge base. The head commit — the kind a coding agent opens — adds a
refund tool to the MCP export:
{ "name": "support.search_kb",
"annotations": { "readOnlyHint": true, "idempotentHint": true },
"auth": { "type": "oauth2", "scopes": ["support:kb:read"] } },
+ { "name": "stripe.create_refund",
+ "annotations": { "readOnlyHint": false, "destructiveHint": true },
+ "auth": { "type": "oauth2", "scopes": ["stripe:*"] } }The manifest beside it still declares one scope, support:kb:read, and no
approval policy for the new tool.
agents-shipgate fixture run ai_generated_refund_prThe fixture builds the base/head git history in a temporary directory and runs
verify across it. Nothing in your own repository is touched — not even a
report directory:
Fixture: ai_generated_refund_pr
Mode: verify
Merge verdict: blocked
Decision: blocked
Can merge without human: false
Reports: /tmp/shipgate-fixture-ai_generated_refund_pr-<random>/ai_generated_refund_pr/reports
Verifier: …/reports/verifier.json
PR comment: …/reports/pr-comment.md
Static-verdict boundary: This verdict covers deterministic static evidence
only. Agents Shipgate did not execute the agent or prove runtime behavior,
tool routing, credential enforcement, or safety.
The Reports: line names the directory it wrote; every path below is
relative to it. --out <absolute-path> puts them somewhere you choose instead
— a relative --out resolves inside the fixture copy, not your shell's
directory.
Open pr-comment.md — the same text the GitHub Action posts on a pull
request. It leads with the capability delta, by subject:
- Capability delta (analysed surface): 2 subjects across 6 changes (+1 added, 2 modified, -0 removed)
- Top capability changes by subject:
- `stripe.create_refund`: added action — blocks release; added tool — blocks
release; broadened action destructive — high-risk effect destructive added;
blocks release; …
- `support.search_kb`: broadened action (support:kb:read -> support:kb:read) —
Capability binding_hash changed without a proven direction; review required
That is the answer to "what did this capability change": one tool added, one existing tool whose binding moved. You did not have to read the identity model to get it — the subject is the tool you would open.
On the published v0.15.0 build this step reads differently. Per-subject
grouping postdates that release, so its pr-comment.md counts changes rather
than subjects and lists them flat, with no support.search_kb row:
- Capability delta: +2, 3 modified, -0
- Top capability changes:
- `stripe.create_refund`: blocks release; Capability added. (tools.json)
- `stripe.create_refund`: blocks release; tool added (tools.json)
- `stripe.create_refund`: blocks release; high-risk effect financial_action added (tools.json)
- `stripe.create_refund`: blocks release; high-risk effect destructive added (tools.json)
- `stripe.create_refund:stripe:*`: blocks release; scope added (tools.json)
Same change, same verdict; you read it per change instead of per subject.
report.md in that same directory orders findings by subject, most urgent
first:
Decision: blocked
Reason: 4 active findings block release.
Blockers (4):
- CRITICAL SHIP-POLICY-APPROVAL-MISSING — stripe.create_refund lacks a declared approval policy
- CRITICAL SHIP-ACTION-DESTRUCTIVE-ROLLBACK-MISSING — stripe.create_refund has destructive capability without required controls
- CRITICAL SHIP-ACTION-FINANCIAL-WRITE-CONTROL-MISSING — stripe.create_refund has financial write capability without required controls
- CRITICAL SHIP-ACTION-WILDCARD-SCOPE — stripe.create_refund declares a broad action scope
The top blocker is not "a new tool appeared". It is that a tool which moves
real money can be called with no declared approval gate, under a stripe:*
scope the manifest never granted. The published v0.15.0 build names the same
four check IDs in the same order; two of the messages are worded differently
("adds destructive capability without rollback controls", "adds financial write
capability without required controls"). Match on the check ID, not the
sentence. Each finding names the evidence it rests on —
in this fixture, the MCP export at tools.json#/tools/1. Run
agents-shipgate explain SHIP-POLICY-APPROVAL-MISSING for the check's own
description, or see checks.md for the catalog.
On current main, capability summaries distinguish conservative effect
projections from their static evidence. A provisional write and unknown
effect evidence are compatible answers; see effect projections and
evidence for the existing JSON fields and review meaning.
report.md carries its own coverage limit, and that limit is part of the
answer:
Evidence coverage: static (2/2 catalog tools reachable; 1 semantic review
concern(s); 2/2 actions pass-eligible; human review recommended)
On the published v0.15.0 build that line carries no counts — it reads
Evidence coverage: static (human review recommended). The limit is stated;
what it is a limit on is not enumerated. That enumeration is the part of this
step you need a newer build for.
The same boundary is repeated in pr-comment.md:
- Static-verdict boundary: This verdict covers deterministic static evidence
only. Agents Shipgate did not execute the agent or prove runtime behavior,
tool routing, credential enforcement, or safety.
v0.15.0 prints that sentence nowhere — not in its comment, not in its
artifacts, not on stdout. What it carries instead is the ## Disclaimer
section at the foot of report.md, which makes the same point in the report
rather than beside the verdict: "Runtime behavior, actual tool routing, and
output interpretation are not verified."
Nothing here proves the refund tool behaves as described at runtime, that the credential is really scoped, or that a human is really in the loop. It proves what the declared and statically discoverable surface says, and it names the one semantic concern it could not close.
When static extraction cannot enumerate a tool surface at all — a toolkit
factory, a config-bound allowlist, tools assembled at runtime — the release
decision is insufficient_evidence rather than a guess. That is the intended
failure mode. It routes to a person; it does not claim the agent is unsafe.
Ask the shell what the run returned:
agents-shipgate fixture run ai_generated_refund_pr > /dev/null; echo $?0
report.md says why:
Fail policy: ci_mode=advisory, fail_on=[none], new_findings_only=false, would_fail_ci=false (exit 0)
Advisory mode never fails the job. The verdict is blocked anyway. Exit
status is a CI policy choice; the verdict is the answer. A gate wired to the
process exit code, in advisory mode, gates on nothing. Advisory
CI below is where you make the verdict load-bearing.
verifier.json carries fix_task. One of its five instructions — an
evidence-gap line — and the allowed_repairs[] array beside them are elided
here:
{
"actor": "human",
"safe_to_attempt": false,
"instructions": [
"4 active findings block release.",
"Declare an approval policy for stripe.create_refund or remove this tool from the release.",
"Declare approval.required, confirmation policy, and safeguards.rollback for this destructive action.",
"Declare approval.required, safeguards.audit_log, and safeguards.idempotency for this financial write action."
]
}actor: "human" is the operative field, and safe_to_attempt: false says the
same thing to a machine. Every instruction is a declaration about approval,
control or safety — a claim only a person can make. A coding agent may report
this, and may not resolve it. The published v0.15.0 build carries the same
fix_task with the same actor and safe_to_attempt; its instruction list
differs in wording and adds one about replacing the wildcard scope. This step
reads the same on either channel.
On a build that implements contract 21 or newer, the same answer arrives as
an explicit state machine. Validate the pointer before you read anything it
binds — current-control.json is an ordinary file and goes stale silently.
Run it against the workspace and report directory the fixture printed, not
against your own:
cd /tmp/shipgate-fixture-ai_generated_refund_pr-<random>/ai_generated_refund_pr
agents-shipgate agent control --workspace . --reports-dir reportsThe first path is the Fixture copy left at line from step 2; reports is the
last segment of its Reports: line. Both flags matter: the fixture's history
lives in that copy, and its artifacts are under reports/, not the default
agents-shipgate-reports/. Run it from your own directory instead and it exits
3 with Current control is unavailable (missing) — or, worse, validates an
unrelated run you happen to have lying there.
It exits 0 here and returns the compact shipgate.agent_control/v1
envelope, whose fields are top-level:
"control_state": "human_review_required",
"next_actor": "human",
"decision": "blocked",
"permissions": { … all false … }
That command checks the pointer against every artifact it binds and against
the live repository, and exits 4 when the workspace has moved since the
decision (workspace_changed, naming the commit the pointer was published
against). Reading the file directly cannot tell you that — after an unrelated
commit it still reports the old state. Only once agent control succeeds do
you read current-control.json — whose own spelling is nested, control.state
rather than control_state — and then agent-handoff.json (control.state,
then gate.merge_verdict).
On the published v0.15.0 build none of this exists: no agent control
command, no pointer, no control block. Read agent-handoff.json's gate
block and report.json's release_decision.decision instead. See
Which build you get.
Where a human-review route still lets an agent work. Contract v20 separates
two states. human_review_required is a full stop. review_publishable means
the agent may still commit, push and update the pull request so the review can
happen — merge and completion stay denied. Updating a pull request is not
merging it.
There is no in-repo approval mechanism. A comment, a human_ack added by
the same change, or an acknowledgement in conversation does not clear a
human-review route. A trusted coding host can sign an external authorization
for one exact command; that protocol is described in
agent-contract-current.md,
Agents Shipgate ships no signing or approval command, and it does not turn the
verdict into a pass.
Findings split cleanly, and the split is in the artifact rather than in your judgement.
Reversible, containment-checked, and safe to run without further approval:
agents-shipgate apply-patches \
--from agents-shipgate-reports/report.json \
--confidence high --applyAt --confidence high this fires on exactly three stale-manifest removals
(SHIP-MANIFEST-STALE-{SUPPRESSION,POLICY,RISK-OVERRIDE}) and refuses to
mutate anything outside the manifest's directory. It is dry-run by default.
Scope-coverage appends require an explicit --confidence medium and a reviewer.
Adding advisory CI and adding agents-shipgate-reports/ to .gitignore are
mechanical too.
An agent must never supply, infer or auto-assert an action's effect, its
authority, its bindings, the agent's purpose, or an approval,
confirmation, idempotency, broad-scope, or prohibited-action claim — including
by reading them out of a prompt, a README, or a docstring. Prose is not
evidence of authority. agent.declared_purpose and every value like it must be
supplied by a human, because Shipgate never invents a declaration nobody made;
init --json and fix_task mark each one with actor: "human", naming the
fields and lines. The counterpart matters as much: agent.name and the
tool_sources[] rows are facts in the repository, and an agent that escalates
those to a person has stopped a turn for nothing.
agent-autofix-boundary.md is the full boundary,
with the check-ID mapping and the exact phrasing an agent should hand to a
reviewer. autofix-policy.md is the mechanical filter
behind apply-patches.
Two routes. Pick by what the repository is, not by which looks lighter.
For any repository that declares what its coding agents may do: .mcp.json,
.claude/settings.json, .codex/, .cursor/, hooks, workflow scopes. There is
no manifest and no policy authoring.
shipgate audit --host --json --out agents-shipgate-reports/host-grants.jsonThe default audit is deterministic repository scope. Use --scope local-static only when you explicitly want supported user and file-based
managed configuration included. Both modes publish per-host coverage and
excluded sources; neither proves session or runtime authority. Inspect
host_coverage, issues and excluded_scopes before relying on the result,
and see the static host-boundary support matrix.
Record a baseline once, on the default branch, then check drift per change:
shipgate audit --host --save-baseline # then commit .agents-shipgate/
shipgate audit --host --drift --json --out shipgate-drift.jsonNever --save-baseline from the changed checkout to get past a missing
baseline. That acknowledges the very expansion under review, and the drift
then reports nothing.
Before a coding agent reports an agent-capability change complete, it runs the local boundary check:
shipgate check --agent codex --workspace . --format agent-boundary-json
shipgate check --agent claude-code --workspace . --format agent-boundary-json
shipgate check --agent cursor --workspace . --format agent-boundary-jsonThe --agent value is caller identity, not a coverage selector: every
recognized changed host boundary is evaluated on every invocation. Parse the
stdout shipgate.agent_boundary_result/v2 object, switch on control.state,
and follow control.next_action, control.allowed_next_commands and
control.human_review. Treat decision as diagnostic context, never as the
control signal, and never infer control from prose. check is necessary but
not sufficient for a capability-expanding diff: if a change adds dynamic,
undeclared or otherwise ambiguous tool capability, do not read
decision="allow" as merge readiness — run verify and read
release_decision.decision. On the published v0.15.0 build this command
accepts only --format codex-boundary-json and emits no control block; see
Which build you get.
For repositories that define and declare a tool surface in-repo: MCP or OpenAPI exports, framework tool definitions, a Codex plugin package.
agents-shipgate verify --preview --json # what would be read, before anything is written
agents-shipgate detect --json # classify the workspace
agents-shipgate init --write --ci --json # manifest + advisory workflowinit --write --ci produces a schema-valid shipgate.yaml with
framework-specific tool_sources populated, plus a GitHub Actions workflow.
Then verify the change:
# local, uncommitted work — omit --base/--head so working-tree edits are scanned
agents-shipgate verify --workspace . --config shipgate.yaml \
--ci-mode advisory --format json
# committed PR/CI refs — make the base ref available first; verify never fetches
agents-shipgate verify --workspace . --config shipgate.yaml \
--ci-mode advisory --format json --base origin/main --head HEADThe short shipgate verify alias remains invokable for compatibility;
agent-facing PR-gate guidance uses agents-shipgate verify.
init reports what it could not infer, and routes the whole turn.
Two fields, and only one of them decides. placeholders[] is a location
list — each entry is path, current and line, with no owner on it.
control.next_action.actor routes the next step rather than assigning
ownership to individual fields. A human-owned placeholder normally selects
"human" with permissions.edit: false, and why names where to start, like
In ./shipgate.yaml: line 13 (agent.declared_purpose[0]) must be supplied by a
human. These fields declare what this agent is for and what it is permitted to
do; Shipgate never invents them, and a value a coding agent supplied is a
declaration nobody made.
That sentence is a starting point, not a partition. It is fitted to the
envelope's prose budget: with seven unresolved declarations it names three
paths and then and 4 more in placeholders[], and the four it dropped are
human-owned as well. While actor is "human", surface the required review
and stop. When it is "coding_agent", perform only the exact
control.next_action, using its kind, path or command, then rerun the
stated check. A blocking setup repair can take precedence while human-owned
declarations remain unresolved. An agent route does not prove every remaining
placeholder belongs to the agent.
Underneath the routing, the ownership split is real in both directions:
- A coding agent owns the facts it can read out of the repository:
agent.name,project.name, and thetool_sources[]rows. Sending these to a person stops a turn for work the agent owns. - A person owns every declaration: purpose, prohibited actions, effect,
authority, binding, approval, confirmation, idempotency, safeguards, accepted
debt and its owner/reason/expiry, and the blocks that are declarations end to
end (
action_surface,permissions,policies,agent_bindings,tool_identity,checks,baseline,human_ack,risk_overrides,validation,organization). These values must be supplied by a human — Shipgate never invents a declaration nobody made. Their review route remains due after any blocking setup repair and returnscontrol.next_action.actor: "human"withpermissions.edit: false.
Those names illustrate the rule; control.next_action.actor is what decides a
given turn, and it is the field to act on. Neither list nor prose overrides it.
Do not let a coding agent fill a human-owned value from a prompt, the main agent file or the repository README. A purpose or authority claim lifted out of prose is not a guess to be corrected later; the engine will treat it as evidence, which is exactly the gap this tool exists to catch.
When discovery reads no tool surface at all. init --json reports
tool_surface_origin: "detected" when every source in the manifest was read
out of the workspace, "scaffold" when none was, and null when this run's
render reached neither disk nor the payload. On "scaffold" the tool_sources
block is a placeholder — id, type and path are all CHANGE_ME, all three
are in placeholders[], and manifest_message says nothing was inferred. That
can happen when registration depends on a server factory that static
discovery cannot resolve.
A scaffold says discovery could not read a surface — not that there is none,
and not that you are on Route H. This checkout detects a direct
server = FastMCP(...) with @server.tool functions, including under an import
package. A server created through app = create_server() with @app.tool
functions still returns "scaffold": the static reader cannot resolve that
factory. audit --host reviews coding-host configuration and never reads those
tools. Switching routes there
stops reviewing the capability you came to review. Stay on Route A and give
verify something it can read — an exported MCP tool list, an OpenAPI spec, or
a local tool inventory (see Choose your first
source) — or report that static extraction cannot
reach this surface. Switch to Route H only when the thing you actually want
reviewed is the coding-host configuration.
| Symptom | Next action |
|---|---|
detect says is_agent_project: false, but suggested_sources includes MCP or OpenAPI files |
Proceed to init. MCP/OpenAPI-only repos are valid tool-surface targets even without Python framework detection. |
detect says is_agent_project: false, but codex_plugin_candidates is non-empty |
Proceed to init. Codex plugin repos are valid static plugin-surface targets. |
doctor shows zero tools |
Check tool_sources[].path, MCP tools[], OpenAPI paths, optional source warnings, and dynamic ADK/MCP toolsets. |
| Tools are created by factories, wrappers, or dynamic toolsets | Provide an explicit MCP export, OpenAPI spec, local tool inventory artifact, or a broader OpenAI SDK source directory when tools are static but split across files. |
init --write --json returns placeholders[] entries |
Switch on control.next_action.actor, not the array. "human" means surface all of them and stop — why names where to start and says when it elided more. See Route A. |
| Install fails in a Python 3.10/3.11 project | Install the CLI outside the project env with pipx or uv using Python 3.12+. |
Reports appear in git status |
Add agents-shipgate-reports/ to .gitignore; reports are local release-review artifacts. |
Point the manifest at the clearest tool boundary you already have:
| Source | Use when | Manifest path |
|---|---|---|
| OpenAI Agents SDK Python | Tools are defined with @function_tool in local Python files. |
tool_sources[].type: openai_agents_sdk; path may be one Python file or a directory |
| MCP export | You can export the MCP server's tool list to JSON. | tool_sources[].type: mcp |
| MCP server source | This checkout can read the static TypeScript, Go or Python registration shape, including direct FastMCP @server.tool decorators. |
tool_sources[].type: mcp_server_source |
| OpenAPI spec | The agent calls HTTP APIs described by OpenAPI 3.x. | tool_sources[].type: openapi |
| Codex plugin package | The repo contains .codex-plugin/plugin.json or .agents/plugins/marketplace.json. |
tool_sources[].type: codex_plugin |
Google ADK, LangChain/LangGraph, CrewAI, n8n workflow JSON, Conductor OSS
workflow JSON, Anthropic Messages API artifacts, and simple OpenAI API
artifacts are supported inputs too. Start with the tool surface closest to the
release boundary. Framework-by-framework minimal manifests are in
minimal-real-configs.md; agent-driven recipes are
in agent-recipes.md.
The release gate is report.json → release_decision.decision:
| Decision | Meaning | Next action |
|---|---|---|
blocked |
Active, unaccepted blockers exist. | Fix blockers or remove the risky tool surface. |
insufficient_evidence |
The scan cannot confidently gate release from the available static evidence; this does not prove the agent is unsafe. | Follow the first structured evidence-gap action. Supported frameworks name the generated local inventory and exact manifest route; unidentified source shapes receive the generic source guidance. Then rerun. |
review_required |
Human review is needed for accepted debt or evidence gaps below the blocked threshold. | Review the listed items before promotion. |
passed |
No active blocker or review signal was found. | Keep the report artifact with the PR/release record. |
merge_verdict is the PR-facing projection of that decision, on verifier.json
and agent-handoff.json:
| Merge verdict | Meaning | Next action |
|---|---|---|
blocked |
Active, unaccepted blockers exist. | Fix blockers or remove the risky capability. |
insufficient_evidence |
Static evidence is too weak to gate release confidently. | Add better sources and rerun; do not auto-merge. |
human_review_required |
A person must review accepted debt, trust-root changes, or authority-bearing gaps. | Surface the required review; a coding agent must not self-approve it. |
mergeable |
No active blocker or review signal was found. | Keep verifier/report artifacts with the PR record. |
unknown |
Verify could not produce a reliable head scan or diff context. | Fix the setup, fetch the base ref, or rerun with usable inputs. |
Common review signals: missing confirmation, missing idempotency evidence, broad-scope permissions, prohibited-action policy gaps, and trust-root changes such as weakened CI or manifest policy.
The first run is setup. The second is the one that tells you whether this is worth keeping.
Nothing from the first run has to be redone. On the next change that touches tools, prompts, scopes, policy or the gate itself, run the same command with the PR's refs:
agents-shipgate verify --workspace . --config shipgate.yaml \
--ci-mode advisory --format json --base origin/main --head HEADOn Route H, the same shape: shipgate audit --host --drift against the
baseline you already committed, then shipgate check. Not every PR needs a
run — triggers.json is the machine-readable rule set for
deciding, and verify --preview --json answers it for one workspace.
Drop this into .github/workflows/agents-shipgate.yml. It runs on every PR,
posts the verdict as a comment, uploads artifacts, and never fails the job:
name: Agents Shipgate
on:
pull_request:
permissions:
contents: read
pull-requests: write
jobs:
agents-shipgate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- id: shipgate
uses: ThreeMoonsLab/agents-shipgate@v0.15.0
with:
config: shipgate.yaml
ci_mode: advisory
diff_base: target
pr_comment: "true"
shipgate_version: "0.15.0"The action delegates to verify and never fetches — keep fetch-depth: 0.
Advisory mode is where you start, not where you stop: as § 6 showed, an
advisory job exits zero on a blocked verdict. Make the verdict load-bearing
with an explicit policy step:
- name: Block only blocked verdicts
if: steps.shipgate.outputs.merge_verdict == 'blocked'
run: exit 1- name: Require no human authority gap
if: steps.shipgate.outputs.can_merge_without_human != 'true'
run: exit 1GitLab, CircleCI, Jenkins and pre-commit equivalents, the full output catalog,
and the other merge policies are in integrations.md and
examples/github-actions/.
Strict mode exits 20 on unsuppressed critical findings. On an existing
project, save a baseline first so strict CI fails only on new findings:
agents-shipgate baseline save --config shipgate.yaml --out .agents-shipgate/baseline.json
agents-shipgate scan --config shipgate.yaml --baseline .agents-shipgate/baseline.json --ci-mode strictbaseline.md covers baseline lifecycle, suppression hygiene and
noise triage.
Assessing someone else's repository without adopting Shipgate into it is an
explicit opt-in. Point preview at the reserved manifest so it emits the
local-review setup route instead of the durable init --write route:
agents-shipgate verify --preview --workspace . \
--config .agents-shipgate-local-review.yaml --json
agents-shipgate init --workspace . --local-review --json
agents-shipgate verify --workspace . \
--config .agents-shipgate-local-review.yaml \
--base origin/main --head HEAD --json--local-review writes an ephemeral manifest at the workspace root so its
relative source paths still resolve, and privately excludes that file plus
agents-shipgate-reports/ through .git/info/exclude. Tracked project files
stay untouched and generated reports stay out of git status. The JSON result
enumerates every effect plus an executable cleanup command;
init --local-review --undo --json removes those local setup effects while
preserving reports.
Verification marks the reserved manifest as local_review and always keeps the
result provisional. The same fail-safe applies to any differently named
uncommitted or Git-unproven manifest: passed, mergeable, and
human-authorization evidence all require a manifest that Git proves is present
in the evaluated repository tree. Git presence alone does not prove human
review. Durable adoption remains init --write.
For a design-partner review or a false-positive report, export the small redacted feedback artifact:
agents-shipgate feedback export \
--from agents-shipgate-reports/verifier.json \
--redact \
--out shipgate-feedback.jsonThe export includes the merge verdict, top capability changes, finding IDs,
next action, fix_task, and reviewer prompts. It does not include raw finding
evidence.
agent-recipes.md— copy-pasteable AI-agent workflows for verify-first PRs and first adoptionagent-autofix-boundary.md— what an agent may fix, and what it must never assertminimal-real-configs.md— framework-by-framework minimal manifest referencesmanifest-v0.1.md— manifest schema in prose formchecks.md— what the scanner looks formental-model.md— the five-minute version of the decision modelcategory.md— what an "agent release gate" is, and what it is notfaq.md— common questionstroubleshooting.md— when a run does not do what you expectedglossary.md— category vocabulary