ci: bump actions/setup-python from 5 to 7 - #3
dependabot[bot] wants to merge 5 commits into
Conversation
A research prototype of a control architecture for maintaining behavioural, authority and operational integrity across autonomous AI-agent ecosystems. Research question: can an agent ecosystem detect and contain compromised agents without transferring unrestricted authority to the agents responsible for defence? Architecture: - control plane (identity, capabilities, authority, task contracts, policy, risk, delegation, containment, revocation, emergency controls, hash-chained audit, execution gateway) as the single enforcement boundary - behavioural observatory (telemetry, behavioural profiles, agent identity and behaviour graphs, deterministic anomaly detectors, authority drift engine) - immune system (threat classification, bounded-prior immune memory, verification ledger, graduated response planning) - defensive agents with disjoint capability sets: DETECT, AUDIT, CONTAIN, FORENSICS, RECOVERY - quarantine sandbox with deterministic replay - simulated ecosystem, named experiments, metrics and benchmark harness Enforced invariants: - authority is external: self-grant raises SelfGrantDenied for every principal - delegation attenuates: EffectiveAuthority(B) subset of DelegableAuthority(A), with depth attenuation, cycle detection and attenuated sub-contracts - detection does not grant containment authority; containment requires an independent, subject-bound verification re-derived from the audit chain - offensive capabilities are unrepresentable at parse time - emergency controls are out-of-band and unreachable from any agent - default deny for anything no policy matches Results (20 steps, 213 decisions, deterministic): detection 1.00, containment false-positive rate 0.00, mean time to detection 2.17 steps, 18/18 unauthorized actions prevented, 33/33 authority amplification attempts blocked, 0/11 defender escape attempts successful, 12/12 section 17 escape paths blocked, audit chain intact, identical audit head hash across runs. Stdlib only. 114 tests.
Research framing - evaluation.md rewritten: explicit closed-world assumptions, what the metrics do and do not claim, 13 falsifying experiments with the test that reports each, 4 open experiments that are not yet implemented, 8 ablations - claims-map.md: every claim -> mechanism -> code -> test -> metric - research-thesis.md: one-page invariant -> research question -> metric map - getting_started.md: run it, add a policy, write a scenario, read the chain - results.json committed with metrics, audit head hash and hard-check block Adaptive and relationship-level adversaries - SLOWDRIFT-01: four-phase drift over 48 steps; detection latency is the result (first probe step 18, first verified flag step 24) - INSIDER-01 / COURIER-01: split-capability collusion where neither agent exceeds its own authority; detected by relationship analysis with capability complementarity and answered with ESCALATE rather than containment - compromised detector fabricating evidence against an innocent agent - decision-time race experiment: restriction, expiry and revocation all take effect on the next submission, with no grace window - quarantine escape suite: nine bypasses, each with the mechanism that stops it - benchmark sweeps: observation loss, defensive cycle period, population scale Enforcement - InvariantChecker re-checks P1-P8 as state predicates after every decision; detects an in-process gateway bypass that no policy can prevent - multi-party attestation: revocation requires two distinct verifiers (POL-110) - context-bound grants: a grant scoped to a task contract is unusable elsewhere - POL-060 escalates acquire/transfer/exfiltrate steps while a collusion chain is open, without touching either agent's authority - HOLD and ESCALATE no longer count as agent violations, so escalation cannot manufacture the evidence used to justify containment Tests and tooling - stdlib property-test harness (no third-party dependency) with reproducible seeds; properties for delegation attenuation, self-grant denial, revocation cascade, default deny, defensive containment gates - 148 tests (from 114), including the adaptive suite and invariant tests - CI workflow makes the defender-escape and quarantine-escape experiments build gates and asserts the benchmark hard checks - trace tooling: JSON timeline, Graphviz DOT, single-file HTML report; ActionResult.narrative() renders one decision as a sentence - demo rewritten with a narrative and a --trace mode Results (20 steps, 213 decisions, 368 invariant checks): detection 1.00, containment false positives 0.00, prevention 18/18, amplification blocked 33/33, defender escapes 0/11, invariant violations 0, audit head hash identical across runs. Under 90% telemetry loss detection falls to 0.83 while prevention stays 1.00 and containment false positives stay 0.00.
SENTINEL-AUDIT was logically independent - it read only the audit chain - but shared address space with the components it audited. It can now run as its own OS process that receives a serialised audit chain and returns a signed verdict. New package ais/verifier/ - reconstruction.py: pure, chain-only evidence derivation. Recomputes the hash chain, derives each subject's grants, revocations and restrictions from authority.* events, rebuilds cross-agent collusion chains, and decides a verdict against the shared requirement table. The in-process auditor now uses the same module, so the two transports cannot drift apart. - service.py: the worker process (python -m ais.verifier.service). No control plane, no authority service, no observatory, no gateway - a chain and a claim. - protocol.py: newline-delimited JSON, HMAC-signed in both directions with operator key material that never passes through a gateway request. - client.py: owns the child process, binds request id and nonce, verifies the response signature, and fails closed on timeout, crash or bad signature. SENTINEL-AUDIT-REMOTE (agent.observe, agent.analyze, immune.read) ships the chain once per audit length and records verdicts through the verification ledger with transport="cross-process". Sentinel(remote_verification=True) adds it to the cycle; POL-110 counts it as a distinct verifier, so the two verifiers required for revocation can now live in two processes. Ten checks, all passing (verification/cross-process): - verdicts agree with the in-process auditors, from a different pid - a doctored chain supplied by the caller is REJECTED at the broken index - a truthful chain with a falsely claimed head is REJECTED - a false claim about an innocent agent is INCONCLUSIVE with unsatisfied predicates named; an empty chain cannot confirm anything - a verdict signed with another key is rejected; an unsigned request is refused - a crash respawns and still refuses an unsupported claim: respawning is not a laundering route - an outage fails closed: no verdict, verification.unavailable audited, containment stays at HOLD, victim authority unchanged Cost: 9.8 ms mean round trip for a 124-record chain against 2.7 ms for the same reconstruction in-process. Also - confidence now counts amplification persistence as well as breadth, so a delegation attacker that retries is not scored like a single probe - defender metrics split: "policy violations" means acting outside declared authority (0), while a containment request held for insufficient evidence is reported separately as defensive_actions_held_for_evidence - 171 tests (from 148); CI gains a cross-process job; docs updated across architecture, threat model (A14-A16), evaluation, claims map and thesis; the remaining gap is stated plainly: same host, same user
Learning adversary (adaptive/learning-adversary)
An adversary that observes the outcome of its own requests - which any real
agent can - and adapts: it never retries a refused capability, widens its probe
interval after each refusal, dilutes refusals with in-contract activity, and
when every probe has been refused it stops probing and pursues the objective
through capability it actually holds.
Three arms, identical configuration otherwise:
naive learner learner (undefended)
retries of denied capabilities 28 0 0
denial ratio 0.45 0.05 0.36
steps to verified detection 2 9 never (cycle disabled)
real unauthorized effects 0 0 0
prevention rate 1.00 1.00 1.00
invariants clean clean clean
Four findings, reported in order of how uncomfortable they are:
- adaptation works, and costs the defence 7 steps of detection latency;
- adaptation buys latency, not authority: prevention stays at 1.00 and real
production effects stay at 0 in every arm;
- given a long horizon the adversary converges to acting inside its grant, at
which point the relationship layer produces a score-1.0 collusion signal and
POL-060 escalates the chain at the gateway, with the defensive cycle disabled
- escalation is a policy decision, not a SENTINEL action;
- quarantine deceives it: the sandbox answers its probe with synthetic data, so
it records success and stops probing while producing no real effect.
Ground truth for "did it succeed" is computed outside the agent, from decisions
that reached the production registry.
Trace report
Rebuilt the HTML report with a light, modern interface: status pills for chain
and policy integrity, metric cards, containment and drift above the decision
log, decision badges, evidence shown where it decides something (refusals),
client-side row filtering, sticky headers, responsive layout. Still a single
self-contained file with no external assets.
Screenshots in docs/assets are generated from real runs by
scripts/make_screenshots.py (headless Chromium); the terminal images contain
actual stdout, nothing is mocked up.
Project structure
pyproject.toml, CITATION.cff, SECURITY.md (ten in-scope findings, each mapped to
a falsification entry), CODE_OF_CONDUCT.md, CHANGELOG.md, Makefile, .editorconfig,
.gitattributes, CODEOWNERS, dependabot for CI actions, three issue templates
(bug, broken invariant, experiment proposal) and a pull request template that
asks which invariant a change protects. README rewritten with badges, the three
screenshots, the results table and a claims-first reading order.
180 tests (from 171). results.json is now tracked as a research artefact.
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 5 to 7. - [Release notes](https://github.com/actions/setup-python/releases) - [Commits](actions/setup-python@v5...v7) --- updated-dependencies: - dependency-name: actions/setup-python dependency-version: '7' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
LabelsThe following labels could not be found: Please fix the above issues or remove invalid values from |
ECC Tools / PR Config AuditCommit: PR audit blocked before scanning (action_required) Private repository analysis requires Pro or Enterprise. No audit was performed. Check installation entitlement and usage with ECC Tools support. Check publication was denied or unavailable. An app owner must enable Checks: read and write, and the installation owner must approve the updated permission. |
a9b4226 to
8abc3bd
Compare
|
OK, I won't notify you again about this release, but will get in touch when a new version is available. If you'd rather skip all updates until the next major or minor version, let me know by commenting If you change your mind, just re-open this PR and I'll resolve any conflicts on it. |
Bumps actions/setup-python from 5 to 7.
Release notes
Sourced from actions/setup-python's releases.
... (truncated)
Commits
5fda3b9Pin SHA commits and update docs with latest versions (#1338)4ab7e95Merge pull request #1337 from actions/philip-gai/bump-actions-cache-6-2-00f3a009Remove the pip-install input (#1336)f8cf429Migrate to ESM and upgrade dependencies (#1330)54baeeaValidate and retry manifest fetch to prevent silent failures (#1332)c709277Annotation code fix (#1335)6849080remove EOL Python versions and Bumps numpy text fixture (#1333)0903b46Bump certifi from 2020.6.20 to 2024.7.4 in /tests/data (#1328)ece7cb0Fix pip cache error handling on Windows. (#1040)1d18d7aUpdate advanced-usage.md (#811)Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting
@dependabot rebase.Dependabot commands and options
You can trigger Dependabot actions by commenting on this PR:
@dependabot rebasewill rebase this PR@dependabot recreatewill recreate this PR, overwriting any edits that have been made to it@dependabot show <dependency name> ignore conditionswill show all of the ignore conditions of the specified dependency@dependabot ignore this major versionwill close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this minor versionwill close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this dependencywill close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)