diff --git a/.claude/notes/agents.md b/.claude/notes/agents.md index b36815af..74f486bc 100644 --- a/.claude/notes/agents.md +++ b/.claude/notes/agents.md @@ -434,7 +434,10 @@ agent had produced; bounding it only by a fixed cycle count disconnected from `t path unreachable — the watchdog always wins, and the same spurious-orphan turn burns the full turn timeout before crashing with zero criteria evaluated. -The cycle cap is the SOLE bound when a task sets no timeout at all. It is deliberately not +The cycle cap applies alongside that deadline, whichever comes first, and is the SOLE +bound when a task sets no timeout at all. Under a long `turn_timeout` (1800 s gives a +1440 s deadline) the deadline alone let a never-ending command (a dev server, an +unanswered prompt) idle the turn for 24 minutes. It is deliberately not "break after N consecutive empty polls": `receive_steps()` returns identically empty whether a backgrounded job is still running or will never resolve, and there is no signal that tells the two apart except waiting. A count small enough to matter would abort real @@ -447,6 +450,36 @@ answer in a headless eval), CANCELED and UNKNOWN — none of which the closed se done, and none of which the poll loop should wait out, since they will never become DONE on their own. +It is also an allowlist on the TOOL: only a `run_command` can be backgrounded. Any other +tool left ACTIVE never resolved in practice: an `edit_file` on the read-only skill mount, +a `view_file` paged at a large `content_offset`, a `start_subagent`. In a 370-task +SkillSpec run, 12 of 21 turns that polled to the deadline were waiting on one of those, +after the model had already finished. `finalize` force-closes them as unresolved. + +## Antigravity tool allowlist + +`allowed_tools` / `disallowed_tools` map onto the harness's builtin toolset +(`CapabilitiesConfig.enabled_tools` / `disabled_tools`) through +`_CLAUDE_TO_ANTIGRAVITY_TOOL_MAP`, so an experiment's allowlist confines Antigravity the +way it confines Claude Code. Without it Antigravity ran with every builtin, including +`start_subagent` and `search_web`, under a list that gave Claude Code neither. `Skill` and +`TodoWrite` have no builtin (skills load through `skills_paths`) and are skipped. `finish` +stays on under any allowlist: it returns structured output, not a capability. +`enable_subagents` is a separate switch from the toolset and follows whether +`start_subagent` survives. The SDK takes an allowlist OR a denylist, so with both set the +denied tools are removed from the allowlist. + +## Antigravity non-interactive commands + +`run_command` runs in a real terminal: a command that blocks becomes a background task the +model can check on and type into. A command that stops to ask (`npx` installing a missing +package, git credentials, an apt/pip confirmation, a pager) and that the model then leaves +behind keeps the turn waiting on a prompt no one answers. `_NONINTERACTIVE_ENV` +(`CI`, `npm_config_yes`, `GIT_TERMINAL_PROMPT=0`, `DEBIAN_FRONTEND`, `PIP_NO_INPUT`, +`PAGER`/`GIT_PAGER=cat`) rides the per-agent `env` seam, each variable only where the +environment does not already set it, so those commands answer themselves or fail fast. It +cannot close the terminal: a bare `read` or a server still runs until the poll cap. + ## The receive_steps re-entrancy window `receive_steps()` is two nested async generators: the public one delegates to the diff --git a/docs/agents/ANTIGRAVITY.md b/docs/agents/ANTIGRAVITY.md index ccc37db1..ad6aefba 100644 --- a/docs/agents/ANTIGRAVITY.md +++ b/docs/agents/ANTIGRAVITY.md @@ -27,7 +27,7 @@ working directory — both required for an unattended eval. pip install 'coder-eval[antigravity]' ``` -This pulls in `google-antigravity` (pinned to `0.1.18`), whose wheel bundles the +This pulls in `google-antigravity` (pinned to `0.1.20`), whose wheel bundles the platform `localharness` binary. As with the other agents the SDK is imported lazily — a base install without the extra still runs end-to-end; Antigravity tasks fail at dispatch with a clear hint to install the extra. @@ -136,18 +136,28 @@ can't be resolved or if zero skills are discovered. ## Permissions & tools — important differences -**Antigravity ignores `permission_mode`, `allowed_tools`, and `disallowed_tools`.** -The local harness runs in a single unconditional mode: every tool call (including -`run_command`) is approved via an allow-all policy, and file tools are restricted to -the configured `workspaces` (the sandbox working directory plus any skill roots). +**Antigravity ignores `permission_mode`.** The local harness runs in a single +unconditional mode: every tool call (including `run_command`) is approved via an +allow-all policy, and file tools are restricted to the configured `workspaces` (the +sandbox working directory plus any skill roots). + +**`allowed_tools` / `disallowed_tools` are enforced** by mapping the Claude tool names +onto the harness's builtin tools (`Bash` → `run_command`, `Read` → `view_file`, `Write` → +`create_file`, `Edit` → `edit_file`, `Glob` → `find_file`, `Grep` → `search_directory`, +`Task` → `start_subagent`, `WebSearch` → `search_web`, `WebFetch` → `read_url_content`). +`Skill` has no builtin and is skipped; `finish` always stays on. + +`run_command` also gets non-interactive environment variables (`CI=1`, +`npm_config_yes=true`, `GIT_TERMINAL_PROMPT=0`, `DEBIAN_FRONTEND=noninteractive`, +`PIP_NO_INPUT=1`, `PAGER=cat`), each only where the environment does not already set it. The trust boundary for an Antigravity run is therefore the **sandbox**, not the agent config. Run untrusted tasks under the [Docker driver](../DOCKER_ISOLATION.md); the `tempdir` driver is not a security boundary. This mirrors the reality that the `bypassPermissions`-equivalent behavior is always on for this backend. -Those inherited fields still exist on the config for schema uniformity but have no -runtime effect here — don't rely on them to gate Antigravity. +`permission_mode` still exists on the config for schema uniformity but has no runtime +effect here — don't rely on it to gate Antigravity. ## Telemetry @@ -187,19 +197,18 @@ as every other agent. 4. **`permission_mode` does not confine the harness.** Every mode runs `policy.allow_all()`; coder_eval's write boundary is the sandbox driver, and a headless eval has no human to approve anything. -5. **`allowed_tools` / `disallowed_tools` are not read.** The harness runs with its - full builtin tool set, so an Antigravity run has tools (web search, subagents, - URL fetch) that the same task file denies on Claude Code and Codex. -6. **`max_turns` is counted by the harness.** One `communicate()` is a single SDK turn +5. **`max_turns` is counted by the harness.** One `communicate()` is a single SDK turn here, so the harness counts model API calls itself (a MODEL step at a new `step_index` opens one) and enforces the cap on the step loop. See [Run-Limit Parity](HARNESS_PARITY.md). -7. **Shell commands over ~10s are moved to the background.** The localharness has a +6. **Shell commands over ~10s are moved to the background.** The localharness has a 10-second maximum synchronous wait; past it the command becomes a background task and the model gets a task id, not a result. The turn polls for that result instead - of finalizing on an idle step stream, so slow work does complete — but the wait is - bounded by 80% of `turn_timeout`, and a job that outlives it is force-closed as - `result_status: unknown` and graded as an ordinary low score rather than a timeout. + of finalizing on an idle step stream, so slow work does complete — but only an + orphaned `run_command` is waited on, the wait is bounded by 10 minutes or 80% of + `turn_timeout` (whichever is shorter), and a job that outlives it (typically a server + the model left running) is force-closed as `result_status: unknown` and graded + normally rather than as a timeout. Measured in [Run-Limit Parity](HARNESS_PARITY.md). ## Running in Docker diff --git a/docs/agents/HARNESS_PARITY.md b/docs/agents/HARNESS_PARITY.md index 86fb07bb..e3714df2 100644 --- a/docs/agents/HARNESS_PARITY.md +++ b/docs/agents/HARNESS_PARITY.md @@ -718,9 +718,10 @@ both. See [OpenCode](OPENCODE.md) and [Pi § plugins](PI.md#known-limitations). lowercase (`bash`/`read`/`write`/`edit`/`grep`/`find`/`ls`), but the shared config default (`experiments/default.yaml`) sets Claude-namespaced names (`Bash`/`Read`/`Write`/…). Forwarding those to `--tools` would allowlist tools that - do not exist in Pi and strip the agent of ALL tools — so, like OpenCode (drops them), - Codex (forwards `disallowed_tools` without SDK enforcement), and Antigravity (does not - read them), Pi ignores them and runs with its full native toolset. A task that needs a + do not exist in Pi and strip the agent of ALL tools — so, like OpenCode (drops them) + and Codex (forwards `disallowed_tools` without SDK enforcement), Pi ignores them and + runs with its full native toolset. (Antigravity maps them onto its builtin tools; see + [Antigravity](ANTIGRAVITY.md).) A task that needs a restricted Pi toolset would have to name Pi's lowercase tools — a documented follow-up. - **`permission_mode` is NOT enforced** — Pi headless print mode auto-runs tools and exposes only project-file trust (`--approve` / `--no-approve`), no tool-approval diff --git a/pyproject.toml b/pyproject.toml index 26724d7a..2c3fee5d 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -131,7 +131,7 @@ codex = [ # Without this extra the framework still installs and runs; Antigravity-dependent # code paths fail at start() with a clear hint pointing back here. antigravity = [ - "google-antigravity==0.1.18", + "google-antigravity==0.1.20", ] # Optional extra that enables OpenCode agent support. # diff --git a/src/coder_eval/agents/antigravity_agent.py b/src/coder_eval/agents/antigravity_agent.py index 5c82e488..3f28cb78 100644 --- a/src/coder_eval/agents/antigravity_agent.py +++ b/src/coder_eval/agents/antigravity_agent.py @@ -96,12 +96,20 @@ _POLL_DEADLINE_TIMEOUT_FRACTION = 0.8 # Cap on poll *cycles* -- the SOLE bound when a task sets no timeout at all, and -# a backstop against a very large one. 120 * 5s = 10 minutes, ~2x the worst real +# a backstop against a very large one (applied alongside the deadline, whichever +# is reached first). 120 * 5s = 10 minutes, ~2x the worst real # backgrounded-job duration observed (60-300s). Deliberately NOT "break after N # consecutive empty polls". # Rationale: .claude/notes/agents.md § Antigravity Step interleaving and the background poll _MAX_BACKGROUND_POLLS = 120 +# The only tool a model can background: a `run_command` left ACTIVE is a job +# that may still finish. Any other tool left ACTIVE (edit_file on a read-only +# mount, a paged view_file, start_subagent) never resolves, so it does not +# justify polling. +# Rationale: .claude/notes/agents.md § Antigravity Step interleaving and the background poll +_BACKGROUNDABLE_TOOL = "run_command" + # Antigravity builtin tool name -> the canonical (Claude) vocabulary every # criterion is written against. Unmapped names pass through unchanged. # Rationale: .claude/notes/agents.md § Tool-name and argument normalization @@ -120,6 +128,45 @@ "finish": "Finish", } +# Canonical (Claude) tool name -> the Antigravity builtin it enables, for +# `allowed_tools` / `disallowed_tools`. "Skill" and "TodoWrite" have no builtin +# (skills load through `skills_paths`), so they map to nothing; a name that is +# already an Antigravity builtin passes through. +# Rationale: .claude/notes/agents.md § Antigravity tool allowlist +_CLAUDE_TO_ANTIGRAVITY_TOOL_MAP: dict[str, str] = { + "Bash": "run_command", + "Read": "view_file", + "Write": "create_file", + "Edit": "edit_file", + "MultiEdit": "edit_file", + "Glob": "find_file", + "Grep": "search_directory", + "LS": "list_directory", + "Task": "start_subagent", + "WebSearch": "search_web", + "WebFetch": "read_url_content", + "AskUserQuestion": "ask_question", +} + +# Kept on under any allowlist: `finish` is how a turn returns structured +# output, not a capability an allowlist is meant to grant or withhold. +_ALWAYS_ENABLED_TOOLS: frozenset[str] = frozenset({"finish"}) + +# Set on every run_command unless the environment already sets them, so a +# command that would stop to ask (`npx` installing a package, git credentials, +# apt/pip confirmations, a pager) answers itself or fails instead of waiting on +# a terminal no one types into. +# Rationale: .claude/notes/agents.md § Antigravity non-interactive commands +_NONINTERACTIVE_ENV: dict[str, str] = { + "CI": "1", + "npm_config_yes": "true", + "GIT_TERMINAL_PROMPT": "0", + "DEBIAN_FRONTEND": "noninteractive", + "PIP_NO_INPUT": "1", + "PAGER": "cat", + "GIT_PAGER": "cat", +} + # Tool-call arg keys the harness ADDS at completion (the result payload), not # model-supplied inputs. The STATIC backstop; ``_params`` also strips any key # that first appears at DONE. A leaked result would false-positive @@ -303,18 +350,57 @@ def _resolve_workspaces(self, skills_paths: list[str]) -> list[str]: def _harness_env(self) -> dict[str, str] | None: """Per-agent environment for the localharness subprocess (``LocalAgentConfig.env``). - Returns the mock-CLI PATH prepend as a one-key overlay, or ``None`` when no - mock dirs are configured. The SDK merges it over ``os.environ`` at spawn, - so naming only ``PATH`` leaves every other inherited variable untouched. + An overlay of the non-interactive variables (each only where the + environment does not already set it) plus, when mock dirs are configured, + the mock-CLI PATH prepend; ``None`` when there is nothing to add, so the + SDK spawns with a plain inherited env. The SDK merges it over + ``os.environ`` at spawn, so every other inherited variable is untouched. The same overlay becomes the harness's ``run_command`` environment. """ - if not self._env_path_prepend: + env = {k: v for k, v in _NONINTERACTIVE_ENV.items() if k not in os.environ} + if self._env_path_prepend: + # Match the parent's own casing (Windows exports ``Path``) so the merge + # overrides the inherited entry instead of adding a sibling key. + path_key = next((k for k in os.environ if k.upper() == "PATH"), "PATH") + env[path_key] = os.pathsep.join([*self._env_path_prepend, os.environ.get(path_key) or ""]) + return env or None + + def _tool_capabilities(self, types: Any) -> Any: + """``CapabilitiesConfig`` restricting the builtin tools to the configured + ``allowed_tools`` / ``disallowed_tools``, or ``None`` when neither is set. + + Names are mapped through ``_CLAUDE_TO_ANTIGRAVITY_TOOL_MAP``; names with no + Antigravity builtin (``Skill``, ``TodoWrite``, MCP tools) are skipped. The + SDK takes an allowlist OR a denylist, so with both set the denied tools are + removed from the allowlist. + + Rationale: .claude/notes/agents.md § Antigravity tool allowlist + """ + allowed, disallowed = self.config.allowed_tools, self.config.disallowed_tools + if not allowed and not disallowed: return None - # Match the parent's own casing (Windows exports ``Path``) so the merge - # overrides the inherited entry instead of adding a sibling key. - path_key = next((k for k in os.environ if k.upper() == "PATH"), "PATH") - merged = os.pathsep.join([*self._env_path_prepend, os.environ.get(path_key) or ""]) - return {path_key: merged} + builtin = {t.value for t in types.BuiltinTools} + + def to_builtin(names: list[str] | None) -> set[str]: + mapped = {_CLAUDE_TO_ANTIGRAVITY_TOOL_MAP.get(n, n) for n in names or []} + return mapped & builtin + + # `enable_subagents` is a separate switch from the toolset; it follows + # whether `start_subagent` survives the filter. + subagent = types.BuiltinTools.START_SUBAGENT.value + if allowed: + enabled = (to_builtin(allowed) | _ALWAYS_ENABLED_TOOLS) - to_builtin(disallowed) + self._log.debug("Enabled builtin tools: %s", ", ".join(sorted(enabled))) + return types.CapabilitiesConfig( + enabled_tools=[types.BuiltinTools(t) for t in sorted(enabled)], + enable_subagents=subagent in enabled, + ) + disabled = to_builtin(disallowed) - _ALWAYS_ENABLED_TOOLS + self._log.debug("Disabled builtin tools: %s", ", ".join(sorted(disabled))) + return types.CapabilitiesConfig( + disabled_tools=[types.BuiltinTools(t) for t in sorted(disabled)], + enable_subagents=subagent not in disabled, + ) async def start( self, @@ -370,9 +456,11 @@ async def start( system_instructions=self.config.system_prompt or None, # Skill discovery: the search-path roots that parent the skill dirs. skills_paths=skills_paths, - # Mock-CLI PATH shadowing, per agent: two concurrent tasks never - # see each other's mock dirs. + # Mock-CLI PATH shadowing plus the non-interactive variables, per + # agent: two concurrent tasks never see each other's mock dirs. env=self._harness_env(), + # The builtin toolset, narrowed to allowed_tools / disallowed_tools. + capabilities=self._tool_capabilities(types), ) # Thinking level onto every resolved model's endpoint. The SDK # validates the model list in a model_validator, so options are set on @@ -547,11 +635,8 @@ def _on_turn_timeout() -> None: and not state.max_turns_hit and not state.timeout_hit and state.has_orphaned_tool_call() - and ( - poll_count < _MAX_BACKGROUND_POLLS - if poll_deadline is None - else time.monotonic() < poll_deadline - ) + and poll_count < _MAX_BACKGROUND_POLLS + and (poll_deadline is None or time.monotonic() < poll_deadline) ): poll_count += 1 self._log.debug("Polling for backgrounded work (orphaned tool call); attempt %d", poll_count) @@ -577,9 +662,9 @@ def _on_turn_timeout() -> None: # stop/timeout: the call is force-closed as unresolved and # the turn is still graded normally on everything else. bound = ( - f"poll_deadline ({_POLL_DEADLINE_TIMEOUT_FRACTION:.0%} of {timeout:g}s turn timeout)" - if poll_deadline is not None - else f"_MAX_BACKGROUND_POLLS ({_MAX_BACKGROUND_POLLS})" + f"_MAX_BACKGROUND_POLLS ({_MAX_BACKGROUND_POLLS})" + if poll_count >= _MAX_BACKGROUND_POLLS + else f"poll_deadline ({_POLL_DEADLINE_TIMEOUT_FRACTION:.0%} of {timeout:g}s turn timeout)" ) msg = "Poll budget exhausted (%s, poll_count=%d) with a tool call still ACTIVE." self._log.warning(msg, bound, poll_count) @@ -761,6 +846,8 @@ def __init__( # Most recently seen StepStatus per tool id, for has_orphaned_tool_call. # Separate from _closed_tools, which tracks only DONE/ERROR. self._tool_last_status: dict[str, Any] = {} + # Raw (Antigravity) tool name per tool id, for has_orphaned_tool_call. + self._tool_raw_names: dict[str, Any] = {} # Content blocks accumulated since the last per-generation flush. self._blocks: list[ContentBlock] = [] # Where the CURRENT generation started, advanced only by a flush that @@ -865,6 +952,7 @@ def _handle_tool_call(self, call: Any, step: Any, done: bool, sstatus: Any, call step_key = f"{trajectory_id}:{step.step_index}" if trajectory_id else str(step.step_index) cid = call.id or f"{raw_name}_{step_key}_{call_index}" self._tool_last_status[cid] = sstatus + self._tool_raw_names[cid] = raw_name if cid not in self._seen_tools: self._seen_tools.add(cid) seq = self._next_seq @@ -990,8 +1078,8 @@ def _agent_output(self) -> str: return "" def has_orphaned_tool_call(self) -> bool: - """True if any NOT-YET-CLOSED tool call's most recently seen status is - ACTIVE — the structural signature of a backgrounded task the model went + """True if any NOT-YET-CLOSED ``run_command``'s most recently seen status + is ACTIVE — the structural signature of a backgrounded task the model went idle on without waiting for. See ``communicate``'s poll loop. An ALLOWLIST on ACTIVE, never a denylist on "not yet closed": the SDK also @@ -999,9 +1087,18 @@ def has_orphaned_tool_call(self) -> bool: should wait out. The `not in _closed_tools` guard is layered on top as a monotonicity backstop, not a substitute. + Also an allowlist on the TOOL: only ``run_command`` can be backgrounded. + Any other tool left ACTIVE never resolves, so waiting on it only burns + the poll budget; ``finalize`` force-closes it as unresolved. + Rationale: .claude/notes/agents.md § Antigravity Step interleaving and the background poll """ - return any(cid not in self._closed_tools and s == _STATUS_ACTIVE for cid, s in self._tool_last_status.items()) + return any( + cid not in self._closed_tools + and s == _STATUS_ACTIVE + and self._tool_raw_names.get(cid) == _BACKGROUNDABLE_TOOL + for cid, s in self._tool_last_status.items() + ) def finalize(self, status: AgentEndStatus, *, crashed: bool = False, crash_reason: str | None = None) -> None: """Close orphaned tools, flush leftover blocks, emit TurnEnd + AgentEnd. diff --git a/tests/test_antigravity_agent.py b/tests/test_antigravity_agent.py index 5c0e963c..c5ddfebe 100644 --- a/tests/test_antigravity_agent.py +++ b/tests/test_antigravity_agent.py @@ -5,6 +5,7 @@ """ import asyncio +import enum import inspect import os import sys @@ -489,6 +490,25 @@ async def test_communicate_requires_started_agent(): await agent.communicate("x") +class _FakeBuiltinTools(enum.StrEnum): + """The SDK's ``types.BuiltinTools`` values, for the stubbed SDK.""" + + LIST_DIR = "list_directory" + SEARCH_DIR = "search_directory" + FIND_FILE = "find_file" + VIEW_FILE = "view_file" + CREATE_FILE = "create_file" + EDIT_FILE = "edit_file" + RUN_COMMAND = "run_command" + ASK_QUESTION = "ask_question" + START_SUBAGENT = "start_subagent" + GENERATE_IMAGE = "generate_image" + SEARCH_WEB = "search_web" + READ_URL_CONTENT = "read_url_content" + SCHEDULE = "schedule" + FINISH = "finish" + + def _install_fake_sdk(monkeypatch, sdk_agent_cls) -> None: """Stub ``google.antigravity`` in sys.modules so ``start()`` runs without the extra. @@ -502,6 +522,8 @@ def _install_fake_sdk(monkeypatch, sdk_agent_cls) -> None: ThinkingLevel=lambda level: level, GeminiAPIEndpoint=type("GeminiAPIEndpoint", (), {}), GeminiModelOptions=SimpleNamespace, + BuiltinTools=_FakeBuiltinTools, + CapabilitiesConfig=SimpleNamespace, ) hooks = ModuleType("google.antigravity.hooks") hooks.policy = SimpleNamespace( @@ -533,6 +555,7 @@ def test_has_orphaned_tool_call_detects_active_vs_other_statuses(): state = _AntigravityTurnState.__new__(_AntigravityTurnState) state._closed_tools = set() state._tool_last_status = {} + state._tool_raw_names = {"t1": "run_command", "t2": "run_command"} assert state.has_orphaned_tool_call() is False # no tool calls at all state._tool_last_status = {"t1": "ACTIVE"} @@ -557,6 +580,14 @@ def test_has_orphaned_tool_call_detects_active_vs_other_statuses(): state._tool_last_status = {"t1": "ACTIVE"} assert state.has_orphaned_tool_call() is False + # Only a run_command can be backgrounded: any other tool left ACTIVE never + # resolves on its own, so it must not arm the poll loop. + state._closed_tools = set() + for tool in ["edit_file", "view_file", "start_subagent"]: + state._tool_raw_names = {"t1": tool} + state._tool_last_status = {"t1": "ACTIVE"} + assert state.has_orphaned_tool_call() is False, f"an ACTIVE {tool} must not trigger polling" + async def test_communicate_fast_path_when_no_orphaned_tools(monkeypatch): """A normal turn closes its tool call before the stream exhausts -- the poll @@ -852,6 +883,67 @@ async def _record_sleep(seconds: float) -> None: assert bash.result_status == "unknown" # force-closed as UNRESOLVED by finalize() +async def test_communicate_poll_cap_also_bounds_a_turn_with_a_large_timeout(monkeypatch): + """The cycle cap applies alongside the timeout-derived deadline, not only when + no timeout is set: under a large turn_timeout (1800s gives a 1440s deadline) a + never-closing job stops at _MAX_BACKGROUND_POLLS instead of the deadline.""" + from coder_eval.agents import antigravity_agent + + monkeypatch.setattr(antigravity_agent, "_MAX_BACKGROUND_POLLS", 3) + sleep_calls: list[float] = [] + + async def _record_sleep(seconds: float) -> None: + sleep_calls.append(seconds) + + monkeypatch.setattr(antigravity_agent.asyncio, "sleep", _record_sleep) + + never_closing = [ + _step( + "TOOL_CALL", + "ACTIVE", + target="TARGET_ENVIRONMENT", + tool_calls=[_tc("run_command", "stuck", {"command_line": "node server.js"})], + ), + _step("TEXT_RESPONSE", "DONE", content="server started", complete=True, usage=_usage(10, 0, 1, 0)), + ] + agent = _agent_with_steps([never_closing]) + tr = await agent.communicate("start it", timeout=1800.0) + + assert len(sleep_calls) == 3 # the cap, long before the 1440s deadline + bash = next(c for c in tr.commands if c.tool_name == "Bash") + assert bash.result_status == "unknown" + assert tr.agent_output == "server started" + + +async def test_communicate_does_not_poll_a_non_command_tool_left_active(monkeypatch): + """Only a run_command can be backgrounded. An edit_file left ACTIVE (seen live: + an edit of a skill script on the read-only /work/plugins mount) never resolves, + so the turn finalizes at once instead of waiting out the poll budget.""" + from coder_eval.agents import antigravity_agent + + async def _sleep_should_not_be_called(_seconds: float) -> None: + raise AssertionError("asyncio.sleep must not be called for a non-command tool left ACTIVE") + + monkeypatch.setattr(antigravity_agent.asyncio, "sleep", _sleep_should_not_be_called) + + steps = [ + _step( + "TOOL_CALL", + "ACTIVE", + target="TARGET_ENVIRONMENT", + tool_calls=[_tc("edit_file", "e1", {"file_path": "/work/plugins/0/skills/s/scripts/x.py"})], + ), + _step("TEXT_RESPONSE", "DONE", content="done", complete=True, usage=_usage(10, 0, 1, 0)), + ] + agent = _agent_with_steps(steps) + tr = await agent.communicate("fix it", timeout=1800.0) + + assert agent._sdk_agent.conversation.receive_steps_call_count == 1 # poll loop never entered + edit = next(c for c in tr.commands if c.tool_name == "Edit") + assert edit.result_status == "unknown" # force-closed as UNRESOLVED by finalize() + assert tr.agent_output == "done" + + async def test_communicate_finalizes_gracefully_under_a_realistic_turn_timeout(monkeypatch): """A never-resolving orphan under a realistic 300s timeout finalizes gracefully. @@ -1172,18 +1264,55 @@ async def cancel(self): # never mutated, which is what lets two tasks start harnesses concurrently. +def _ni() -> dict[str, str]: + """The non-interactive variables the overlay adds: those the (possibly + monkeypatched) process env does not already set.""" + from coder_eval.agents import antigravity_agent + + return {k: v for k, v in antigravity_agent._NONINTERACTIVE_ENV.items() if k not in antigravity_agent.os.environ} + + async def test_harness_env_prepends_path_in_order(monkeypatch): """Mock dirs land at the FRONT of the overlay PATH, in order, ahead of the parent's.""" monkeypatch.setenv("PATH", "/parent/bin") agent = AntigravityAgent(parse_agent_config(type="antigravity")) agent._env_path_prepend = ["/sandbox/mocks", "/sandbox/bins"] - assert agent._harness_env() == {"PATH": f"/sandbox/mocks{os.pathsep}/sandbox/bins{os.pathsep}/parent/bin"} + assert agent._harness_env() == { + **_ni(), + "PATH": f"/sandbox/mocks{os.pathsep}/sandbox/bins{os.pathsep}/parent/bin", + } -async def test_harness_env_none_without_prepend(monkeypatch): - """No mock dirs → no overlay at all, so the SDK spawns with a plain inherited env.""" - monkeypatch.setenv("PATH", "/parent/bin") +async def test_harness_env_without_prepend_is_the_noninteractive_overlay(monkeypatch): + """No mock dirs → the overlay is just the non-interactive variables; PATH is inherited.""" + from coder_eval.agents import antigravity_agent + + monkeypatch.setattr(antigravity_agent.os, "environ", {"PATH": "/parent/bin"}) + agent = AntigravityAgent(parse_agent_config(type="antigravity")) + + assert agent._harness_env() == antigravity_agent._NONINTERACTIVE_ENV + + +async def test_harness_env_keeps_an_inherited_noninteractive_value(monkeypatch): + """A variable the environment already sets is left to inheritance, not overridden.""" + from coder_eval.agents import antigravity_agent + + monkeypatch.setattr(antigravity_agent.os, "environ", {"PATH": "/parent/bin", "CI": "false"}) + env = AntigravityAgent(parse_agent_config(type="antigravity"))._harness_env() + + assert env is not None and "CI" not in env + assert env["npm_config_yes"] == "true" + + +async def test_harness_env_none_when_nothing_to_add(monkeypatch): + """Every non-interactive variable already set and no mock dirs → no overlay at all, + so the SDK spawns with a plain inherited env.""" + from coder_eval.agents import antigravity_agent + + monkeypatch.setattr( + antigravity_agent.os, "environ", {"PATH": "/parent/bin", **antigravity_agent._NONINTERACTIVE_ENV} + ) agent = AntigravityAgent(parse_agent_config(type="antigravity")) assert agent._harness_env() is None @@ -1233,7 +1362,7 @@ async def test_harness_env_resolves_path_key_case_insensitively(monkeypatch): agent = AntigravityAgent(parse_agent_config(type="antigravity")) agent._env_path_prepend = ["/sandbox/mocks"] - assert agent._harness_env() == {"Path": f"/sandbox/mocks{os.pathsep}/parent/bin"} + assert agent._harness_env() == {**_ni(), "Path": f"/sandbox/mocks{os.pathsep}/parent/bin"} async def test_harness_env_handles_absent_path(monkeypatch): @@ -1244,7 +1373,7 @@ async def test_harness_env_handles_absent_path(monkeypatch): agent = AntigravityAgent(parse_agent_config(type="antigravity")) agent._env_path_prepend = ["/sandbox/mocks"] - assert agent._harness_env() == {"PATH": f"/sandbox/mocks{os.pathsep}"} + assert agent._harness_env() == {**_ni(), "PATH": f"/sandbox/mocks{os.pathsep}"} async def test_concurrent_starts_get_isolated_mock_dirs(monkeypatch, tmp_path): @@ -1288,8 +1417,8 @@ async def __aexit__(self, *exc): envs = [c.env for c in configs] assert envs == [ - {"PATH": f"/a/mocks{os.pathsep}/parent/bin"}, - {"PATH": f"/b/mocks{os.pathsep}/parent/bin"}, + {**_ni(), "PATH": f"/a/mocks{os.pathsep}/parent/bin"}, + {**_ni(), "PATH": f"/b/mocks{os.pathsep}/parent/bin"}, ] assert os.environ["PATH"] == "/parent/bin" # process env untouched throughout @@ -1318,12 +1447,18 @@ async def __aexit__(self, *exc): await agent.start(str(tmp_path), env_path_prepend=["/sandbox/mocks", "/sandbox/bins"]) assert agent._env_path_prepend == ["/sandbox/mocks", "/sandbox/bins"] - assert configs[0].env == {"PATH": f"/sandbox/mocks{os.pathsep}/sandbox/bins{os.pathsep}/parent/bin"} + assert configs[0].env == { + **_ni(), + "PATH": f"/sandbox/mocks{os.pathsep}/sandbox/bins{os.pathsep}/parent/bin", + } assert os.environ["PATH"] == "/parent/bin" # never mutated -async def test_start_omits_env_when_no_mock_dirs(monkeypatch, tmp_path): - """Without mock dirs the SDK gets env=None, so the harness inherits os.environ verbatim.""" +async def test_start_passes_only_the_noninteractive_env_when_no_mock_dirs(monkeypatch, tmp_path): + """Without mock dirs the SDK gets just the non-interactive overlay; PATH is inherited.""" + from coder_eval.agents import antigravity_agent + + monkeypatch.setattr(antigravity_agent.os, "environ", {"PATH": "/parent/bin"}) class _FakeSdkAgent: def __init__(self, cfg): @@ -1341,7 +1476,7 @@ async def __aexit__(self, *exc): agent = AntigravityAgent(parse_agent_config(type="antigravity")) await agent.start(str(tmp_path)) - assert configs[0].env is None + assert configs[0].env == antigravity_agent._NONINTERACTIVE_ENV # --- permission_mode ---------------------------------------------------------------- @@ -1383,6 +1518,96 @@ async def __aexit__(self, *exc): assert [p.kind for p in configs[0].policies] == ["allow_all"] +# --- allowed_tools / disallowed_tools -------------------------------------------------- +# +# Mapped onto the harness's builtin toolset (CapabilitiesConfig), so an experiment's +# allowlist confines Antigravity the way it confines Claude Code and Codex. + + +def _capabilities(**cfg): + return _agent(**cfg)._tool_capabilities( + SimpleNamespace(BuiltinTools=_FakeBuiltinTools, CapabilitiesConfig=SimpleNamespace) + ) + + +def test_tool_capabilities_none_without_tool_lists(): + """No allowed_tools / disallowed_tools → no CapabilitiesConfig, so the harness keeps its defaults.""" + assert _capabilities() is None + + +def test_allowed_tools_map_to_enabled_builtins(): + """The default experiment allowlist enables exactly the matching builtins, plus `finish`. + + `Skill` has no builtin (skills load through skills_paths) and is skipped; subagents, + web search and the rest stay off, as they do for Claude Code under the same list. + """ + caps = _capabilities(allowed_tools=["Bash", "Read", "Write", "Edit", "Glob", "Grep", "Skill"]) + + assert [t.value for t in caps.enabled_tools] == [ + "create_file", + "edit_file", + "find_file", + "finish", + "run_command", + "search_directory", + "view_file", + ] + assert caps.enable_subagents is False + + +def test_allowed_task_keeps_subagents(): + caps = _capabilities(allowed_tools=["Bash", "Task"]) + + assert {t.value for t in caps.enabled_tools} == {"run_command", "start_subagent", "finish"} + assert caps.enable_subagents is True + + +def test_disallowed_tools_map_to_disabled_builtins(): + caps = _capabilities(disallowed_tools=["Task", "WebSearch", "TodoWrite"]) + + assert [t.value for t in caps.disabled_tools] == ["search_web", "start_subagent"] + assert caps.enable_subagents is False + + +def test_disallowed_tools_are_removed_from_the_allowlist(): + """The SDK takes an allowlist OR a denylist, so with both the denied tools leave the allowlist.""" + caps = _capabilities(allowed_tools=["Bash", "Read", "Task"], disallowed_tools=["Task"]) + + assert {t.value for t in caps.enabled_tools} == {"run_command", "view_file", "finish"} + assert caps.enable_subagents is False + + +async def test_start_passes_tool_capabilities_to_sdk_config(monkeypatch, tmp_path): + configs: list[Any] = [] + + class _FakeSdkAgent: + def __init__(self, cfg): + configs.append(cfg) + + async def __aenter__(self): + return self + + async def __aexit__(self, *exc): + return False + + _install_fake_sdk(monkeypatch, _FakeSdkAgent) + + await _agent(allowed_tools=["Bash", "Read"]).start(str(tmp_path)) + + assert {t.value for t in configs[0].capabilities.enabled_tools} == {"run_command", "view_file", "finish"} + + +def test_installed_sdk_accepts_the_tool_capabilities(): + """Pin the SDK-side half: the real CapabilitiesConfig takes what we build and + every mapped name is a real builtin, so a renamed tool fails here, not live.""" + types = pytest.importorskip("google.antigravity").types + + assert set(agent_module._CLAUDE_TO_ANTIGRAVITY_TOOL_MAP.values()) <= {t.value for t in types.BuiltinTools} + caps = _agent(allowed_tools=["Bash", "Read", "Write", "Edit", "Glob", "Grep", "Skill"])._tool_capabilities(types) + assert isinstance(caps, types.CapabilitiesConfig) + assert types.BuiltinTools.START_SUBAGENT not in caps.enabled_tools + + # --- max_turns cap ------------------------------------------------------------------- # # max_turns caps main-thread model API calls, Claude Code's unit. A call opens at its diff --git a/uv.lock b/uv.lock index e5721665..7c30d3da 100644 --- a/uv.lock +++ b/uv.lock @@ -577,7 +577,7 @@ requires-dist = [ { name = "claude-agent-sdk", specifier = ">=0.2.157" }, { name = "click", specifier = ">=8.3.3" }, { name = "defusedxml", marker = "extra == 'dev'", specifier = ">=0.7.1" }, - { name = "google-antigravity", marker = "extra == 'antigravity'", specifier = "==0.1.18" }, + { name = "google-antigravity", marker = "extra == 'antigravity'", specifier = "==0.1.20" }, { name = "harbor", marker = "extra == 'harbor'", specifier = "==0.23.0" }, { name = "httpx2", specifier = ">=2.12.0,<3.0.0" }, { name = "jmespath", specifier = ">=1.1.0" }, @@ -960,7 +960,7 @@ wheels = [ [[package]] name = "google-antigravity" -version = "0.1.18" +version = "0.1.20" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "absl-py" }, @@ -972,12 +972,12 @@ dependencies = [ { name = "websockets" }, ] wheels = [ - { url = "https://files.pythonhosted.org/packages/4e/57/c53855740f8372635d266acfeaad7cd7f7a446d3d98c8b0f57b21443117f/google_antigravity-0.1.18-py3-none-macosx_11_0_arm64.whl", hash = "sha256:c13d51210945d3a154febb900c2c3077625c99322ceafe48f6a2e009254f1361", size = 37943163, upload-time = "2026-09-22T22:03:24.685Z" }, - { url = "https://files.pythonhosted.org/packages/d5/ef/088996f82e4e636dc94d58225d0a7e855b3026b3925f259c7b15b229e55d/google_antigravity-0.1.18-py3-none-macosx_11_0_x86_64.whl", hash = "sha256:16b4069be48b83e0ff01cfd79761e8408ec164ef21aec6c03485e63c5cd94ffb", size = 40251073, upload-time = "2026-09-22T22:03:28.331Z" }, - { url = "https://files.pythonhosted.org/packages/7d/1b/a261fb20a1554611dad99b4b5e8d2674f0f4861c42ff5902ed69a22b19c3/google_antigravity-0.1.18-py3-none-manylinux_2_17_aarch64.musllinux_1_1_aarch64.whl", hash = "sha256:14772cc17f7f89f1d47d26ff3aa1cb443596122f8b9844df1dbed7e18d35b73c", size = 38828002, upload-time = "2026-09-22T22:03:31.459Z" }, - { url = "https://files.pythonhosted.org/packages/6c/4e/f13bd23d26667abf5d90aa94e01e7902536343a417bc389de1b58ce986b3/google_antigravity-0.1.18-py3-none-manylinux_2_17_x86_64.musllinux_1_1_x86_64.whl", hash = "sha256:d22943eece2ba465101a65cb431440e037d3d00d5948babb5a271b68f8e1973d", size = 42782040, upload-time = "2026-09-22T22:03:35.077Z" }, - { url = "https://files.pythonhosted.org/packages/d6/14/6ebe59d4442e8a16cf63cc59a7d31b849161cc3252b6ed46ae2a1e5aadb0/google_antigravity-0.1.18-py3-none-win_amd64.whl", hash = "sha256:b6a37e4924fb01d5f66b4f3d2a49f7ee56672a6038a44ab556d68df9bef6c0c5", size = 44211098, upload-time = "2026-09-22T22:03:38.79Z" }, - { url = "https://files.pythonhosted.org/packages/4e/2e/24412c27b5fe85adf9509cc7c44fccc661be7af0526d085d333d00cae3fe/google_antigravity-0.1.18-py3-none-win_arm64.whl", hash = "sha256:ad790a8b9abcaa756a737584d8333728a01328a772d8002e6aa0297da1781a89", size = 40072554, upload-time = "2026-09-22T22:03:42.204Z" }, + { url = "https://files.pythonhosted.org/packages/77/1c/115e346cf2b10eb9632e862de5e2a130fc7f33ac0a352558067e9e863617/google_antigravity-0.1.20-py3-none-macosx_11_0_arm64.whl", hash = "sha256:a9df0f117407a7d39c4a03de97f4e082e1c2f24650dcca876706c23574118159", size = 38385383, upload-time = "2026-09-27T22:19:21.628Z" }, + { url = "https://files.pythonhosted.org/packages/1d/b3/49c09f0645d6d9c08e559bbb0a5adc45d22e299c541439ff36b8d8ed869a/google_antigravity-0.1.20-py3-none-macosx_11_0_x86_64.whl", hash = "sha256:8beba12288733e9f3f53e11103318c86dbd6814115890b8767eb4b5976f363a6", size = 40737091, upload-time = "2026-09-27T22:19:24.774Z" }, + { url = "https://files.pythonhosted.org/packages/2a/b2/f0ebc284dbdf3217ed62abd3a6b290effa7ae025555512aa9700fd20bb22/google_antigravity-0.1.20-py3-none-manylinux_2_17_aarch64.musllinux_1_1_aarch64.whl", hash = "sha256:1583b59ab45952231403da823e706a91d67f827cb2fbf27f97c2b9a26a803f09", size = 39260558, upload-time = "2026-09-27T22:19:28.063Z" }, + { url = "https://files.pythonhosted.org/packages/a1/38/a93e755d9504455f29ec483cd24e60c790ca82101191fe663e546f48bea1/google_antigravity-0.1.20-py3-none-manylinux_2_17_x86_64.musllinux_1_1_x86_64.whl", hash = "sha256:ff95642ea4e0a65d6e8d8d8b991ce22868d04717005cdc4a60abf2a6af52c2fc", size = 43262816, upload-time = "2026-09-27T22:19:30.851Z" }, + { url = "https://files.pythonhosted.org/packages/38/86/cce6b1576170f80f746999659e405c82a79113945f48520f23d0099d9b48/google_antigravity-0.1.20-py3-none-win_amd64.whl", hash = "sha256:2c931b0a5cdd5c121a7ae663969ee9f9b9109fc219c71c56273a4d646e30c114", size = 44781977, upload-time = "2026-09-27T22:19:33.723Z" }, + { url = "https://files.pythonhosted.org/packages/7a/9d/6c5b1ee7512f86c089f6a2e46b46b2372d9fa713aedd1f0730e3d506b43f/google_antigravity-0.1.20-py3-none-win_arm64.whl", hash = "sha256:5138da9ec775f7c2c0ad13e0cd93949d76430fc2f972c4005dac16a29c065125", size = 40601510, upload-time = "2026-09-27T22:19:36.635Z" }, ] [[package]] @@ -2400,15 +2400,14 @@ wheels = [ [[package]] name = "python-discovery" -version = "1.2.0" +version = "1.6.1" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "filelock" }, - { name = "platformdirs" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/9c/90/bcce6b46823c9bec1757c964dc37ed332579be512e17a30e9698095dcae4/python_discovery-1.2.0.tar.gz", hash = "sha256:7d33e350704818b09e3da2bd419d37e21e7c30db6e0977bb438916e06b41b5b1", size = 58055, upload-time = "2026-03-19T01:43:08.248Z" } +sdist = { url = "https://files.pythonhosted.org/packages/0c/57/250bd238b966cece44328235eb85290045d059265fdaf7527a3a958123db/python_discovery-1.6.1.tar.gz", hash = "sha256:cf87d3627dfb4412437fdd5b13eae402607722998d21567993aedbc59b23c15e", size = 84338, upload-time = "2026-09-18T01:31:53.971Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/c2/3c/2005227cb951df502412de2fa781f800663cccbef8d90ec6f1b371ac2c0d/python_discovery-1.2.0-py3-none-any.whl", hash = "sha256:1e108f1bbe2ed0ef089823d28805d5ad32be8e734b86a5f212bf89b71c266e4a", size = 31524, upload-time = "2026-03-19T01:43:07.045Z" }, + { url = "https://files.pythonhosted.org/packages/16/7d/e9ffbadfbf89c93848412d04594135c4ae8c1d37d9e053b9c3ed718fabc4/python_discovery-1.6.1-py3-none-any.whl", hash = "sha256:d43fcdef879fe795352bd13ccf8d185ba5a9f86f36cfcd00529f596e737442b3", size = 38664, upload-time = "2026-09-18T01:31:52.448Z" }, ] [[package]] @@ -3225,11 +3224,11 @@ wheels = [ [[package]] name = "urllib3" -version = "2.7.0" +version = "2.8.0" source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/53/0c/06f8b233b8fd13b9e5ee11424ef85419ba0d8ba0b3138bf360be2ff56953/urllib3-2.7.0.tar.gz", hash = "sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c", size = 433602, upload-time = "2026-05-07T16:13:18.596Z" } +sdist = { url = "https://files.pythonhosted.org/packages/e3/05/b17359e1cefb4f909b5e40b1b90a496d987258916dbbf88e842c729f510e/urllib3-2.8.0.tar.gz", hash = "sha256:63bf2ead4c879426ebf22ef2a781eeb4aa3b4ae798a0435506f8687fd5bb9b63", size = 458972, upload-time = "2026-09-15T19:29:36.253Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/7f/3e/5db95bcf282c52709639744ca2a8b149baccf648e39c8cc87553df9eae0c/urllib3-2.7.0-py3-none-any.whl", hash = "sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897", size = 131087, upload-time = "2026-05-07T16:13:17.151Z" }, + { url = "https://files.pythonhosted.org/packages/92/9d/c4e665119135114480843e7ab388fa94d8480650450e6f8e26b70d323a4c/urllib3-2.8.0-py3-none-any.whl", hash = "sha256:0cf3cae568d36aa9576b28dfb35f11328f1cb974ca7647d9475ebb86c75ac6e3", size = 135717, upload-time = "2026-09-15T19:29:34.577Z" }, ] [[package]] @@ -3259,17 +3258,18 @@ wheels = [ [[package]] name = "virtualenv" -version = "21.2.0" +version = "21.14.2" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "distlib" }, { name = "filelock" }, + { name = "packaging" }, { name = "platformdirs" }, { name = "python-discovery" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/aa/92/58199fe10049f9703c2666e809c4f686c54ef0a68b0f6afccf518c0b1eb9/virtualenv-21.2.0.tar.gz", hash = "sha256:1720dc3a62ef5b443092e3f499228599045d7fea4c79199770499df8becf9098", size = 5840618, upload-time = "2026-03-09T17:24:38.013Z" } +sdist = { url = "https://files.pythonhosted.org/packages/67/57/630a01cf5ab58f33b9c7dc8a7f13464cb5740b5227f8a08cab9d798bd532/virtualenv-21.14.2.tar.gz", hash = "sha256:571930928b11e43db690073ad8228162eca8a3f8fd3a87acdeae07df7dd57068", size = 5463798, upload-time = "2026-10-01T06:22:46.965Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/c6/59/7d02447a55b2e55755011a647479041bc92a82e143f96a8195cb33bd0a1c/virtualenv-21.2.0-py3-none-any.whl", hash = "sha256:1bd755b504931164a5a496d217c014d098426cddc79363ad66ac78125f9d908f", size = 5825084, upload-time = "2026-03-09T17:24:35.378Z" }, + { url = "https://files.pythonhosted.org/packages/1c/e0/10bb71ddf2946ccbe4ae06fca4e951212531e148daa3f45fa6b907bcafcf/virtualenv-21.14.2-py3-none-any.whl", hash = "sha256:6a8b7fd4e4ee00ebcdc1c5ee8cffd035d1fe130e354e8c9a44f510aac8affdc1", size = 5487107, upload-time = "2026-10-01T06:22:44.56Z" }, ] [[package]]