Skip to content

plugin: fewer wasted tool calls; workspace snapshot on the first turn - #88

Merged
mjwsolo merged 7 commits into
mainfrom
feat/fewer-steps
Sep 27, 2026
Merged

mjwsolo merged 7 commits into
mainfrom
feat/fewer-steps

Conversation

@mjwsolo

@mjwsolo mjwsolo commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Plugin-only. Measured with the evals suite (Qwen3.8-27B Q4, one trial per task, same idle machine, main vs branch):

  • 16-task regression set: 16/16 on both. Tool calls 97 -> 81 (-16%), wall 1,381 s -> 1,171 s. On the original 14 tasks: 81 -> 62 calls (-23%), median steps 7 -> 4.
  • Ten-turn session re-reads 22-29 tokens per turn after the first; topic-change task answers the unrelated question with zero tool calls and resumes the code work.

What changed:

  1. Planning rule: orient from the snapshot instead of listing directories; plan only for 3+ steps and update todos once per step; run the check once after the edit; one edit per file per step.
  2. A WORKSPACE SNAPSHOT part on the session's first user turn (200 entries, labelled as session-start), and on later turns a WORKSPACE CHANGED SINCE YOUR LAST TURN part for files added/modified/removed outside the agent's own edits. Nothing in the system prompt, so the cached prefix is unchanged (Checks 10/10).
  3. Per-session state (snapshot sent, file tree, agent-touched paths) persists in <workspace>/.localcode-agent/ws-<session>.json, because a headless run is one process per turn and users reopen the UI.
  4. The todo and placeholder gates show a 'Continuing' toast.

Trade-off seen in the data: with the check-after-edit rule the agent runs the test suite after each edit in multi-turn sessions (ten-turn task: 10 -> 14 calls), which is correct behaviour but not free.

Tests: 46 bun tests pass (4 known pre-existing failures), new layout tests cover constancy, caps, restart survival and external-change reporting.

Also in this PR (commit 98bcd89): the plateau breaker no longer aborts the session. It measured progress only by edits, checks and completed todos, so a research turn with no edits was stopped after 14 rounds as "no progress", and the abort killed the running tool and showed a bare "Interrupted". Now a turn with nothing delivered yet counts a never-seen tool call as progress (only repetition stalls it), and a stop refuses further tool calls with the reason so the model writes up and the turn ends on its own; abort only after three ignored refusals. 49 bun tests pass; Checks 10/10.

…cipline rules, gate toasts

Measured on the regression evals (Qwen3.8-27B): ~40% of tool calls did no work — directory
reads to orient, globs to find the tests after the edit, checks run before any change,
todowrite before and after the same edit. Changes, all plugin-side:
- a WORKSPACE LAYOUT block (two levels, 60 entries / 2,500 chars cap, computed once per
  session so the cached prefix is unchanged by files created mid-session)
- rules: orient from the layout, plan only for 3+ steps and update todos once per step,
  run the check once after the edit, one edit per file per step
- the todo and placeholder gates show a 'Continuing' toast so a pause is never a mystery
Unit test pins the layout block's constancy and limits.
… on later turns

The layout block moved out of the system prompt (Gemini CLI and Cline put it in the
first message; aider keeps the map after the stable prefix). Now: a WORKSPACE SNAPSHOT
part on the first user turn (200 entries, labelled as session-start), and on later turns a
WORKSPACE CHANGED SINCE YOUR LAST TURN part listing files added, modified or removed
outside the agent's own edits (mtime scan, 20 per kind). The rule now says: glob for the
specific file if it is not in the snapshot or a change report. Prefix unchanged.
A headless run is one process per turn and a user may reopen the UI mid-session; with
in-memory state the snapshot was re-sent on 33 of 36 turns in the evals and no change was
ever reported. State (sent flag, file tree, agent-touched paths) now persists under
.localcode-agent/ws-<session>.json in the workspace.
…dowrite calls in a row

The previous wording ('once per step, never as a separate call between edits') made
the model stop updating the plan altogether: one todowrite at the start, none after,
contradicting the tool's own 'update as you go'. Now: one call when a step's last edit
lands (it may set the next item in_progress too), never twice in a row without an edit
or check between, never before AND after the same edit.
…calls are progress until something is delivered

The breaker measured progress only by edits, passing checks and completed
todos, so a research or investigation turn that kept reading, fetching and
trying new commands was stopped after 14 rounds as no progress, and the stop
was a session.abort that killed the running tool and showed a bare
Interrupted. Now a turn with nothing delivered yet counts a round with a
never-seen tool call as progress (only repetition stalls it), and a stop
refuses further tool calls with the reason so the model writes up and the
turn ends on its own; abort remains only after three ignored refusals.
@mjwsolo
mjwsolo merged commit 2cd5dc7 into main Sep 27, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant