plugin: fewer wasted tool calls; workspace snapshot on the first turn - #88
Merged
Merged
Conversation
…cipline rules, gate toasts Measured on the regression evals (Qwen3.8-27B): ~40% of tool calls did no work — directory reads to orient, globs to find the tests after the edit, checks run before any change, todowrite before and after the same edit. Changes, all plugin-side: - a WORKSPACE LAYOUT block (two levels, 60 entries / 2,500 chars cap, computed once per session so the cached prefix is unchanged by files created mid-session) - rules: orient from the layout, plan only for 3+ steps and update todos once per step, run the check once after the edit, one edit per file per step - the todo and placeholder gates show a 'Continuing' toast so a pause is never a mystery Unit test pins the layout block's constancy and limits.
… on later turns The layout block moved out of the system prompt (Gemini CLI and Cline put it in the first message; aider keeps the map after the stable prefix). Now: a WORKSPACE SNAPSHOT part on the first user turn (200 entries, labelled as session-start), and on later turns a WORKSPACE CHANGED SINCE YOUR LAST TURN part listing files added, modified or removed outside the agent's own edits (mtime scan, 20 per kind). The rule now says: glob for the specific file if it is not in the snapshot or a change report. Prefix unchanged.
A headless run is one process per turn and a user may reopen the UI mid-session; with in-memory state the snapshot was re-sent on 33 of 36 turns in the evals and no change was ever reported. State (sent flag, file tree, agent-touched paths) now persists under .localcode-agent/ws-<session>.json in the workspace.
…dowrite calls in a row
The previous wording ('once per step, never as a separate call between edits') made
the model stop updating the plan altogether: one todowrite at the start, none after,
contradicting the tool's own 'update as you go'. Now: one call when a step's last edit
lands (it may set the next item in_progress too), never twice in a row without an edit
or check between, never before AND after the same edit.
…calls are progress until something is delivered The breaker measured progress only by edits, passing checks and completed todos, so a research or investigation turn that kept reading, fetching and trying new commands was stopped after 14 rounds as no progress, and the stop was a session.abort that killed the running tool and showed a bare Interrupted. Now a turn with nothing delivered yet counts a round with a never-seen tool call as progress (only repetition stalls it), and a stop refuses further tool calls with the reason so the model writes up and the turn ends on its own; abort remains only after three ignored refusals.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Plugin-only. Measured with the evals suite (Qwen3.8-27B Q4, one trial per task, same idle machine, main vs branch):
What changed:
<workspace>/.localcode-agent/ws-<session>.json, because a headless run is one process per turn and users reopen the UI.Trade-off seen in the data: with the check-after-edit rule the agent runs the test suite after each edit in multi-turn sessions (ten-turn task: 10 -> 14 calls), which is correct behaviour but not free.
Tests: 46 bun tests pass (4 known pre-existing failures), new layout tests cover constancy, caps, restart survival and external-change reporting.
Also in this PR (commit 98bcd89): the plateau breaker no longer aborts the session. It measured progress only by edits, checks and completed todos, so a research turn with no edits was stopped after 14 rounds as "no progress", and the abort killed the running tool and showed a bare "Interrupted". Now a turn with nothing delivered yet counts a never-seen tool call as progress (only repetition stalls it), and a stop refuses further tool calls with the reason so the model writes up and the turn ends on its own; abort only after three ignored refusals. 49 bun tests pass; Checks 10/10.