A cache-aware context trimming plugin for opencode.
"Long sessions stop costing money while you watch."
┌──────────────────────────────────────────────────────┐
│ opencode-token-min │
│ │
│ server hook ──▶ trim + digest each step's context │
│ cache-aware budgets keep the whole session lean │
│ every task logged: cost · tokens · est. saved │
│ right-sidebar Context box now tells the truth │
│ │
│ /╳\ MIT · one plugin │
└──────────────────────────────────────────────────────┘
License MIT · Branch main
opencode-token-min is a single-purpose plugin that stops the runaway-token grind in long interactive sessions. Every step in a chat re-sends your whole context to the model — history, tool dumps, cached prefixes — and by step 30 that is a lot of money per keystroke. This plugin rewrites the context before it ships: it trims history down to configurable per-role budgets, digests huge tool outputs into short summaries, and respects prompt caching (no cache breaks, so skipped-in tokens stay skipped-in). It tracks what it kept vs. trimmed for every step, records a per-task ledger, and the TUI plugin turns that into the Context box you actually wanted: tokens, % used, $ spent, total saved, % saved — and now saved on the last task, not a blurry lifetime total.
The leading goal, per the awesome LLM token-optimization guide: cut the re-sent-context bill by 80–99%. That is ambitious, measurable, and exactly what long sessions need.
Most cost is not the model. It is your own conversation being re-invoiced on every single step:
- You re-pay for all of history, every step. A 20-step task with 50k tokens of context isn't 50k tokens — it's ~1M tokens by the end. That's the cost you actually watch happen.
- Tool dumps dominate.
git status, file reads, 2,000-line logs all go into the prompt verbatim, then get re-sent forever. Trimming them is the single biggest lever. - The sidebar lies. opencode's built-in Context box shows only the last step's tokens. You can't see the trend, the burn, or what was saved — so the problem stays invisible until the invoice.
- Fixable, locally, with one hook. The
experimental.chat.messages.transformhook runs before every model call. Token-min is that hook, plus a ledger, plus an honest sidebar.
One-liner from npm (installs the server + TUI side by side and registers
them in both opencode.json and tui.json — nothing to copy by hand):
opencode plugin opencode-token-min --global
# → adds "opencode-token-min" to your config's plugin array
# → later updates: opencode plugin opencode-token-min --global --forcePrefer the source version? Then it's the same manual copy:
mkdir -p ~/.config/opencode/plugins ~/.config/opencode/tui
# server plugin (trim + digest + ledger + /cost tool)
cp plugins/token-min.ts ~/.config/opencode/plugins/
# TUI plugin (the honest Context box on the right)
cp tui/token-min.tsx ~/.config/opencode/tui/
# config file for sidebar (registers token-min, renders the honest Context box)
cp tui.json ~/.config/opencode/tui.json
# note: if you already have a tui.json (e.g. other TUI plugins), merge the
# "plugin" arrays instead of overwriting it.Either way: restart opencode, or just /tui --dev to hot-reload the sidebar.
Zero config. Sanely conservative defaults. Everything is tunable via the
constants at the top of plugins/token-min.ts:
| Constant | Default | Meaning |
|---|---|---|
MODE |
"auto" |
auto (default) = trim + digest every step. watch = measure only (safe). trim = same as auto. cached = trims + widens budgets unconditionally |
MAX_USER / MAX_ASSISTANT / MAX_TOOL |
12 / 12 / 12 |
max messages per role kept in the tail |
MAX_TOTAL |
30 |
total messages kept (history + tail) |
MAX_TOTAL_CACHED |
60 |
tail cap when the session has an active cache |
CACHED_MULT |
2 |
multiplier applied to per-role budgets once a cache is confirmed |
PRESERVE_FIRST |
3 |
instructions/system messages never trimmed |
MIN_KEEP |
2 |
floor regardless of budgets |
OLD_MULT |
2 |
loosen budgets while cache state is unknown |
CHARS_PER_TOKEN |
4 |
chars÷this ≈ tokens for saved estimates |
KEEP_TAIL_MSGS |
4 |
recent tool parts skipped by digesting |
TOOL_DIGEST_BYTES |
4000 |
output limit before a tool part becomes a digest |
TOOL_DIGEST_HEAD / TOOL_DIGEST_TAIL |
800 / 800 |
chars kept at each end when digesting |
The default auto mode trims and digests from the very first install. If you'd
rather measure first, set MODE = "watch" — it only records ~saved numbers
without touching a single token. Sessions that confirm an active cache switch
to cached budgets (MAX_TOTAL_CACHED / CACHED_MULT): keeping more of the
warm prefix intact is cheaper than forcing a rewrite, so trimming steps aside.
Set TOKEN_MIN_LOG=1 (or true) to enable per-step console output. By
default the plugin is quiet — LOG is opt-in, so nothing prints unless you ask.
The ledger (token-usage.jsonl) is always written regardless. (Output appears
in the TUI status area above the chat field when enabled.)
every step's messages
┌──────────────────────────────────────────────┐
│ ▼
│ experimental.chat.messages.transform ──▶ trimMessages()
│ │ │
│ │ per-role budgets + total cap │ preserve first N
│ │ (loosened while cache unknown) ▼
│ │ digest oversized tool outputs ──▶ keep-then-drop
│ │ │
│ ▼ ▼
│ token-usage.jsonl ◀──── step-finish ──▶ ledger append
│ │ per-task rows: cost · tokens · estSaved
│ ▼
│ sidebar Context box (tui/token-min.tsx)
│ context · tokens · % used · $ · saved · % saved · last task
└──────────────────────────────────────────────────────────────
- Measure first. Every step records chars-before vs. chars-after and
converts the difference to an estimated token saving (
chars/4 ≈ tokens), plus cache-read tokens when the provider reports them. - Cache-aware. While the cache state of a session is unknown the budgets are
doubled (
OLD_MULT), so trimming can't silently break a warm cache. Once the session confirms an active cache the per-role budgets widen again (CACHED_MULT) and the tail targetsMAX_TOTAL_CACHED: keeping the warm prefix intact beats forcing a rewrite, since re-sending cached tokens is far cheaper than new ones. - Tool outputs get digested, not deleted. Oversized outputs are rewritten
to a short digest (
[digested N-byte tool output → summary]), keeping the tail of the conversation useful instead of just small. - Per-task ledger. One JSONL row per step, tagged with the originating user
message (
taskID), so "how much did that last attempt cost and save" is a real number — not the whole-session ballpark. - The sidebar tells the truth. Context / tokens / % used / $ spent / ~saved / % saved, then ~N tokens saved · last task — the two numbers that matter mid-session, from the ledger, not from API internals.
- Ask the agent. The
costtool prints session / today / all-time totals in chat; a/cost-style right-menu item is planned.
| In the sidebar | What it means |
|---|---|
8,460 tokens |
tokens in the last step, as the API reports them |
7% used |
side panel heading as stock opencode |
$1.23 spent |
cost of the last step (if the provider reports it) |
~3,063,221 tokens saved |
cumulative est. tokens not re-sent, from the ledger |
81% saved |
ΣestSaved ÷ ΣbeforeTok across the session — e.g. 9,852,160 of 12,235,272 (live 7kg47KC0 session) |
~622,254 tokens saved · last task |
savings from the most recent task only |
93% saved |
same formula, but for the last task's own rows (lestSaved ÷ lastBeforeTok) |
$ opencode /tui --dev # hot-reload the sidebar while you tweakNumbers from live sessions (default auto mode):
| Run | Input tokens sent per step |
|---|---|
| No token-min, other context plugin active | 441,934 (one request) · a heavy test session sustained 700k–838k on every step |
Same-class session, token-min auto |
~4,000 shipped (+ ~14k cache-read that stays warm) |
Where a pluginless/other-plugin step was about to ship ~178k tokens, token-min
sent ~3k — the ledger logged estSaved ≈ 175k per step, i.e. a ~98% cut.
The gap is why this exists: the untrimmed case re-invoices the entire session
every keystroke; the trimmed case only pays for what actually changed.
Two long-running sessions this week (auto mode, GLM-5.3), names and plan
details kept out. One session started on a local non-billable model before
being switched back to GLM; those rows are excluded, so the table counts GLM
steps only. Both ran from an afternoon into the next day.
| Session | Timeframe (UTC) | GLM steps | GLM cost | input sent | cache-read | est. tokens saved | est. $ saved* | % saved† |
|---|---|---|---|---|---|---|---|---|
| Private site + API + apps build | Sep 20–21, 16:06–17:11 | 3,022 | $92.17 | 52.5M | 49.1M | 1,511.4M | $435.82 | ~97.5% |
| App modernisation session | Sep 20–21, 18:09–17:11 | 945 | $27.54 | 15.8M | 14.1M | 416.5M | $101.84 | ~96.5% |
| Both combined | 48.1 h total‡ | 3,967 | $119.71 | 68.3M | 63.2M | 1,927.9M | $537.66 | ~97.3% |
*$ saved = est. saved tokens × that session's fitted GLM cache-read rate
(~$0.29/M and $0.25/M) — the cheapest possible re-send, so the figures are
conservative. At each session's actual blended billed rate ($0.90/M) the same
tokens are ~$1,363 and $380; at full input price ($1.43/M), ~$2,161 and
~$603.
†% saved = ΣestSaved ÷ ΣbeforeTok, zero-trim rows included, so this is
conservative. Combined, the sessions billed ~132M effective tokens (input +
cache-read) while the same steps would have re-sent ~1,982M context tokens —
~97% of the re-sent context never reached the API, and the context that
did ship was only ~54M tokens.
‡Sum of each session's elapsed span — the two ran partly in parallel and
include idle time, so the per-hour rates below are diluted (conservative).
Take those two sessions — 48 h of session time combined, $119.71 billed, ~$538–2,763 of context not re-sent — and run them like a job: 8-hour day, 7-day week, 4-week month, 12-month year. That's 2,688 hours, about 29% more than a standard 40-h-week year (2,080 h). Linear extrapolation from two real sessions (elapsed spans include idle time), so treat it as the shape of the number, not a quote:
| Period | Hours | Billed with token-min | $ saved (floor*) | $ saved (blended) | $ saved (full input) |
|---|---|---|---|---|---|
| One day (8 h) | 8 | $20 | $89 | $290 | $460 |
| One week (7 days) | 56 | $139 | $626 | $2,029 | $3,217 |
| One month (4 weeks) | 224 | $557 | $2,503 | $8,115 | $12,866 |
| One year (12 months) | 2,688 | ~$6,700 | ~$30,000 | ~$97,000 | ~$154,000 |
*Floor / blended / full input reuse the per-session rates from the table
above: the trimmed context priced at each session's fitted cache-read
($0.25–0.29/M), at the sessions' actual blended billed rate ($0.90/M), and at
full input price (~$1.43/M) — the cheapest, the likely, and the no-cache worst
case. Savings scale with context weight: tool-heavy agentic sessions like
these trim the most; chat-shaped sessions save less. The honest year-end
comparison: ~$6.7k billed with token-min against ~$30k–154k of context that
was never re-sent.
- Usage and cost are provider-truth; savings are an estimate. The
input/output/reasoning/cacheRead/cacheWritetokens and the$figure are copied verbatim from the provider's step-finish report — the same numbers an invoice would be based on. Only what the sidebar calls saved is an estimate: there is no invoice for tokens you didn't re-send, so the plugin measures the context before vs. after trimming (chars ÷ 4, conservative upper bound, floored at 0). That part is for watching the trend; the usage and$columns are for accounting. - The default (
auto) trims from day one. If you want to measure first, setMODE = "watch"— nothing is touched until you opt in withtrim/cached. - Cache safety is measured, not guaranteed — the plugin errs on the side of smaller context, which is never wrong on cost.
tail -5 ~/.local/share/opencode/token-usage.jsonlOne row per step-finish: ts, sessionID, taskID, messageID, model, cost, tokens {input, output, reasoning, cacheRead, cacheWrite, estSaved, beforeTok, afterTok}.
estSaved (≈ chars-saved ÷ 4) is conservative and floored at 0; beforeTok /
afterTok are estimates of the prompt size before vs. after trimming. Sum
estSaved per taskID and you have exactly what a long session really burned.
bun build plugins/token-min.ts --outdir /tmp/token-min-build # syntax check
# then: run a long session → inspect the ledger (MODE="watch" to observe first)MIT — do whatever, keep the name.
Built by Big Pickle, directed by the boss.