Preserve the agent's working state—goals, files, decisions, errors, constraints, and open loops—not just a vague recap of the conversation.
Deterministic facts · fail-closed verification · branch-safe continuity · local-first privacy
Install · Why · Pipeline · Safety · Configuration · Development
Requires the Pi Coding Agent on Node.js 22.19 or newer. The published extension uses Pi's host packages and does not bundle a second Pi runtime.
pi install npm:pi-smart-compactThen run /smart-compact for an explainable preflight before anything changes.
/smart-compact # interactive preflight
/smart-compact auto # adaptive mode selection
/smart-compact anthropic/claude-sonnet-4 fast # explicit model + mode
/smart-compact balanced --focus=auth # preserve extra auth detail
/smart-compact metrics # text metrics report
/smart-compact dashboard # interactive dashboard
/smart-compact restore # browse and restore backups
/smart-compact loops # manage persisted open loopsAt 60% context usage by default, the extension also participates in Pi's native
compaction flow. Long-running agents can call smart_compact, smart_recall,
and smart_save_memory directly.
Important
The tool path only stages a verified pending summary for Pi's next natural compact. It never compacts the active conversation in the middle of an agent turn.
| Native-style recap | pi-smart-compact |
|---|---|
| Summarizes prose | Preserves operational coding state |
| Trusts one LLM response | Extracts deterministic ground truth first |
| File/error omissions can be silent | Verifies coverage and repairs known gaps |
| One strategy for every session | Chooses single-pass or hierarchical synthesis |
| No quality feedback | Tracks provenance, damage signals, and metrics |
| No scoped cross-session recall | Searches a project-isolated SQLite FTS5 context graph |
The design principle is simple:
Facts first. Synthesis second. Verification before apply.
Any unresolved verification gap rejects the custom summary before staging or apply. A zero-gap deterministic fallback is preferred over unverifiable model output; only failure of that fallback rejects the run. A successful compaction must also meet its mode target and at least 10% estimated net savings both before synthesis and after the final summary is measured. Automatic failures leave Pi free to use its native compactor, while manual failures leave the conversation unchanged.
Pi conversation
│
▼
┌───────────┐ ┌───────────┐ ┌────────────┐ ┌───────────┐
│ Extract │ → │ Explore │ → │ Synthesize │ → │ Verify │
│ 0 LLM │ │ adaptive │ │ 1-pass or │ │ + repair │
│ calls │ │ │ │ hierarchical│ │ │
└───────────┘ └───────────┘ └────────────┘ └───────────┘
│
▼
staged/applied by Pi
| Stage | Responsibility |
|---|---|
| Extract | Deterministically catalogs files, errors, decisions, constraints, topics, media metadata, and open loops. This is the verification ground truth. |
| Explore | Runs only in thorough mode (or when auto selects it); cheaper modes use deterministic boundaries. |
| Synthesize | Uses adaptive single-pass or bounded hierarchical synthesis with per-mode call, prompt-token, chunk, and output budgets. |
| Verify | Applies deterministic repairs to a bounded fixed point, then uses a verified deterministic quality floor. Only thorough may spend one additional LLM repair call before that fallback. |
- The current goal and user constraints
- Modified, read, and deleted files
- Unresolved and resolved error history; free-form goal changes never claim an unfixed error was resolved
- Explicit and implicit decisions
- Open follow-ups, blockers, priorities, and pinned loops
- Next actions and critical continuation context
- Changes since the previous compaction
- A bounded Continuity Ledger carrying prior decisions, constraints, unresolved errors, and open loops across follow-ups and goal wording changes. Goal shifts are recorded as context; facts retire only through positive resolution evidence or an explicit override.
Summaries use a canonical H1/H2/H3-aware structure, collision-safe file matching, typed verification gaps, and persisted repair provenance.
Applied compactions also index their verified scoped state into a bounded,
project-isolated SQLite FTS5 context graph. smart_recall searches goals,
decisions, constraints, unresolved errors, open loops, files, and critical
context across this project's sessions; recall and resolution use the complete
visible branch ancestry, never sibling-branch state. File relationships add
one-hop graph recall without an embedding service or extra LLM call.
smart_save_memory persists or explicitly resolves one user-confirmed decision,
constraint, preference, warning, procedure, or context fact. It fails closed
when the working directory is exactly HOME or the filesystem root, and
requires an interactive confirmation showing the complete scrubbed title,
content, and paths. Each project may have at most 500 active manual memories.
It rejects empty inputs, scrubs configured secrets/PII, deduplicates exact facts,
and must not be used for guesses, transient progress, secrets, or code that is
cheap to re-read. Set contextGraphEnabled to false to disable indexing and
both tools.
| Surface | Behavior |
|---|---|
/smart-compact |
Explicit manual run. Opens a target-first preflight or accepts direct args, dry-run, focus, and budgets. |
session_before_compact |
Auto path. Returns/stages a verification-scored summary under pressure; durable state waits for matching session_compact. |
smart_compact tool |
Agent path. Produces a pending summary for Pi's next natural compact; does not compact mid-turn. |
/smart-compact loops |
Project-level open-loop manager: resolve/reopen, priority, pin/unpin. |
The interactive command uses the configured summary route and exact execution
planner before spending LLM tokens. Its compact decision card compares the
three modes by estimated after-size and saving, highlights the recommendation,
and keeps only the selected plan plus hard tool-pair/zero-gap guarantees in the
primary view. The plan reserves 25% of the LLM summary allowance for verified
state/delta/continuity sections added after synthesis. Technical estimator,
target, route, and boundary data stays under D instead of crowding the decision.
↑/↓changesFast,Balanced, orThoroughand recalculates the plan.Enterruns only a viable plan;Esccancels without mutation.Dtoggles calibrated estimator, target, boundary, and route details.Mopens Advanced model selection and replans with that route's calibration.
Values remain estimates until the next provider turn reports usage. Smart
Compact therefore uses ~/≤ language, measures the completed summary again
before staging, and reports final success only after Pi confirms the matching
session_compact run ID. A single long user turn may be split at a safe message
boundary: its older prefix is verified into the summary while the budgeted
working tail stays raw. Tool exchanges are summarized or retained as complete
call/result pairs, never split. If Verify, yield, provider, or native apply
fails, the UI shows one bounded actionable line without evidence text or a
JavaScript stack. A successful 100/100 is labeled verification coverage;
the source score and deterministic/LLM/fallback provenance remain visible so
repaired coverage is never presented as raw synthesis quality. Stack diagnostics
are opt-in with DEBUG=smart-compact.
During execution a two-line live brief shows the EESV phase chain and the
meaningful current action; it states that the conversation remains unchanged
until verified Apply. Routine phase toasts and raw per-batch watchdog/provider
errors are suppressed by default: handled fallbacks appear as one content-free
brief. verbose restores routine phase notices; full stack diagnostics require
DEBUG=smart-compact.
/smart-compact balanced --focus=authentication
/smart-compact fast --max-input-tokens=120000
/smart-compact fast --focus=src/auth.ts --max-calls=3
/smart-compact thorough--focusassigns more synthesis/exploration budget to a topic or path. It does not attempt unsupported non-contiguous compaction.--max-callsaccepts1–100.--max-input-tokensaccepts10000–1000000aggregate prompt tokens.--max-latencyaccepts5000–600000milliseconds as a cancellation deadline; cancellation waits for safe pipeline unwind before returning.- Budget exhaustion degrades to deterministic summaries instead of dropping context.
The tool exposes equivalent focus, max_calls, max_input_tokens, and
max_latency_ms parameters.
| Mode | Calls | Prompt cap | Output cap | Behavior |
|---|---|---|---|---|
fast |
3 | 100K | 20K | Quickest recovery; 3K summary, 10K recent tail, 30% context target |
balanced |
6 | 200K | 40K | Default quality/speed trade-off; 6K summary, 20K recent tail, 40% target |
thorough |
8 | 300K | 80K | Deepest analysis; 10K summary, 30K recent tail, 50% target, Explore and optional LLM repair |
These are the only three execution modes. Automatic runs choose among them
from context pressure and deterministic session risk; auto is a selector,
not a fourth execution policy. Fast can use a zero-call deterministic summary
when extraction confidence is high; otherwise it keeps the bounded LLM path.
The mode token target is binding: recent user turns, pi-toolkit checkpoints,
and topical grouping remain raw only when they fit the planned tail; otherwise
the verified summary carries them forward. Automatic risk refinement may deepen
analysis/repair strategy after extraction, but it does not mutate the profile
allowance or retention window that was already used to prove the target.
Output caps stop subsequent calls after observed usage reaches the threshold.
The ChatGPT Codex subscription endpoint rejects max_output_tokens,
max_tokens, and max_completion_tokens; Smart Compact therefore enforces a
15–90 second per-call watchdog plus a streamed visible-output ceiling and falls
back deterministically on abort. Custom Codex endpoints receive
max_output_tokens through Pi AI's payload hook.
Legacy compression profiles remain as advanced/backwards-compatible policy:
light maps to thorough, balanced maps to balanced, and aggressive maps
to fast with a deprecation warning. The selected model never changes
automatically; M changes the summary route inside preflight and recalculates
all three plans.
All stages use the selected Pi model by default. Routing is explicit and independent of modes:
| Stage | Config key | Default |
|---|---|---|
| Explore / segmentation | segmentationModel |
selected model |
| Synthesis / assembly | summaryModel |
selected model |
| Verification repair | verificationModel |
summary/selected model |
Every run persists per-stage provider, model, reliability, latency, and token
telemetry with schema-versioned verifier quality. bun run provider-eval
builds an advisory matrix by context pressure and tool density; it never edits
configuration or selects a model. Legacy rows contribute operational evidence
but not quality because old verifier score semantics are incompatible.
A reproducible paid-API probe is opt-in only:
bun run provider-eval:live --live \
--models=openai/gpt-5.4,anthropic/claude-sonnet-4-6It runs three bounded, identical coding-continuity scenarios and reports verification score, latency, and token usage. Apply a route manually only after representative evidence. See the dated provider evaluation baseline.
Raw local JSONL remains available to the interactive dashboard, while
bun run telemetry-report emits aggregate-only telemetry: no session/project
IDs, prompts, summaries, paths, or error text. Failures use a stable taxonomy
(cancelled, timeout, rate limit, authentication, budget, output limit,
provider, persistence, validation, verification, yield, internal).
Verification and yield failures retain only content-free diagnostics.
Set telemetryChannel to canary only on the externally selected canary
cohort. The report shows total/applied counts, but only non-dry, host-confirmed
applied runs satisfy promotion evidence. A deterministic green check never
implies PROMOTE; the report compares schema-v2 canary runs with the stable
baseline and returns HOLD, ROLLBACK, or PROMOTE. Rollback triggers are: failure rate
+5pp and ≥10%, verifier quality −5 points, p95 latency +50%, tokens +50%,
heuristic fallback +10pp, or post-compaction damage +10pp. Promotion requires
20 non-dry applied canary runs, a stable baseline, ≥70% verifier-quality coverage,
≥70% run-correlated damage-observation coverage in both cohorts, ≥85 absolute
canary quality, and ≥95% success. The extension reports the decision; it never
edits config or deploys automatically.
The interactive and HTML dashboards make trust evidence explicit: a Data Confidence score (target ≥85) combines recent sample size (25 points), schema-v2 coverage (25), verifier-quality coverage (20), field completeness (20), and seven-day freshness (10). Separate views show repair gain and quality bands, stage/provider/model reliability with quality coverage, stable-vs-canary deltas, rollback triggers, and the failure taxonomy. Low confidence is shown as low—not silently filled from incompatible legacy scores—and includes concrete guidance for reaching the target.
- Tool-call-aware recent-tail budgeting
- Exact access-call pruning—different reads, searches, offsets, and patterns do not collapse
- Tool-call/tool-result pair integrity at the compaction boundary
- Collision-safe modified-file verification for monorepos
- Bounded fixed-point repair for patchable verification gaps, followed by a zero-gap deterministic quality floor
- Cross-session guard and five-minute TTL for pending summaries
- Session-log recovery for older, truncated tool results
- Full selected pre-prune conversation backups, scrubbed before a 0600 write; marker-owned retention leaves foreign files in custom directories untouched
- Private artifact directories are enforced as 0700 and files as 0600
High-confidence secret scrubbing is enabled by default at every relevant trust boundary:
provider request · extraction cache · backup · state · context graph · pending summary
It covers common API keys, cloud/GitHub/Slack tokens, JWTs, bearer tokens,
private keys, and credential assignments. Optional email/phone/payment-card
scrubbing is available through scrubPii.
Secret scrubbing is defense in depth, not a replacement for proper secret handling or a dedicated DLP system. See the security policy.
- The manual preflight is the default decision point.
requireApproval: trueadditionally shows a fail-closed verified-summary Apply / Cancel review;falseavoids a redundant second modal. Fingerprint, continuity state, context graph, and success telemetry commit only after the host confirms the matching nativesession_compactevent, and only then does the UI reportApplied. - Online damage monitoring observes the first post-compaction messages and records re-read files or repeated context. Observations join the originating compaction by a local run id; missing evidence lowers coverage rather than counting as a clean run. Remediation hints feed affected files into the next compaction.
adaptiveDamageFeedbackcan opt a project into larger preservation budgets after repeated high-damage reports.
/smart-compact loopsThe manager operates on the project's persisted CompactionState:
- resolve or reopen a loop
- change priority
- pin or unpin it across later compactions
Overrides use normalized summary identity instead of positional IDs, so a loop cannot accidentally inherit another loop's state on a later run.
Add smartCompact to ~/.pi/agent/settings.json:
{
"smartCompact": {
"mode": "auto",
"profile": "balanced",
"summaryModel": null,
"segmentationModel": null,
"verificationModel": null,
"summaryThinkingLevel": "minimal",
"segmentationThinkingLevel": "minimal",
"autoTrigger": true,
"minContextPercent": 60,
"backupEnabled": true,
"scrubSecrets": true,
"scrubPii": false,
"requireApproval": false,
"maxLlmCalls": 8,
"maxLlmInputTokens": 0,
"codexMaxCallMs": 0,
"maxLatencyMs": 0,
"focusWeighting": true,
"zeroCallEnabled": true,
"contextGraphEnabled": true,
"telemetryChannel": "stable",
"onlineDamageMonitor": true,
"adaptiveDamageFeedback": false,
"pinPaths": []
}
}Exploration can use a cheaper reasoning level while final synthesis and repair use a stronger one:
{
"smartCompact": {
"segmentationThinkingLevel": "low",
"summaryThinkingLevel": "high"
}
}segmentationThinkingLevel applies to exploration; summaryThinkingLevel
applies to synthesis, assembly, and repair. Both default to minimal because
reasoning tokens from multi-call compaction add up quickly. Supported values
are minimal, low, medium, high, xhigh, and max. Set either value to
null to restore the provider's default behavior. An explicit call-level
reasoning option takes precedence.
Automatic and tool-triggered runs operate on Pi's current active context, not
the append-only session history. A same-session staged summary is reused, the
exploration loop is limited to three rounds, provider and outer retries are
disabled, and every mode has finite call plus aggregate prompt-token budgets.
Complete tool-call/result pairs are the hard window boundary. Recent user
turns, pi-toolkit checkpoints, and topical grouping are soft and may expand the
raw tail only while remaining inside the selected budget. Automatic/tool runs
normally return to Pi's native compactor without an LLM call when a hard
boundary cannot meet the target. Overflow is the safety exception: EESV keeps
chunked recovery rather than resending an oversized one-shot prompt to native
compaction. Manual /smart-compact uses an absolute adaptive tail rather than
a percentage of a large model window. A plan below 10% projected savings never
starts; if the measured final summary misses the same yield/target contract,
the run fails closed before staging or apply.
All configuration keys
| Key | Type | Default | Notes |
|---|---|---|---|
mode |
auto | fast | balanced | thorough |
auto |
Automatic selector or one of the three execution modes |
profile |
light | balanced | aggressive |
balanced |
Legacy/advanced compression profile; used when mode is absent |
summaryModel |
string | null |
null |
Uses the active session model when null |
segmentationModel |
string | null |
null |
Optional explicit model for Explore |
verificationModel |
string | null |
null |
Optional explicit model for LLM verification repair |
summaryThinkingLevel |
minimal | low | medium | high | xhigh | max | null |
minimal |
Reasoning level for synthesis and repair; provider default when null |
segmentationThinkingLevel |
minimal | low | medium | high | xhigh | max | null |
minimal |
Reasoning level for exploration; provider default when null |
autoTrigger |
boolean |
true |
Participate in Pi's native compact hook |
autoTriggerTimeoutMs |
number |
120000 |
Auto cancellation deadline; waits for safe pipeline unwind, never an unsafe hard return |
minContextPercent |
number |
60 |
Auto/tool context gate; manual /smart-compact warns and bypasses it |
backupEnabled |
boolean |
true |
Write a pre-compaction backup |
backupDir |
string |
~/.pi/agent/compact-backups |
Empty config value uses this path |
profiles |
object | built-ins | Per-profile numeric overrides |
pinPaths |
string[] |
[] |
Always preserve matching paths |
requireApproval |
boolean |
false |
Manual UI only; cancel/error fails closed |
scrubSecrets |
boolean |
true |
High-confidence credential redaction |
scrubPii |
boolean |
false |
Email/phone/card-shaped redaction |
maxLlmCalls |
integer 0–100 |
8 |
Global ceiling combined with the selected mode |
maxLlmInputTokens |
integer 0–1000000 |
0 |
0 uses the selected mode's aggregate prompt-token cap |
codexMaxCallMs |
integer 0 or 5000–300000 |
0 |
ChatGPT Codex per-call watchdog; 0 derives 15–90s from requested output tokens |
maxLatencyMs |
0 or 5000–600000 |
0 |
Pipeline cancellation deadline; 0 means unlimited |
focusWeighting |
boolean |
true |
Weight focused topics/paths higher |
zeroCallEnabled |
boolean |
true |
Use deterministic synthesis for high-confidence Fast runs |
contextGraphEnabled |
boolean |
true |
Index verified state and enable project-scoped recall/save tools |
telemetryChannel |
stable | canary |
stable |
Tag local schema-v2 metrics for external canary comparison |
onlineDamageMonitor |
boolean |
true |
Observe post-compaction regression signals |
adaptiveDamageFeedback |
boolean |
false |
Increase preservation after repeated damage |
The legacy semanticCompact root key is still accepted for compatibility.
Show canonical output
## Goal
Tighten aggregate token budgets without breaking cancellation.
## Constraints & Preferences
- [requirement] Never compact mid-turn from the tool path.
## Progress
### Done
- [x] Reserved concurrent output budgets before provider calls.
### In Progress
- [ ] Collect canary evidence for the new limits.
### Blocked
- None.
## Key Decisions
- **Charge failed streams conservatively**: an interrupted stream consumes its output reservation.
## Files Modified
- src/infra/services.ts
- src/utils/cache.ts
## Open Loops
- [high] Verify provider usage reconciliation across cache-read/write responses.
## Changes Since Last Compaction
- Concurrent output accounting now fails closed.
## Next Steps
1. Run the adversarial release gate.
## Critical Context
- Input accounting includes uncached input, cache reads, and cache writes./smart-compact metrics # text report
/smart-compact dashboard # interactive TUI; can write a local HTML report
/smart-compact restore # browse, inspect, and restore backupsMetrics include effective mode, profile, provider, phase timing, token/call estimates, verification quality, cache behavior, redactions, adaptation, fallbacks, and cancelled runs.
Runtime artifacts
Default artifacts live under ~/.pi/agent/. Smart Compact normalizes the
private directories it creates to 0700 and its files to 0600; a custom
backupDir receives the same protection. settings.json remains host-owned
and read-only to the extension.
| Path | Purpose |
|---|---|
settings.json |
Configuration (read only) |
compact-backups/ |
Full selected pre-prune conversation backups, scrubbed before write and retention-pruned |
.cache/compact-extraction-<session>.json |
Incremental extraction cache |
.cache/compact-metrics.jsonl |
Tail-retained metrics log; 5 MiB cap |
.cache/smart-compact-report.html |
Local HTML dashboard |
.cache/smart-compact/projects/<projectId>.json |
Project fingerprint |
.cache/smart-compact/states/<projectId>/<sessionId>.json |
Scoped compaction state and loop overrides |
.cache/smart-compact/run-locks/ |
0600 cross-process session/global concurrency leases |
.cache/smart-compact/native-continuity/ |
0600 one-shot project/session/branch handoffs |
.cache/smart-compact/context-graph.sqlite |
Project-isolated FTS5 context graph and explicit saved memory |
.cache/smart-compact/damage-reports.jsonl |
Damage reports; 5 MiB cap |
.cache/smart-compact/remediation-<projectId>.json |
Files to preserve after damage |
Pi core packages are host-provided wildcard peers and are excluded from the
published bundle. The lockfile gives contributors a reproducible baseline,
while CI validates the latest Pi release daily without changing the manifest.
An exact version can be checked with bun run compat:pi <version>.
pi-smart-compact is designed to coexist with
pi-toolkit: toolkit handles daily
context hygiene; smart-compact handles high-pressure verified compaction. If
another extension also owns session_before_compact or rewrites branch history,
coordinate hook order or prefer a single automatic compaction owner.
bun install --frozen-lockfile
bun run release:check # typecheck + tests + adversarial gate + build + package audit
bun run bench
bun run compat:pi # isolated latest-Pi compatibility checkPull requests run the same deterministic checks in GitHub Actions. See CONTRIBUTING.md for focused test commands and docs/RELEASE.md for publication and canary gates.
MIT © Alper Tarhan