VT Code uses a local-first performance workflow. Performance checks are measured manually and are not hard CI gates. The default stance is simple: do not guess, measure first, and only keep complexity that pays for itself.
- Keep release artifacts portable.
- Improve runtime without hurting day-to-day iteration speed.
- Optimize only measured hotspots.
- Do not guess where time goes. Capture a baseline before changing code that claims a performance win.
- Measure before tuning. Keep before/after numbers from
baseline.sh, targeted timers, or benchmarks. - Prefer simple algorithms when input sizes are small or not yet proven large.
- Avoid fancy algorithms and broad refactors unless measurements justify their constant-factor and maintenance cost.
- Start with data structures and layout. In VT Code, the right cache shape, queue boundary, or representation usually matters more than clever control flow.
These rules apply to product code and refactors alike. The burden of proof is on the optimization, not on the simpler baseline.
Independent code-search backends (literal, declaration, and path search) are
started together with tokio::join!; filesystem reads, tree-sitter parsing,
and candidate aggregation run in one spawn_blocking task. Keep the async
coordinator responsible for ordering and cancellation, not synchronous disk
work. For synchronous side-channel APIs such as progress monitoring, use a
bounded coalescing writer so callers replace stale snapshots and never wait on
filesystem latency.
Avoid combining #[serde(flatten)] with #[serde(untagged)] on frequent,
discriminator-driven protocol payloads. Serde must buffer the surrounding map
to decide which flattened shape applies; direct wire structs can decode the
known fields once and construct the tagged payload afterward. VT Code uses
this for OpenResponses and ACP streaming notifications. Keep flattening when
it is the actual contract, such as trace metadata's vendor-extension map.
Keep streaming payloads as borrowed SSE text until a consumer needs an owned
payload. The normalized Responses adapter now avoids the old
Value -> JSON -> Value round trip; common text, reasoning, tool, and
lifecycle events use typed decoding, with full payload materialization reserved
for completion and compatibility fallbacks.
Session-log index rebuilds have an even narrower requirement: they need the
versioned envelope and event.type, not the full event payload. The rebuild
path therefore skips nested payload materialization, while turn reconstruction
continues to use the canonical VersionedThreadEvent decoder.
# 1) Capture baseline
./scripts/perf/baseline.sh baseline
# 2) Make a targeted change
# 3) Capture latest
./scripts/perf/baseline.sh latest
# 4) Compare results
./scripts/perf/compare.shArtifacts are written to .vtcode/perf/ and include JSON metrics plus raw logs.
The perf harness builds and measures target/release/vtcode, not cargo run
or the debug binary. It clears RUSTC_WRAPPER and
CARGO_BUILD_RUSTC_WRAPPER by default for its cargo steps so local
measurements still work when sccache is configured but unavailable. Set
PERF_KEEP_RUSTC_WRAPPER=1 only when you explicitly want to keep the wrapper.
Use this loop for any non-trivial performance change. Change one thing at a time so the comparison stays attributable.
vtcode's startup-critical work lives in StartupContext::from_cli_args
(src/startup/mod.rs). The perf harness captures separate process and
interactive metrics:
cold_startup_ms— three--versionlaunches from fresh copies in/tmp. This measures a fresh-copy loader/process path; it does not evict the operating system's page cache.warm_startup_ms(also reported asstartup_ms) — eight warm release--versionruns. This is the stable binary/loader signal and does not enterStartupContext::from_cli_args.first_user_io_ms— eight credential-free releasetool-policy statusruns using temporaryHOME, config, and data paths. It exercises the non-interactive startup path without a provider request or real user credentials.interactive_first_render_ms— three interactive release runs through a PTY. The harness answers terminal capability queries and stops timing when the firstType a requestprompt is rendered; provider response time is not included.
baseline.sh writes raw samples for each metric and compare.sh reports
before/after deltas. Use the same machine and workload configuration for both
runs.
For phase-level diagnostics, set the opt-in trace before launching the binary:
VTCODE_STARTUP_TRACE=1 target/release/vtcode --provider ollama --model llama3The trace is silent when unset and reports only duration records for bootstrap, CLI parsing, runtime creation, config, validation, authentication, session setup, and first UI render. It is initialized before tracing is configured so early startup work is observable without adding work to normal launches.
- Join independent disk I/O.
initialize_dot_folder,init_global_guardian,determine_theme, andresolve_runtime_provider_authonly depend on config that is already resolved; run them throughtokio::join!so their disk reads overlap instead of running serially. - Gate inits behind
command_skips_provider_auth. Commands that never run tools (Login, Logout, Auth, ToolPolicy, AppServer, Notify, Pods, Schedule) do not need the guardian, file/command caches, gatekeeper, session-archive, or perf-telemetry init — skip them entirely. - Keep file reads bounded. The dotfile audit log (
audit.rs::read_last_hash) is append-only and grows unbounded; read only the tail window so startup cost staysO(window), notO(file size). - Defer non-critical background work. Temp-spool cleanup
(
cleanup_old_temp_spools) runs inspawn_blockingso a cold~/.vtcode/tmpnever blocks first user I/O.
The shipped release profile remains tuned for launch size and dead-code
removal: opt-level = "z", full LTO, codegen-units = 1, stripping, and an
abort-on-panic runtime. macOS release scripts and .cargo/config.toml also
apply -Wl,-dead_strip. Verify the effective profile and the measured binary
size before attributing a result to Rust startup code; a debug binary is not a
valid proxy for the shipped launch path.
Cold and warm results answer different questions. Warm results isolate process and loader overhead after the binary is resident. Fresh-copy results expose the size and relocation cost paid by a newly spawned process, which is the relevant signal for subprocess-heavy workflows. Interactive results additionally include configuration, authentication, terminal initialization, and session setup through the first usable frame.
The default binary links heavy subsystems that most invocations never use:
vtcode-eval— eval framework (onlyvtcode evalcommands).vtcode-acp— Agent Client Protocol (onlyvtcode acp).- transitively via
vtcode-core:vtcode-indexer,vtcode-mcp,vtcode-a2a,vtcode-skills.
These are potential binary-size levers. Cutting them requires feature-gating them out of the default binary (and behind an opt-in feature for the commands that need them). That is a product decision, so it is intentionally not done silently. Measure cold-start impact with:
./scripts/perf/baseline.sh latestUse this when collecting profiler traces:
./scripts/perf/profile.shThis builds release with:
-C force-frame-pointers=yesCARGO_PROFILE_RELEASE_DEBUG=line-tables-only
Then profile target/release/vtcode with your preferred tool.
For local experiments only:
./scripts/perf/native-build.sh
./scripts/perf/native-run.sh -- --versionThese scripts append -C target-cpu=native for local runs only. They do not change portable release defaults.
Current Criterion benches:
cargo bench -p vtcode-core --bench tool_pipeline
cargo bench -p vtcode-core --bench agent_harnessUse benches when a hotspot is stable and repeatable. Use the baseline/profile scripts when the question is broader end-to-end behavior.
The agent_harness target measures the repeated work that affects interactive
requests: warm prompt-resource cache hits, few-shot tag selection, tool
definition sorting during catalog refresh, and scoring against a warm indexed
file list. Its fixtures are deterministic and keep filesystem setup outside the
timed iterations.
Prompt resources use canonical source paths, a five-minute bounded cache, and a
two-second metadata polling interval. Cache misses perform scans, reads, and
parsing on Tokio's blocking pool; warm prompt assembly does not reread or
reparse unchanged resources. Indexed searches read an immutable path-text table
by StringId; incremental updates publish a new table so searches holding an
older index remain safe.
Workspace tool registration also retains the effective persistent-memory
configuration from its startup parse. Memory tool calls no longer reload and
reparse vtcode.toml; this follows the same snapshot policy as the web-tool
and output-spooler settings while preserving the disabled-memory guard.
Basic directory-list cache keys include the canonical workspace and every response-shaping filter and pagination value. A cached listing therefore stays local to its workspace and cannot satisfy a request with a different list shape; the async path reuses its single metadata result for directory checks.
The same target also includes uncached filesystem workloads:
agent_harness_file_search_uncached measures parallel traversal and bounded
candidate aggregation, while agent_harness_file_index_build measures a full
index construction per iteration. agent_harness_tool_catalog_projection_repeat
measures repeated schema/model-tool projection after the catalog is warm; its
projection cache is private to an immutable catalog and keyed by documentation
mode. These benchmarks expose repeated work and synchronization cost rather
than serving as universal CI thresholds.
Code-search changes should be checked for both backend overlap and blocking pool behavior; progress-ledger changes should be checked for bounded queue growth and latest-snapshot semantics before comparing end-to-end medians.
Compare repeated local medians rather than adding a noisy hard gate:
./scripts/perf/baseline.sh baseline
./scripts/perf/baseline.sh latest
./scripts/perf/compare.shRustc-specific AST shrinking, compiler incremental-cache changes, and PGO are outside this runtime-focused wave. Revisit them only with a confirmed VT Code profile hotspot and a separate build-performance budget.
- Change one thing at a time.
- Keep changes surgical and behavior-preserving.
- Prefer simple, safe single-pass reductions over broad refactors.
- Revisit data structures before introducing algorithmic sophistication.
- Keep the simplest implementation until measured workload data proves it insufficient.
- For hashers, follow the selective policy in
performance-hasher-policy.md.