Skip to content

Latest commit

 

History

History
255 lines (193 loc) · 11.3 KB

File metadata and controls

255 lines (193 loc) · 11.3 KB

Performance Optimization

VT Code uses a local-first performance workflow. Performance checks are measured manually and are not hard CI gates. The default stance is simple: do not guess, measure first, and only keep complexity that pays for itself.

Goals

  • Keep release artifacts portable.
  • Improve runtime without hurting day-to-day iteration speed.
  • Optimize only measured hotspots.

Performance & Simplicity Rules

  • Do not guess where time goes. Capture a baseline before changing code that claims a performance win.
  • Measure before tuning. Keep before/after numbers from baseline.sh, targeted timers, or benchmarks.
  • Prefer simple algorithms when input sizes are small or not yet proven large.
  • Avoid fancy algorithms and broad refactors unless measurements justify their constant-factor and maintenance cost.
  • Start with data structures and layout. In VT Code, the right cache shape, queue boundary, or representation usually matters more than clever control flow.

These rules apply to product code and refactors alike. The burden of proof is on the optimization, not on the simpler baseline.

Bounded I/O on the agent hot path

Independent code-search backends (literal, declaration, and path search) are started together with tokio::join!; filesystem reads, tree-sitter parsing, and candidate aggregation run in one spawn_blocking task. Keep the async coordinator responsible for ordering and cancellation, not synchronous disk work. For synchronous side-channel APIs such as progress monitoring, use a bounded coalescing writer so callers replace stale snapshots and never wait on filesystem latency.

Serialization and event-log replay

Avoid combining #[serde(flatten)] with #[serde(untagged)] on frequent, discriminator-driven protocol payloads. Serde must buffer the surrounding map to decide which flattened shape applies; direct wire structs can decode the known fields once and construct the tagged payload afterward. VT Code uses this for OpenResponses and ACP streaming notifications. Keep flattening when it is the actual contract, such as trace metadata's vendor-extension map.

Keep streaming payloads as borrowed SSE text until a consumer needs an owned payload. The normalized Responses adapter now avoids the old Value -> JSON -> Value round trip; common text, reasoning, tool, and lifecycle events use typed decoding, with full payload materialization reserved for completion and compatibility fallbacks.

Session-log index rebuilds have an even narrower requirement: they need the versioned envelope and event.type, not the full event payload. The rebuild path therefore skips nested payload materialization, while turn reconstruction continues to use the canonical VersionedThreadEvent decoder.

Local Workflow

# 1) Capture baseline
./scripts/perf/baseline.sh baseline

# 2) Make a targeted change

# 3) Capture latest
./scripts/perf/baseline.sh latest

# 4) Compare results
./scripts/perf/compare.sh

Artifacts are written to .vtcode/perf/ and include JSON metrics plus raw logs.

The perf harness builds and measures target/release/vtcode, not cargo run or the debug binary. It clears RUSTC_WRAPPER and CARGO_BUILD_RUSTC_WRAPPER by default for its cargo steps so local measurements still work when sccache is configured but unavailable. Set PERF_KEEP_RUSTC_WRAPPER=1 only when you explicitly want to keep the wrapper.

Use this loop for any non-trivial performance change. Change one thing at a time so the comparison stays attributable.

Startup budget

vtcode's startup-critical work lives in StartupContext::from_cli_args (src/startup/mod.rs). The perf harness captures separate process and interactive metrics:

  • cold_startup_ms — three --version launches from fresh copies in /tmp. This measures a fresh-copy loader/process path; it does not evict the operating system's page cache.
  • warm_startup_ms (also reported as startup_ms) — eight warm release --version runs. This is the stable binary/loader signal and does not enter StartupContext::from_cli_args.
  • first_user_io_ms — eight credential-free release tool-policy status runs using temporary HOME, config, and data paths. It exercises the non-interactive startup path without a provider request or real user credentials.
  • interactive_first_render_ms — three interactive release runs through a PTY. The harness answers terminal capability queries and stops timing when the first Type a request prompt is rendered; provider response time is not included.

baseline.sh writes raw samples for each metric and compare.sh reports before/after deltas. Use the same machine and workload configuration for both runs.

For phase-level diagnostics, set the opt-in trace before launching the binary:

VTCODE_STARTUP_TRACE=1 target/release/vtcode --provider ollama --model llama3

The trace is silent when unset and reports only duration records for bootstrap, CLI parsing, runtime creation, config, validation, authentication, session setup, and first UI render. It is initialized before tracing is configured so early startup work is observable without adding work to normal launches.

Patterns that pay off on the startup path

  • Join independent disk I/O. initialize_dot_folder, init_global_guardian, determine_theme, and resolve_runtime_provider_auth only depend on config that is already resolved; run them through tokio::join! so their disk reads overlap instead of running serially.
  • Gate inits behind command_skips_provider_auth. Commands that never run tools (Login, Logout, Auth, ToolPolicy, AppServer, Notify, Pods, Schedule) do not need the guardian, file/command caches, gatekeeper, session-archive, or perf-telemetry init — skip them entirely.
  • Keep file reads bounded. The dotfile audit log (audit.rs::read_last_hash) is append-only and grows unbounded; read only the tail window so startup cost stays O(window), not O(file size).
  • Defer non-critical background work. Temp-spool cleanup (cleanup_old_temp_spools) runs in spawn_blocking so a cold ~/.vtcode/tmp never blocks first user I/O.

Release artifact assumptions

The shipped release profile remains tuned for launch size and dead-code removal: opt-level = "z", full LTO, codegen-units = 1, stripping, and an abort-on-panic runtime. macOS release scripts and .cargo/config.toml also apply -Wl,-dead_strip. Verify the effective profile and the measured binary size before attributing a result to Rust startup code; a debug binary is not a valid proxy for the shipped launch path.

Cold and warm results answer different questions. Warm results isolate process and loader overhead after the binary is resident. Fresh-copy results expose the size and relocation cost paid by a newly spawned process, which is the relevant signal for subprocess-heavy workflows. Interactive results additionally include configuration, authentication, terminal initialization, and session setup through the first usable frame.

The default binary links heavy subsystems that most invocations never use:

  • vtcode-eval — eval framework (only vtcode eval commands).
  • vtcode-acp — Agent Client Protocol (only vtcode acp).
  • transitively via vtcode-core: vtcode-indexer, vtcode-mcp, vtcode-a2a, vtcode-skills.

These are potential binary-size levers. Cutting them requires feature-gating them out of the default binary (and behind an opt-in feature for the commands that need them). That is a product decision, so it is intentionally not done silently. Measure cold-start impact with:

./scripts/perf/baseline.sh latest

Profiling Build

Use this when collecting profiler traces:

./scripts/perf/profile.sh

This builds release with:

  • -C force-frame-pointers=yes
  • CARGO_PROFILE_RELEASE_DEBUG=line-tables-only

Then profile target/release/vtcode with your preferred tool.

Local Native Tuning

For local experiments only:

./scripts/perf/native-build.sh
./scripts/perf/native-run.sh -- --version

These scripts append -C target-cpu=native for local runs only. They do not change portable release defaults.

Benchmarks

Current Criterion benches:

cargo bench -p vtcode-core --bench tool_pipeline
cargo bench -p vtcode-core --bench agent_harness

Use benches when a hotspot is stable and repeatable. Use the baseline/profile scripts when the question is broader end-to-end behavior.

Interactive latency workloads

The agent_harness target measures the repeated work that affects interactive requests: warm prompt-resource cache hits, few-shot tag selection, tool definition sorting during catalog refresh, and scoring against a warm indexed file list. Its fixtures are deterministic and keep filesystem setup outside the timed iterations.

Prompt resources use canonical source paths, a five-minute bounded cache, and a two-second metadata polling interval. Cache misses perform scans, reads, and parsing on Tokio's blocking pool; warm prompt assembly does not reread or reparse unchanged resources. Indexed searches read an immutable path-text table by StringId; incremental updates publish a new table so searches holding an older index remain safe.

Workspace tool registration also retains the effective persistent-memory configuration from its startup parse. Memory tool calls no longer reload and reparse vtcode.toml; this follows the same snapshot policy as the web-tool and output-spooler settings while preserving the disabled-memory guard.

Basic directory-list cache keys include the canonical workspace and every response-shaping filter and pagination value. A cached listing therefore stays local to its workspace and cannot satisfy a request with a different list shape; the async path reuses its single metadata result for directory checks.

The same target also includes uncached filesystem workloads: agent_harness_file_search_uncached measures parallel traversal and bounded candidate aggregation, while agent_harness_file_index_build measures a full index construction per iteration. agent_harness_tool_catalog_projection_repeat measures repeated schema/model-tool projection after the catalog is warm; its projection cache is private to an immutable catalog and keyed by documentation mode. These benchmarks expose repeated work and synchronization cost rather than serving as universal CI thresholds.

Code-search changes should be checked for both backend overlap and blocking pool behavior; progress-ledger changes should be checked for bounded queue growth and latest-snapshot semantics before comparing end-to-end medians.

Compare repeated local medians rather than adding a noisy hard gate:

./scripts/perf/baseline.sh baseline
./scripts/perf/baseline.sh latest
./scripts/perf/compare.sh

Rustc-specific AST shrinking, compiler incremental-cache changes, and PGO are outside this runtime-focused wave. Revisit them only with a confirmed VT Code profile hotspot and a separate build-performance budget.

Optimization Rules

  • Change one thing at a time.
  • Keep changes surgical and behavior-preserving.
  • Prefer simple, safe single-pass reductions over broad refactors.
  • Revisit data structures before introducing algorithmic sophistication.
  • Keep the simplest implementation until measured workload data proves it insufficient.
  • For hashers, follow the selective policy in performance-hasher-policy.md.