diff --git a/docs/agents.md b/docs/agents.md index 681bd9b..aeeba65 100644 --- a/docs/agents.md +++ b/docs/agents.md @@ -12,7 +12,9 @@ runtime, coordination, and SQLite-backed state — the agent itself just reads and writes files via its own tools and calls ferrus's MCP server to drive task transitions. Backend-specific behavior lives in `src/agents/{claude, codex, qwen, opencode, goose}` and is normalized behind shared Supervisor/ -Executor contracts. +Executor contracts. The exception is [Nano](/docs/nano), ferrus's own +native harness. It lives in `src/nano/` and calls ferrus operations directly +instead of going through MCP. ## Backends @@ -23,9 +25,11 @@ Executor contracts. | **Qwen Code** | experimental | `.qwen/settings.json` | | **goose** | experimental | none — attached at launch via `--with-extension` | | **opencode** | experimental | `opencode.json` | +| **Nano** | experimental · headless Executor only | `[hq.executor]` in `ferrus.toml`; creates `~/.ferrus/nano.toml` on first registration | -Each backend loads `ferrus serve` as an MCP server so its tool calls flow -back into the ferrus task state machine. +Each external backend loads `ferrus serve` as an MCP server so its tool +calls flow back into the ferrus task state machine. Nano needs no MCP +server, because it drives the same operations natively. ### goose @@ -56,6 +60,21 @@ provisions — it may operate on the canonical checkout instead. Use opencode for the **supervisor/reviewer** role for now; treat the executor role as not yet supported. +### Nano + +New in 0.5.0-alpha.1. Nano is ferrus's built-in agent harness and talks to +an OpenAI-compatible Chat Completions endpoint, such as a local LM Studio +server. It runs **only as a headless Executor**, so pair it with an external +supervisor: + +```bash +ferrus register --supervisor claude-code --executor nano --executor-model YOUR_LOADED_MODEL_ID +``` + +Prebuilt release archives include Nano. A `cargo install` needs +`--features nano-openai,nano-mcp`. See [Nano](/docs/nano) for setup, +provider settings, and what it can do today. + ## Roles Each task runs up to three roles. A single backend can play all three, or @@ -68,6 +87,9 @@ you can mix and match: exits after approve/reject. In practice, the Reviewer is the Supervisor backend relaunched in review mode. +Nano can only play the **Executor**. Registering it as `--supervisor` is +rejected. + ## Register ```bash @@ -80,7 +102,10 @@ ferrus register \ Model overrides are optional — omit them to use each agent's default. You can also change them interactively from HQ with `/model`, or leave a model -unset and target a local backend (goose/opencode) for cost-free iteration. +unset and target a local backend (goose/opencode/Nano) for cost-free +iteration. Nano is the exception: its first registration needs +`--executor-model`, because ferrus writes that model into the new +`nano.toml`. ## Tools exposed per role @@ -98,6 +123,9 @@ physically cannot call `approve`, and a supervisor physically cannot call `submit`. This is what makes the loop safe to drive from "untrusted" agents. +Nano doesn't use `ferrus serve`. It gets its own executor-only native tool +set, listed in [What Nano can do today](/docs/nano#what-nano-can-do-today). + ### Retrieval tools When the optional [repository graph](/docs/repository-graph) is configured, diff --git a/docs/configuration.md b/docs/configuration.md index 36599af..9587f9a 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -43,7 +43,7 @@ agent = "claude-code" # agent for supervisor/reviewer role: claude-code | cod model = "" # optional override; empty = agent default [hq.executor] -agent = "codex" # agent for executor role: claude-code | codex | qwen-code | goose (experimental); opencode executor is experimental/unstable +agent = "codex" # agent for executor role: claude-code | codex | qwen-code | goose (experimental) | nano (experimental, headless only); opencode executor is experimental/unstable model = "" ``` @@ -237,6 +237,12 @@ Use `/model` inside HQ to update model overrides interactively. See [Supported agents](/docs/agents) for the full backend list, including the experimental `goose` and `opencode` adapters. +`agent = "nano"` is valid only under `[hq.executor]`. Nano's provider +settings (endpoint, model, token budgets) live outside the repository, in an +owner-only `~/.ferrus/nano.toml` or the file named by `FERRUS_NANO_CONFIG`. +The `model` key here is an optional override of that file's model. See +[Nano](/docs/nano#provider-settings). + ## Runtime files Ferrus separates human-readable project artifacts from machine-local runtime diff --git a/docs/hq.md b/docs/hq.md index 1c74b6d..5fb66e4 100644 --- a/docs/hq.md +++ b/docs/hq.md @@ -27,7 +27,7 @@ and retry/cycle counters in real time. | `/run [--limit N]` | Plan a **batch** run: queue tasks for every ready milestone in the selected spec (or up to `--limit`) and let the scheduler dispatch executors for all of them, bounded by `max_parallel_tasks`. | | `/check` | Run the ferrus check gate from HQ. Supports `--force` to run regardless of task status. | | `/supervisor` | Open an interactive supervisor session (no initial prompt). | -| `/executor` | Open an interactive executor session (no initial prompt). | +| `/executor` | Open an interactive executor session (no initial prompt). Not available when the executor is [Nano](/docs/nano), which runs headless only. | | `/resume` | Manually resume the executor headlessly; also recovers consultation by relaunching both supervisor and executor. | | `/review` | Manually spawn supervisor in review mode (escape hatch when automatic spawning failed). | | `/status` | Show task state, agent list, and session log paths. | diff --git a/docs/local-models.md b/docs/local-models.md index fa305dd..c2e41d6 100644 --- a/docs/local-models.md +++ b/docs/local-models.md @@ -7,9 +7,15 @@ slug: /local-models # Tuning local models -`goose` is ferrus's local-model-friendly backend: it's MCP-native, needs no -project config file, and works well against a local LM Studio or Ollama -endpoint set with `goose configure` (see [Supported agents](/docs/agents)). +There are two local-model-friendly executors: + +- **[Nano](/docs/nano)** (new in 0.5): ferrus's own native harness. It talks + directly to LM Studio's OpenAI-compatible endpoint and runs headless as + the executor only. +- **`goose`**: MCP-native, needs no project config file, and works well + against a local LM Studio or Ollama endpoint set with `goose configure`. + See [Supported agents](/docs/agents). + `opencode` can also drive a local model, but only for the supervisor/ reviewer role today. @@ -52,6 +58,12 @@ These numbers are **Qwen3-specific** — they come from Qwen's own guidance, not from ferrus. For another model family (Gemma included) start from that model's own recommended defaults instead of copying this table verbatim. +:::note[Nano] +`nano.toml` accepts only `temperature` from this table. It rejects unknown +keys, so set `top_p`, `top_k`, `min_p`, and `presence_penalty` as the +loaded model's defaults in LM Studio. +::: + A temperature noticeably below 0.6 (e.g. 0.4) is *below* what Qwen3 was tuned for, and the model's own guidance warns that low temperature in thinking mode provokes looping and repeated output. If you're seeing @@ -103,6 +115,13 @@ These are easy to conflate and have opposite recommendations: for anything beyond trivial edge cases — it gives a real correctness boost on subtle contract-level bugs. Double-check your runtime (goose, LM Studio, …) isn't silently disabling it. + + Nano is a deliberate exception. Its generated `nano.toml` sets + `reasoning_effort = "none"`, because reasoning shares one output allowance + with the tool calls and patch text. To turn thinking on, set + `reasoning_effort` (for example `"medium"`). If `ferrus --debug` then + shows frequent output-limit continuations, raise `max_output_tokens`, or + remove it so Nano allocates the output budget automatically. - **Preserving thinking** — carrying previous turns' `` blocks forward in multi-turn history — should stay **off**. Qwen3's own guidance says to strip reasoning from history in multi-turn use: the diff --git a/docs/nano.md b/docs/nano.md new file mode 100644 index 0000000..7f9ae94 --- /dev/null +++ b/docs/nano.md @@ -0,0 +1,277 @@ +--- +id: nano +title: Nano (native executor) +sidebar_position: 6 +slug: /nano +--- + +# Nano + +**Nano** is ferrus's own coding-agent harness, new in **0.5.0-alpha.1**. It is +a small Rust agent runtime built into the `ferrus` binary. External backends +reach ferrus through MCP. Nano calls ferrus operations (claim, heartbeat, +check, submit, repository graph, project memory) directly as Rust functions, +with no loopback MCP server or JSON-RPC hop. It talks to a model through an +OpenAI-compatible Chat Completions endpoint. The first target is a local +[LM Studio](https://lmstudio.ai) server. + +:::caution[Experimental: headless Executor only] +In this release Nano runs **only as a headless Executor** launched by HQ. +It cannot be the Supervisor or Reviewer, and interactive `/executor` +sessions are rejected. Pair it with an external supervisor such as +`claude-code` or `codex`. There are no measured performance claims yet, and +coding quality depends heavily on the model you load. +::: + +## Install + +Prebuilt release archives already include Nano, together with external MCP +tool support. The quick-install scripts use these archives: + +```bash +curl -fsSL https://ferrus.dev/cli/install.sh | sh +``` + +Cargo's default feature set does **not** include Nano. From crates.io, enable +both features explicitly: + +```bash +cargo install ferrus --version 0.5.0-alpha.1 --locked --profile dist --features nano-openai,nano-mcp +``` + +When building from source, use `--features nano-openai,nano-mcp`. You can use +`nano-openai` alone if you don't need external MCP tools. A build without +`nano-openai` reports the missing feature before HQ prepares a task. + +## Set up with LM Studio + +1. Start LM Studio's server on `http://127.0.0.1:1234` and load a model that + supports **tool calls**. +2. Register Nano as the executor with the **exact** model ID LM Studio shows: + + ```bash + ferrus register --supervisor claude-code --executor nano --executor-model YOUR_LOADED_MODEL_ID + ``` + +3. Enter HQ and use the normal `/task` or `/run` workflow: + + ```bash + ferrus + ``` + +On the first registration, ferrus creates an owner-only provider file at +`~/.ferrus/nano.toml` and prints the resolved path. On Windows the file lives +in your user profile, for example `C:\Users\Alice\.ferrus\nano.toml`. The file +starts with three settings: + +```toml +base_url = "http://127.0.0.1:1234/v1" +model = "YOUR_LOADED_MODEL_ID" +reasoning_effort = "none" +``` + +Registration never contacts the provider and never replaces an existing +file. It writes `[hq.executor] agent = "nano"` to `ferrus.toml` and creates +no MCP config entry. On later registrations, `--executor-model` overrides the +model from the file. Leave the flag out to clear the override and use the +file's `model`. + +To use a different provider file, set `FERRUS_NANO_CONFIG` to an absolute, +owner-only path. It must be set in the shell that runs **both** `ferrus +register` and HQ: + +```bash +export FERRUS_NANO_CONFIG=/absolute/private/path/nano.toml +ferrus register --executor nano +ferrus +``` + +### Provider settings + +Keep the file outside the repository. On Unix it must be mode `0400` or +`0600`; on Windows it needs a protected owner-only DACL. Unknown keys fail +explicitly, and so does an inline `api_key`. + +| Key | Default | Meaning | +|---|---|---| +| `base_url` | required | Endpoint ending in `/v1`. Plain HTTP is allowed only on loopback; other hosts need HTTPS. | +| `model` | required | The model ID. Nano never picks a model automatically. | +| `reasoning_effort` | server default | `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, or `max`. `none` can turn off reasoning on local thinking models, so the output budget goes to tool calls. | +| `api_key_file` | none | Absolute path to an owner-only file that holds only a Bearer token. Without it, requests carry no `Authorization` header. | +| `temperature` | server default | Optional sampling override. | +| `context_tokens` | none | Optional per-request window used for Nano's own admission and compaction. It is never sent as a model-loading parameter. | +| `session_tokens` | `1000000` | Hard cumulative input + output budget for the whole work phase, including retries and summaries. | +| `max_output_tokens` | automatic | Optional per-response ceiling. Without it, Nano derives a cap from the context and session budgets. | +| `request_timeout_ms` | `120000` | Read-inactivity timeout. Receiving bytes resets it. | +| `include_usage` | `true` | Request streamed usage. Set it to `false` if your server doesn't support it; Nano then estimates usage. | +| `mcp_config_file` | none | Opt-in [external MCP tools](#external-mcp-tools). | +| `working_set_enabled` / `native_context_enabled` | `true` | Switches for evaluation runs. Leave them on for normal use. | + +Three advanced stream limits are also accepted: `wire_bytes`, +`event_bytes`, and `max_tool_calls`. See the +[provider contract](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-provider.md) +for them. `temperature` is the only sampling parameter Nano sends. Set +`top_p`, `top_k`, and penalties as the model's defaults in LM Studio. For +tuning sampling, context size, and quantization on local models, see +[Tuning local models](/docs/local-models). + +## What Nano can do today + +Nano only gets tools that fit the Executor role. It cannot create, approve, +or reset tasks. Lifecycle mutations always pass through the same SQLite +transactions and guards that external agents use via MCP. + +### Code in the task worktree + +- **`read_file` / `search_text`**: bounded reads of whole lines and literal, + case-sensitive search. Each result includes the file's SHA-256 digest. +- **`apply_patch`**: exact, digest-checked edits. A single-file call replaces + one exact substring that must appear exactly once. A batch can create, + update, or delete up to 16 files. A stale digest returns `conflict` and + writes nothing. There is no fuzzy matching or line-number guessing. +- Nano can't reach `.git`, `.ferrus`, ferrus databases, symlinks, or files + outside the worktree. Git staging, commits, and integration stay with + ferrus. + +### Run commands + +- **`exec`, `read_process`, `stop_process`, `read_output`**: start shell + commands in the background, poll them, cancel the whole process tree, and + page through bounded stdout and stderr. +- Commands run with a scrubbed environment. Provider keys, MCP secrets, and + ferrus launch credentials are not inherited. +- Defaults: 4 concurrent commands, 10 minutes per command, 4 MiB of output + per command. + +### Drive the ferrus lifecycle natively + +- **`check`** runs the configured `[checks]` with the same retry accounting + as external agents. **`submit`** runs the check gate and the final review + gate, then hands the task to Reviewing. +- **`consult`** and **`ask_human`** pause the task and wait for the Supervisor + or for you to answer, and return the actual reply. +- Claiming, lease heartbeats, and waiting for answers are handled by the host, + so they cost no model turns. +- A final model sentence never marks a task done. Only Supervisor approval + does. + +### Use repository context + +- Nano exposes the [repository graph](/docs/repository-graph) and + [project memory](/docs/project-memory) retrieval tools natively: + `repository_graph_status`, `repository_search`, `repository_context`, + `project_memory_status`, `project_context_search`, `project_context`. +- **`repository_fallback`** reads or searches the workspace directly when + graph coverage is missing or stale. The model has to state a reason, and + fallback results are never treated as graph facts. +- A **working set** chooses which earlier evidence to resend each turn, + re-checks source digests before reusing anything, and drops evidence for + files that changed. After edits, it refreshes the task's graph overlay with + a short debounce. +- Nano loads **scoped instructions**: root and nested `AGENTS.md` files for + the paths being edited, plus selected `.agents/skills//SKILL.md` + files, with digests and size caps. + +### Stay within budget + +- Before each request Nano measures the payload and reserves room for the + output. If the context is full, it first swaps large old read-only tool + results for short handles that the model can re-run. If that's not enough, + it summarizes older turns. The task and constraints are never dropped. +- A response cut off by the output limit (`finish_reason = "length"`) + continues in the same session. Tool calls from a truncated response are + never executed. +- Rate limits (429), transient 5xx errors, and broken streams are retried + with backoff, from the same session budget. + +### Survive crashes + +- Each run writes a durable journal under + `~/.ferrus/projects//nano/sessions//`. +- When HQ relaunches an interrupted Executor, Nano reconciles every + in-flight effect before inferring again: + - an unfinished patch is confirmed or ruled out by its before/after + digests; + - a lost `submit` is confirmed against the committed SQLite event. +- Effects that can't be proven, such as an unknown shell command, check, or + MCP call, fail the task for manual reconciliation. Nano never replays them + blindly. + +### External MCP tools + +With `nano-mcp`, Nano can call tools from local **stdio** MCP servers that you +list explicitly. Point `mcp_config_file` in `nano.toml` to a second +owner-only file: + +```toml +[[servers]] +id = "local" +command = "/usr/local/bin/my-mcp-server" # must be absolute +args = ["--stdio"] +allow = ["lookup"] # exact tools the model may call +timeout_ms = 10000 + +[servers.env] +PATH = "/usr/local/bin:/usr/bin" +``` + +The model sees these tools as `mcp__`, and they can't +shadow native tools. Each peer gets only the environment you configure. +Arguments are validated against the pinned input schema, and outputs are +capped at 24 KiB. HTTP transport, OAuth, MCP resources, prompts, sampling, +and elicitation are not supported yet. + +## How HQ runs Nano + +For each dispatch, HQ starts `ferrus nano run` in the task's prepared +worktree. HQ still owns the worktree and baseline, process supervision, +dispatch accounting, review, and recovery. Nano receives versioned JSONL +`start`/`cancel` commands on stdin and emits progress events on stdout. + +- `/stop` sends `cancel` first and gives Nano two seconds to clean up. After + that, HQ falls back to killing the process group. +- `ferrus --debug` also shows tool requests, model failures, output-limit + continuations, and terminal outcomes in the HQ transcript. +- Logs go to `.ferrus/logs/executor___.log`. Look for + `Nano diagnostic` entries. The full event stream is in the session's + `events.jsonl`. +- `ferrus nano --version` prints the bundled ferrus version without loading + provider settings. + +### When a run fails + +These failures mark the task **Failed** with a specific code, and Nano exits +non-zero: + +- a non-retryable provider error: `nano_provider_failed` or + `nano_provider_protocol`; +- an exhausted token, turn, tool-call, retry, or no-progress budget: + `nano_limit_*`. + +Budgets belong to the work phase, so restarting Nano doesn't refill them, and +changing settings doesn't silently restart the task. For longer tasks, raise +`session_tokens` (for example `session_tokens = 2000000`) before the next +work phase. + +## Not there yet + +- Interactive sessions (`/executor`), and the Supervisor and Reviewer roles. +- A standalone `ferrus-nano` executable. +- OS-level sandboxing. The only backend is `trusted_local`, so shell commands + run with your user's permissions. The worktree and scrubbed environment are + not a security boundary. +- Providers other than Chat Completions (for example, the Responses API). +- Published benchmarks against external agents. The pinned evaluation suite + exists, but measured comparisons are still pending. + +## Further reading + +The in-repo design notes go deeper than this page: + +- [Launch and HQ events](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-launch.md) +- [Provider contract](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-provider.md) +- [Workspace tools](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-workspace.md) and [command sessions](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-commands.md) +- [Native context](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-context.md), [working set](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-working-set.md), and [compaction](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-compaction.md) +- [Managed lifecycle](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-lifecycle.md) and [sessions and recovery](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-sessions.md) +- [External MCP tools](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-mcp.md) +- [Architecture](https://github.com/ferrus-dev/ferrus/blob/main/docs/ferrus-nano-architecture.md) diff --git a/docs/quickstart.md b/docs/quickstart.md index a2ece33..940bab9 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -31,6 +31,10 @@ cargo install --path . Requires **Rust 1.95+**. ferrus is currently **alpha** — expect rough edges. +The install scripts ship prebuilt release archives that include the +experimental [Nano](/docs/nano) executor. `cargo install` uses the smaller +default feature set. Add `--features nano-openai,nano-mcp` if you want Nano. + ## 2. Scaffold Inside any project directory: @@ -53,8 +57,9 @@ This creates: Tell ferrus which coding agent plays each role. Supported today: `claude-code`, `codex`, `qwen-code` (experimental), `goose` (experimental — -convenient for local models), and `opencode` (experimental — supervisor/ -reviewer role only for now). +convenient for local models), `opencode` (experimental — supervisor/ +reviewer role only for now), and `nano` (experimental — ferrus's own native +harness, headless executor only; see [Nano](/docs/nano)). ```bash ferrus register --supervisor claude-code --executor codex @@ -64,8 +69,10 @@ This writes the right MCP server config (`.claude/mcp-supervisor.json` / `.claude/mcp-executor.json` for Claude Code, `.codex/config.toml` for Codex, `.qwen/settings.json` for Qwen Code, `opencode.json` for opencode) so agents automatically pick up `ferrus serve` as a tool server. goose needs no config -file — ferrus attaches its role-scoped MCP server at launch instead. See -[Supported agents](/docs/agents) for backend-specific notes. +file — ferrus attaches its role-scoped MCP server at launch instead. Nano +needs no MCP config either: it calls ferrus natively and keeps its provider +settings in `~/.ferrus/nano.toml`. See [Supported agents](/docs/agents) for +backend-specific notes. ## 4. Drop into HQ diff --git a/docusaurus.config.ts b/docusaurus.config.ts index 74e8865..f7eefc3 100644 --- a/docusaurus.config.ts +++ b/docusaurus.config.ts @@ -111,6 +111,7 @@ const config: Config = { {label: 'State Machine', to: '/docs/state-machine'}, {label: 'Repository Graph', to: '/docs/repository-graph'}, {label: 'Project Memory', to: '/docs/project-memory'}, + {label: 'Nano', to: '/docs/nano'}, {label: 'Migrating from 0.2.x', to: '/docs/migration'}, {label: 'Local Models', to: '/docs/local-models'}, ], diff --git a/sidebars.ts b/sidebars.ts index 643a0cc..005613b 100644 --- a/sidebars.ts +++ b/sidebars.ts @@ -14,6 +14,7 @@ const sidebars: SidebarsConfig = { items: ['repository-graph', 'project-memory'], }, 'agents', + 'nano', 'migration', 'local-models', ], diff --git a/src/pages/index.tsx b/src/pages/index.tsx index 07c035c..81bbaa1 100644 --- a/src/pages/index.tsx +++ b/src/pages/index.tsx @@ -39,7 +39,7 @@ function Hero() {
- alpha · v0.4.1 + alpha · v0.5.0 Apache-2.0 Rust 1.95+
@@ -71,7 +71,9 @@ function Install() { {`# stable — published on crates.io cargo install ferrus # or pin an exact version: -cargo install --locked ferrus@0.4.1-alpha.1`} +cargo install --locked ferrus@0.5.0-alpha.1 +# with the experimental Nano executor: +cargo install --locked ferrus@0.5.0-alpha.1 --features nano-openai,nano-mcp`}
@@ -112,9 +114,39 @@ ferrus # enter HQ`} +
+ + $ Nano + +

+ New in 0.5: ferrus's own native agent harness. Nano runs as a + headless Executor against a local OpenAI-compatible model such as + LM Studio. It calls ferrus operations (check, submit, the repository + graph, project memory) directly in Rust, with no MCP round-trip. +

+ {`ferrus register --supervisor claude-code \\ + --executor nano --executor-model # creates ~/.ferrus/nano.toml +ferrus # then /task as usual`} +

+ It ships with digest-checked patches, bounded shell sessions, a + token-budgeted working set with compaction, opt-in stdio MCP tools, + and crash recovery that reconciles every in-flight effect. + Experimental: Executor-only and headless for now. +

+

+ Read the Nano docs → +

+
+ + ); +} + +function RepositoryGraph() { + return ( +
$ Repository graph @@ -152,7 +184,7 @@ function Features() { }, { title: 'Agent-agnostic', - body: 'Claude Code, Codex, Qwen Code, goose, and opencode are interchangeable workers — including local models.', + body: 'Claude Code, Codex, Qwen Code, goose, opencode, and the native Nano executor are interchangeable workers — including local models.', }, { title: 'Crash-safe', @@ -172,7 +204,7 @@ function Features() { }, ]; return ( -
+
$ Why ferrus @@ -200,6 +232,7 @@ export default function Home(): ReactNode {
+