Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 32 additions & 4 deletions docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,9 @@ runtime, coordination, and SQLite-backed state — the agent itself just reads
and writes files via its own tools and calls ferrus's MCP server to drive
task transitions. Backend-specific behavior lives in `src/agents/{claude,
codex, qwen, opencode, goose}` and is normalized behind shared Supervisor/
Executor contracts.
Executor contracts. The exception is [Nano](/docs/nano), ferrus's own
native harness. It lives in `src/nano/` and calls ferrus operations directly
instead of going through MCP.

## Backends

Expand All @@ -23,9 +25,11 @@ Executor contracts.
| **Qwen Code** | experimental | `.qwen/settings.json` |
| **goose** | experimental | none — attached at launch via `--with-extension` |
| **opencode** | experimental | `opencode.json` |
| **Nano** | experimental · headless Executor only | `[hq.executor]` in `ferrus.toml`; creates `~/.ferrus/nano.toml` on first registration |

Each backend loads `ferrus serve` as an MCP server so its tool calls flow
back into the ferrus task state machine.
Each external backend loads `ferrus serve` as an MCP server so its tool
calls flow back into the ferrus task state machine. Nano needs no MCP
server, because it drives the same operations natively.

### goose

Expand Down Expand Up @@ -56,6 +60,21 @@ provisions — it may operate on the canonical checkout instead. Use opencode
for the **supervisor/reviewer** role for now; treat the executor role as
not yet supported.

### Nano

New in 0.5.0-alpha.1. Nano is ferrus's built-in agent harness and talks to
an OpenAI-compatible Chat Completions endpoint, such as a local LM Studio
server. It runs **only as a headless Executor**, so pair it with an external
supervisor:

```bash
ferrus register --supervisor claude-code --executor nano --executor-model YOUR_LOADED_MODEL_ID
```

Prebuilt release archives include Nano. A `cargo install` needs
`--features nano-openai,nano-mcp`. See [Nano](/docs/nano) for setup,
provider settings, and what it can do today.

## Roles

Each task runs up to three roles. A single backend can play all three, or
Expand All @@ -68,6 +87,9 @@ you can mix and match:
exits after approve/reject. In practice, the Reviewer is the Supervisor
backend relaunched in review mode.

Nano can only play the **Executor**. Registering it as `--supervisor` is
rejected.

## Register

```bash
Expand All @@ -80,7 +102,10 @@ ferrus register \

Model overrides are optional — omit them to use each agent's default. You
can also change them interactively from HQ with `/model`, or leave a model
unset and target a local backend (goose/opencode) for cost-free iteration.
unset and target a local backend (goose/opencode/Nano) for cost-free
iteration. Nano is the exception: its first registration needs
`--executor-model`, because ferrus writes that model into the new
`nano.toml`.

## Tools exposed per role

Expand All @@ -98,6 +123,9 @@ physically cannot call `approve`, and a supervisor physically cannot call
`submit`. This is what makes the loop safe to drive from "untrusted"
agents.

Nano doesn't use `ferrus serve`. It gets its own executor-only native tool
set, listed in [What Nano can do today](/docs/nano#what-nano-can-do-today).

### Retrieval tools

When the optional [repository graph](/docs/repository-graph) is configured,
Expand Down
8 changes: 7 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ agent = "claude-code" # agent for supervisor/reviewer role: claude-code | cod
model = "" # optional override; empty = agent default

[hq.executor]
agent = "codex" # agent for executor role: claude-code | codex | qwen-code | goose (experimental); opencode executor is experimental/unstable
agent = "codex" # agent for executor role: claude-code | codex | qwen-code | goose (experimental) | nano (experimental, headless only); opencode executor is experimental/unstable
model = ""
```

Expand Down Expand Up @@ -237,6 +237,12 @@ Use `/model` inside HQ to update model overrides interactively. See
[Supported agents](/docs/agents) for the full backend list, including the
experimental `goose` and `opencode` adapters.

`agent = "nano"` is valid only under `[hq.executor]`. Nano's provider
settings (endpoint, model, token budgets) live outside the repository, in an
owner-only `~/.ferrus/nano.toml` or the file named by `FERRUS_NANO_CONFIG`.
The `model` key here is an optional override of that file's model. See
[Nano](/docs/nano#provider-settings).

## Runtime files

Ferrus separates human-readable project artifacts from machine-local runtime
Expand Down
2 changes: 1 addition & 1 deletion docs/hq.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ and retry/cycle counters in real time.
| `/run [--limit N]` | Plan a **batch** run: queue tasks for every ready milestone in the selected spec (or up to `--limit`) and let the scheduler dispatch executors for all of them, bounded by `max_parallel_tasks`. |
| `/check` | Run the ferrus check gate from HQ. Supports `--force` to run regardless of task status. |
| `/supervisor` | Open an interactive supervisor session (no initial prompt). |
| `/executor` | Open an interactive executor session (no initial prompt). |
| `/executor` | Open an interactive executor session (no initial prompt). Not available when the executor is [Nano](/docs/nano), which runs headless only. |
| `/resume` | Manually resume the executor headlessly; also recovers consultation by relaunching both supervisor and executor. |
| `/review` | Manually spawn supervisor in review mode (escape hatch when automatic spawning failed). |
| `/status` | Show task state, agent list, and session log paths. |
Expand Down
25 changes: 22 additions & 3 deletions docs/local-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,15 @@ slug: /local-models

# Tuning local models

`goose` is ferrus's local-model-friendly backend: it's MCP-native, needs no
project config file, and works well against a local LM Studio or Ollama
endpoint set with `goose configure` (see [Supported agents](/docs/agents)).
There are two local-model-friendly executors:

- **[Nano](/docs/nano)** (new in 0.5): ferrus's own native harness. It talks
directly to LM Studio's OpenAI-compatible endpoint and runs headless as
the executor only.
- **`goose`**: MCP-native, needs no project config file, and works well
against a local LM Studio or Ollama endpoint set with `goose configure`.
See [Supported agents](/docs/agents).

`opencode` can also drive a local model, but only for the supervisor/
reviewer role today.

Expand Down Expand Up @@ -52,6 +58,12 @@ These numbers are **Qwen3-specific** — they come from Qwen's own guidance,
not from ferrus. For another model family (Gemma included) start from that
model's own recommended defaults instead of copying this table verbatim.

:::note[Nano]
`nano.toml` accepts only `temperature` from this table. It rejects unknown
keys, so set `top_p`, `top_k`, `min_p`, and `presence_penalty` as the
loaded model's defaults in LM Studio.
:::

A temperature noticeably below 0.6 (e.g. 0.4) is *below* what Qwen3 was
tuned for, and the model's own guidance warns that low temperature in
thinking mode provokes looping and repeated output. If you're seeing
Expand Down Expand Up @@ -103,6 +115,13 @@ These are easy to conflate and have opposite recommendations:
for anything beyond trivial edge cases — it gives a real correctness
boost on subtle contract-level bugs. Double-check your runtime (goose,
LM Studio, …) isn't silently disabling it.

Nano is a deliberate exception. Its generated `nano.toml` sets
`reasoning_effort = "none"`, because reasoning shares one output allowance
with the tool calls and patch text. To turn thinking on, set
`reasoning_effort` (for example `"medium"`). If `ferrus --debug` then
shows frequent output-limit continuations, raise `max_output_tokens`, or
remove it so Nano allocates the output budget automatically.
- **Preserving thinking** — carrying previous turns' `<think>` blocks
forward in multi-turn history — should stay **off**. Qwen3's own
guidance says to strip reasoning from history in multi-turn use: the
Expand Down
Loading
Loading