Note
Preview only. APIs and design are unstable and subject to change.
A durable Obelisk workflow that is an agent loop. It holds
a provider-neutral chat history and a persistent virtual filesystem, and exposes
exactly one tool to the model: bash.
The shell is a Rust rewrite of just-bash, vendored at
vendor/just-bash-rs. On top of it the workflow adds a VFS that mounts the
active Obelisk deployment and any stateless MCP servers, so their resources show
up as files. Obelisk functions and MCP tools are surfaced as ordinary bash
programs. The same shell is driven by the LLM and, directly, by the user:
type a command with a $ prefix in the UI to run it yourself, e.g. $ pwd.
External programs are Obelisk executions the workflow discovers at session start
from an operator-owned registry (PROGRAMS_JSON), so adding one is a
deployment.toml edit with no workflow rebuild; each entry's description is
surfaced in the system prompt. The one that ships, a GET-only curl, is an
Obelisk activity. The per-turn model invocation limit comes out of the same
registry read: MAX_STEPS, defaulting to 20.
Build the component, pick an LLM catalog, and start the server:
just build
ln -sf models.local.json models.json # pick a catalog
export AGENT_MODELS="$(cat models.json)" # or use direnv (.envrc-example)
export LLM_BASE_URL=http://127.0.0.1:9190 # match the catalog's endpoint
just serve # obelisk server run -d deployment.tomlEverything is pinned by the flake, so nix develop provides the matching
Obelisk and Rust (wasm32) toolchain. The workflow is a native Rust component
(workflow/workflow-rs); no special Obelisk build is required.
Then open http://localhost:9090 (the external/webhook listener; server.toml
keeps the built-in default). Create an empty session to use the shell directly,
or submit a prompt and inspect the same filesystem afterward. The user input
stays live while a completion is pending, so $ commands can edit the session
VFS mid-turn; a prompt sent during the wait is queued for the next model turn.
There are two Obelisk instances in play:
-
The agent instance the workflow-agent runs on. Its own session runs, logs, and the UI's pause/cancel/answer buttons live here, authenticated with
OBELISK__API__TOKEN(andOBELISK_UI_URLfor links). -
The target instance the agent inspects and deploys to. Every control/deploy tool (
obelisk functions|executions|call|deployment ...) and the/workspace/deploymentmount talk to it, configured by three vars that default to the agent instance:export TARGET_OBELISK_API_URL=http://127.0.0.1:5205 export TARGET_OBELISK_API_URL_REGEX="http://127\\.0\\.0\\.1:5205" export TARGET_OBELISK_TOKEN="$OBELISK__API__TOKEN"
A fourth var,
TARGET_OBELISK_WEBHOOK_URL(defaulthttp://127.0.0.1:9290, matchingserver-target.toml), is the target's webhook (external) listener. The shell's GET-only curl is granted access to it, so deployed webhook endpoints can be smoke-tested directly, andmountprints it for discovery.
Point these at a separate Obelisk to deploy somewhere other than the agent's own instance.
Warning
The default targets the agent's own instance, which is recursive and risky.
A deployment the agent applies to itself can remove the agent, or worse remove
the very deployment submit/switch activity that is still running: a
deployment switch closes all executors before it acknowledges, so the
in-flight control execution is left Pending. Re-activating the previous
deployment then immediately switches back to the broken one and re-Pends it. Use
a separate target instance for anything beyond local experimentation.
One endpoint serves the whole catalog, configured by three env vars: the
catalog JSON AGENT_MODELS (required), the origin LLM_BASE_URL (default
http://127.0.0.1:9190), and the bearer LLM_API_KEY (unset for keyless).
Each catalog entry points a model at an OpenAI- or Anthropic-shaped route under
that origin. Three catalogs ship:
-
models.local.json(keyless) : the siblingagent-backed-llm-server, a Claude/Codex subscription in docker on:9190. -
models.exe-integration.json(keyless) : the exe.dev LLM integration —LLM_BASE_URL=https://llm.int.exe.xyz. Inside an attached exe.dev VM, exe.dev authenticates at the network edge and plain OpenAI-compatible requests just work, no key:curl -X POST https://llm.int.exe.xyz/v1/chat/completions \ -H content-type:application/json \ -d '{"model":"...","messages":[{"role":"user","content":"hi"}]}'From outside exe.dev the hostname resolves to a link-local address only exe infra routes, and the https endpoint routes on the Host header, so run an nginx reverse proxy on an attached VM (
examples/exe-llm-proxy.conf) that fixes the header, and tunnel to it:# on <yourinstance>.exe.xyz: mkdir -p /tmp/exe-llm-proxy/{logs,tmp/body} nginx -c $(pwd)/examples/exe-llm-proxy.conf # plain HTTP on 127.0.0.1:7071 # on your machine: ssh -L 7070:127.0.0.1:7071 <yourinstance>.exe.xyz export LLM_BASE_URL=http://localhost:7070 # plain http, no custom CA
-
models.openrouter.json(LLM_API_KEY) : OpenRouter. The key is injected into the outbound header at the edge, never seen by the JS.
Regenerate the exe.dev catalog from the published model list with
node scripts/update-exe-models.mjs. Leave LLM_API_KEY unset.
Any other compatible endpoint (Anthropic/OpenAI directly, vLLM, Ollama) works:
point LLM_BASE_URL at it and add catalog entries.
The agent reads the Obelisk docs from the rendered site, not a GitHub mount.
The pack.describe activity inlines each URL in DOCS_URLS_JSON into the
system prompt as a pointer (never the index body, so the prompt stays slim);
the model fetches the indexes and detail pages itself through the GET-only
curl program, whose allowlist covers https://obeli.sk:
curl https://obeli.sk/docs/latest/js/js-workflows/OBELISK_VERSION pins the doc set to the target runtime's version
(https://obeli.sk/docs/v${OBELISK_VERSION}/llms.txt/); empty means latest.
Override the whole list with DOCS_URLS_JSON (a JSON array of URLs). No GitHub
credential is involved; the former /workspace/{docs,components} mounts and
their GITHUB_TOKEN secret are gone.
The core registers a broad command catalog ported from just-bash (file, path,
text, search, checksum, encoding, and inspection tools; no gzip/gunzip/
zcat). Run help to list commands, or which NAME.
ask-user is deliberately not a general model tool or shell program. It is a
UI-coordinated shell operation for answers needed before the current task can
continue:
obelisk call obelisk-agent:stub/stub.ask-user '["Which deployment?"]'The command blocks until the user answers in the UI, then returns that answer to the agent in the same turn. A normal Markdown response ends the turn and is still the right way to ask a non-blocking conversational question.
The deployment mount under /workspace/deployment is lazy: the tree lists
immediately from digests and byte sizes, and bounded file bodies are fetched
from the content-addressed store on first read. Files over 1 MiB stay
digest-only and read as a placeholder.
Interactive job control is not available: jobs, wait, fg, bg, signals,
and durable background execution with & are unsupported (a trailing & is
rejected rather than run in the foreground). Packs use Obelisk child executions
for durable external work instead.
Sessions on the same Obelisk instance can talk to each other. The chat
program (available to the agent as a shell command and to you with the $
prefix) discovers, inspects, messages, and creates peer sessions:
chat models # LLM catalog: id, label, api type
chat list # sessions, newest first
chat read E_... # normalized transcript (--tail N, --json)
chat state E_... # one JSON line: state, working, offer id
chat send E_... please re-check X # queue a user prompt for a peer
chat create --model claude "prompt" # new session; prints its execution id
chat create '$ ls -la' # new session opened straight in bash
chat create --name research "..." # slug-labeled child (visible in its id)
chat current # this session's identity (JSON)
chat rename my-session # slug-name this session ([a-z0-9-])Notes:
chattalks to the agent's own instance (OBELISK_API_URL), unlike the obelisk-control tools which targetTARGET_OBELISK_API_URL.sendlooks up the peer's open input offer on every call (offers rotate each turn); while the peer is thinking the message queues for its next turn, when idle it is delivered immediately. Never invent or reuse offer ids.current,rename, andcreateare answered by the session workflow itself, not the HTTP activity: only a session knows its own identity, and by defaultcreateschedules the new session durably as a child of the caller (it is cancelled together with its parent). Usecreate --top-levelfor an independent session.- A rename is recorded on its own
session-namejoin set (via the dedicatedstub.session-renamedstub), so readers fetch a session's current slug with one small request; the UI sidebar shows the name, falling back to the execution id. - A prompt starting with
$runs directly in that session's shell instead of reaching the model, inchat createand the composer alike (the space after$is optional). create --name SLUGlabels the child at birth (published as its session name) and names its join set, so the child's execution id carries the slug (E_<parent>.n:<slug>_1). Child sessions start with empty history: pass all context they need in the prompt; a child can read its creator withchat read <parent>(its own system prompt names the parent).chat statereportslast_reply({"turn": N}) when a finished assistant message exists; sessions stay pending on an input offer even after a final answer, so this distinguishes "answered" from "still open". Read exactly that message withchat read ID --turn N.- In the web UI, child sessions are listed indented below their parent.
A dependency-free sample server exposes tools, a prompt, and two resources.
Start it with just sample-mcp-server; the sample's transport block in
deployment.toml, its outbound-host grant in server.toml, and its
MCP_SERVERS_JSON entry are already shipped and enabled, so just build and run
as above. In a new empty session:
mcp list
obelisk-local --help
obelisk-local tools add --a 2 --b 3
obelisk-local prompt greeting --name world
find /workspace/mcp/obelisk-local -type fResources are listed when the session mounts the server and fetched through
resources/read only on first read of their VFS path. See docs/mcp.md.
Apache-2.0. The vendored just-bash port keeps its own LICENSE/NOTICE under
vendor/just-bash-rs.
