Universal Telegram bot with AI assistant, supporting multiple models, multimodality, and advanced Agent Mode with code execution, plus a Web Interface for browser-based chat with the agent.
About: Tech Stack (Rust 1.94), Integrations, and Architecture
This project is a Telegram bot that integrates with various Large Language Model (LLM) APIs to provide users with a multifunctional AI assistant. The bot can process text, voice, video messages, and images, work with documents, manage dialogue history, and perform complex tasks in an isolated sandbox.
The bot is developed using Rust 1.94, the teloxide library, and integrates with 6 Agent Mode LLM providers: ChatGPT/Codex (OAuth), OpenCode Go, OpenCode Zen, Zhipu AI/ZAI, MiniMax, and OpenRouter.
- Modular Workspace: Separation into domain logic (core), orchestration (runtime), and transport layers
- Transport-Agnostic Runtime: Progress rendering and execution model can be adapted for Discord, Slack, etc.
- Web Interface: Browser-based chat with the agent (Leptos SPA, SSE streaming, dark theme, markdown rendering)
- Topic-Scoped Infrastructure: Per-topic agent profiles, hooks, tools, and memory isolation
- Manager Control Plane: Programmatic topic management with RBAC, audit trail, and rollback support
- Sandbox Backends: Docker broker isolation by default, with optional direct Docker access
- Prompt Cache Optimization: Static prefix + dynamic suffix assembly with validated 80%+ cache hit rate on OpenCode Go
-
Workspace Architecture: Modular crate design with clear separation of concerns:
oxide-agent-core- Domain logic, LLM integrations, hooks, compaction, storageoxide-agent-runtime- Session orchestration, execution cycle, tool providers, sandboxoxide-agent-transport-telegram- Telegram transport layer (teloxide integration)oxide-agent-transport-web- Web interface backend (axum HTTP API, SSE, auth) and E2E test transportoxide-agent-web-contracts- Shared web API types: auth, config, events, sessions, tasksoxide-agent-web-ui- Web interface frontend (Leptos SPA): chat UI, SSE streaming, markdown rendering, dark themeoxide-agent-sandboxd- Sandbox broker daemon for Docker access isolation in the default Compose deploymentoxide-agent-telegram-bot- Binary entry point and configuration
-
Agent Mode:
-
Integrated Sandbox: Safe execution of Python code and shell commands in isolated sandbox instances. Docker/broker is the default deployment path.
-
Parallel Tool Execution: Multiple tool calls in one LLM response execute concurrently for faster task completion.
-
Fire-and-Forget Checkpoint: Memory persistence is async, non-blocking for reduced latency.
-
History Repair: Validates tool_call_id before LLM calls; orphaned tool results prevented during compaction.
-
Tools: Read/write files, execute commands, web search, work with video and file hosting.
-
Task Management (Todos):
write_todossystem for planning and tracking progress of complex requests. -
Durable Context: Topic
AGENTS.md, runtime injections, and enabled tools provide deterministic prompt context. -
File Handling: Accept files from user (up to 20MB), send to Telegram (up to 50MB), or upload to cloud (up to 4GB) with link generation.
-
Video Processing:
yt-dlpintegration for downloading video and media files from the internet. -
File Hosting: Upload files from sandbox to public hosting with short retention time.
-
Web Search and Data Extraction: Tavily, Brave Search, CRW, and local
web_markdown/web_crawlerhandle discovery and URL-to-Markdown extraction. -
Hooks System: Extensible architecture for intercepting and customizing agent behavior:
- Completion Check Hook - validates task completion
- Tool Access Policy - enforces profile-level tool allowlists and blocklists
- Hot Context Health - monitors context health during execution
- Search Budget Hook - prevents infinite loops in tool calls
- Soft Timeout Report Hook - provides detailed timeout reporting
- Sub-Agent Safety - ensures safe execution environments
- Registry - centralized hook management
-
Universal Runtime: Transport-agnostic progress rendering system that can be adapted for Discord, Slack, and other transports.
-
Hierarchical Delegation: The Main Agent spawns async Sub-Agents for parallel, independent subtasks. Each sub-agent runs in an isolated ephemeral session with a task-specific tool whitelist, inherits the topic AGENTS.md and parent cancellation, and returns results via background job tracking.
-
Autonomy: Agent plans steps and selects tools itself.
-
Telegram Authorization: Access control via
TELEGRAM_ALLOWED_USERS. -
Long-term Memory and Context: Up to 200K tokens with automatic compression when limit reached.
-
Execution Progress: Interactive display of current working step in Telegram.
-
-
Multi-LLM Support: 6 Agent Mode providers: ChatGPT/Codex (OAuth), OpenCode Go, OpenCode Zen, Zhipu AI/ZAI, MiniMax, and OpenRouter.
-
Native Tool Calling: Efficient use of tools in modern models with ToolCallCorrelation architecture.
-
Web Interface: Browser-based chat with the agent -- Leptos SPA with SSE streaming for real-time responses, dark theme, and markdown rendering.
-
Multimedia Processing:
- Voice and video messages (speech recognition via OpenRouter-hosted Gemini-family models).
- Images (analysis and description via multimodal models).
- Work with documents of various formats.
-
Voice Synthesis: Kokoro TTS for English voice replies and Silero TTS for Russian voice replies.
-
Context Management: Dialogue history saved in SQLx/Postgres with context-scoped isolation per topic.
-
Prompt Cache Optimization: Static prefix + dynamic suffix assembly order maximizes cache hit rate, with validated 80%+ cache hit on OpenCode Go.
The Web Interface is a Leptos SPA with a dark theme, SSE streaming, and markdown rendering.
New session screen with the session list, profile selector, and message input.
Live agent progress with the Thinking indicator and Stop control.
Formatted agent response with tables and the Activity panel showing sub-agent execution.
Settings page for account, default model, agent profiles, and password management.
Agent Mode task list, tool calls, and completed file delivery in Telegram.
Video download via yt-dlp and file sent back to the user in Telegram.
API Keys and Infrastructure
| Provider | Variable | Description |
|---|---|---|
| OpenCode Go | OPENCODE_GO_API_KEY |
Primary Agent Mode provider - recommended route: mimo-v2.5 via opencode-go. OpenCode |
| Telegram | TELEGRAM_TOKEN |
Bot token from @BotFather |
| PostgreSQL | OXIDE_DATABASE_URL |
SQLx durable storage for sessions, memory, web state, reminders, and audit |
| Zhipu AI (ZAI) | OPENAI_BASE_PROVIDERS__1__* |
Configure as OpenAI Base profile zai for GLM routes (glm-4.7, glm-4.5-air). Zhipu AI |
Permanent Life Mode bridge setup is solo-owner and explicit: LIFE_OWNER_WEB_LOGIN names the Web owner account, while LIFE_TELEGRAM_BOT_TOKEN + LIFE_TELEGRAM_CHAT_ID configure a dedicated Telegram Life bot/chat. This is separate from the ordinary Agent Mode TELEGRAM_TOKEN; the bridge milestone is ordinary chat synchronization across interfaces, not a memory/curator UX.
For Supabase Postgres or small deployments, keep the shared SQLx pool conservative (OXIDE_DATABASE_MAX_CONNECTIONS=5), run migrations as a deploy step, and keep the default Postgres task-file byte limit unless WAL/backups have been reviewed. The app image ships /app/migrations, and the Web service enables startup migrations by default so fresh single-instance databases cannot race Web startup reconciliation.
The bot supports these Agent Mode provider routes/profiles with tool calling:
- OpenCode Go (
OPENCODE_GO_API_KEY) - primary (recommended) provider for Agent Mode. Uses subscription OpenAI-compatible API atopencode.ai/zen/go. Recommended Agent Mode model:mimo-v2.5with provideropencode-go. Supports native tool calls (strict), reasoning content parsing, adaptive throttling, and unbounded retry. - OpenCode Zen - Free-tier filtered variant of OpenCode Go. Same provider code, filtered to free-only models via discovery. Provider alias:
opencode-zen. - ChatGPT/Codex (
CHATGPT_AUTH_PATH) - Headless OAuth provider for OpenAI Codex Responses API atchatgpt.com/backend-api/codex/responses. SSE streaming. No audio/image support. Usecargo run -p oxide-agent-telegram-bot --bin chatgpt-login -- loginfor initial auth. - Zhipu AI / ZAI (
OPENAI_BASE_PROVIDERS__1__PROFILE=zai) - Alternative OpenAI Base profile for Agent Mode (glm-4.7orglm-4.5-air). Provides native tool-aware chat completions and reasoning. - MiniMax (
MINIMAX_API_KEY) - Claude SDK-compatible provider via MiniMax API (MiniMax-M2.7). - OpenRouter (
OPENROUTER_API_KEY) - Multimodal/media routes and approved tool-capable Agent Mode routes, including Gemini-family model IDs through OpenRouter.
[!NOTE] Voice recognition requires
AUDIO_STT_BASE_URL; image/video analysis requires an explicitVISION_MODEL_ID/VISION_MODEL_PROVIDERroute.
- Docker - run the default code sandbox (
agent-sandbox:latest) - Sandbox Broker - Unix socket broker for Docker access isolation in Docker Compose
- Tavily API - optional web search provider (
TAVILY_API_KEY) - CRW - optional self-hosted web search and scrape backend (
OXIDE_CRW_BASE_URL,OXIDE_CRW_API_TOKEN) - Local Web Markdown - lightweight single-URL HTTP fetch with HTML-to-Markdown conversion and response/output limits
- Kokoro TTS Server - optional for English voice message synthesis (
KOKORO_TTS_URL) - Silero TTS Server - optional for Russian voice message synthesis (
SILERO_TTS_URL)
Deployment guide: docs/deploy.md.
Quick Docker start:
git clone https://github.com/0FL01/oxide-agent.git
cd oxide-agent
cp .env.example .env
$EDITOR .env
docker compose up --build -dExample Configuration File
# Telegram
TELEGRAM_TOKEN=YOUR_TOKEN
TELEGRAM_ALLOWED_USERS=ID1,ID2
TELEGRAM_MANAGER_ALLOWED_USERS=ID1
ATTACH_DETACH_ENABLED=true
REMINDER_AGENT_PROGRESS_ENABLED=false
REMINDER_SILENT_NO_CHANGE_ENABLED=true
# Agent Configuration
AGENT_TIMEOUT_SECS=300
DEBUG_MODE=false
# PostgreSQL durable storage
OXIDE_DATABASE_URL=postgres://oxide_agent:oxide_agent@localhost:5432/oxide_agent
OXIDE_DATABASE_MAX_CONNECTIONS=5
OXIDE_DATABASE_MIGRATE_ON_STARTUP=false
OXIDE_WEB_TASK_FILE_MAX_BYTES=33554432
# Permanent Life Mode solo bridge (optional/planned rollout)
# LIFE_OWNER_WEB_LOGIN=alice
# LIFE_TELEGRAM_BOT_TOKEN=YOUR_DEDICATED_LIFE_BOT_TOKEN
# LIFE_TELEGRAM_CHAT_ID=123456789
# API Keys
CHATGPT_AUTH_PATH=/app/config/chatgpt/auth.json
OPENROUTER_API_KEY=...
OPENCODE_GO_API_KEY=...
OPENCODE_GO_API_BASE=https://opencode.ai/zen/go/v1/chat/completions
MINIMAX_API_KEY=...
# ZAI / GLM through OpenAI Base
OPENAI_BASE_PROVIDERS__1__NAME=zai
OPENAI_BASE_PROVIDERS__1__API_BASE=https://api.z.ai/api/coding/paas/v4
OPENAI_BASE_PROVIDERS__1__API_KEY=...
OPENAI_BASE_PROVIDERS__1__PROFILE=zai
# Web Search Providers (can be enabled together)
# TAVILY_API_KEY=...
# BRAVE_SEARCH_API_KEY=...
# Brave Search — primary indexed web discovery when BRAVE_SEARCH_API_KEY is configured.
# CRW (self-hosted web search + scrape fallback)
# OXIDE_CRW_BASE_URL=http://127.0.0.1:3000
# OXIDE_CRW_API_TOKEN=...
# `web_search` appears only when at least one indexed backend key/endpoint is configured.
# CRW also backs the rendered path of `web_crawler` when both CRW env vars are configured.
Root docker-compose.yml expects OXIDE_DATABASE_URL from .env or the shell. Keep OXIDE_DATABASE_MIGRATE_ON_STARTUP=true unless a separate migration job is guaranteed to finish before Web startup.
Set explicit agent and media routes through .env.
- Agent and Sub-agent (Recommended Models) The default OpenCode Go route uses mimo-v2.5 for both the Main Agent and Sub-Agent. This route supports strict native tool calling, reasoning content, adaptive throttling, and unlimited retry.
AGENT_MODEL_ID="mimo-v2.5"
AGENT_MODEL_PROVIDER="opencode-go"
AGENT_MODEL_CONTEXT_WINDOW_TOKENS=320000
SUB_AGENT_MODEL_ID="mimo-v2.5"
SUB_AGENT_MODEL_PROVIDER="opencode-go"Alternative (ZAI): If you prefer ZAI/GLM, configure it as an OpenAI Base zai profile and use glm-4.7 for the Main Agent and glm-4.5-air for the Sub-Agent:
OPENAI_BASE_PROVIDERS__1__NAME=zai
OPENAI_BASE_PROVIDERS__1__API_BASE=https://api.z.ai/api/coding/paas/v4
OPENAI_BASE_PROVIDERS__1__API_KEY=...
OPENAI_BASE_PROVIDERS__1__PROFILE=zai
AGENT_MODEL_ID="glm-4.7"
AGENT_MODEL_PROVIDER="openai-base:zai"
SUB_AGENT_MODEL_ID="glm-4.5-air"
SUB_AGENT_MODEL_PROVIDER="openai-base:zai"Alternative (ChatGPT/Codex): OAuth-based provider using the Codex Responses API:
AGENT_MODEL_ID="gpt-5.4"
AGENT_MODEL_PROVIDER="chatgpt"
AGENT_MODEL_MAX_OUTPUT_TOKENS=32000
AGENT_MODEL_CONTEXT_WINDOW_TOKENS=128000Use cargo run -p oxide-agent-telegram-bot --bin chatgpt-login -- login for initial OAuth setup.
Omitting the sub-agent block falls back to the agent model settings.
AUDIO_STT_BASE_URL="http://127.0.0.1:8002"
VISION_MODEL_ID="google/gemini-3.1-flash-lite-preview"
VISION_MODEL_PROVIDER="openrouter"Model Route Catalogs
Configure the models available for manual selection. The first route is the default; an execution never switches models automatically:
# Priority: OpenCode Go (MiMo V2.5) > ZAI/OpenAI Base (GLM-4.7)
AGENT_MODEL_ROUTES__0__ID="mimo-v2.5"
AGENT_MODEL_ROUTES__0__PROVIDER="opencode-go"
AGENT_MODEL_ROUTES__1__ID="glm-4.7"
AGENT_MODEL_ROUTES__1__PROVIDER="openai-base:zai"
Use AGENT_MODEL_ROUTES__N__* for the main-agent catalog and SUB_AGENT_MODEL_ROUTES__N__* for the sub-agent catalog.
| Provider | Description |
|---|---|
| OpenCode Go | Primary (recommended) agent provider, subscription OpenAI-compatible API, DeepSeek V4 Flash, strict tool calling, structured output, reasoning effort |
| OpenCode Zen | Free-tier variant of OpenCode Go, filtered to free-only models via discovery |
| ChatGPT/Codex | Headless OAuth provider for OpenAI Codex Responses API, SSE streaming, no audio/image |
| ZAI (Zhipu AI) | Alternative agent provider, native tool-aware chat, GLM-4.7 / GLM-4.5-Air |
| MiniMax | Claude SDK-compatible, high context (MiniMax-M2.7) |
| OpenRouter | Aggregator for various models, including Gemini-family model IDs |
Note: Gemini-family models are configured through OpenRouter routes, not a direct Google Gemini provider.
Tool Providers
- Web Search Provider (
tool-web-search) - unified indexed web discovery through configured CRW, Tavily, and Brave backends - CRW Scrape Client (
tool-crw) - self-hosted rendered scrape backend forweb_crawler - WebFetch Markdown Provider (
tool-webfetch-md) - single-URL HTTP fetch with HTML-to-Markdown conversion and context-bomb limits
- Sandbox Exec Provider (
tool-sandbox-exec) - command execution in sandbox - Sandbox File Ops Provider (
tool-sandbox-fileops) - file read/write/list/edit in sandbox - Sandbox Lifecycle Provider (
tool-sandbox-recreate) - sandbox recreate
- Kokoro TTS Provider (
tool-tts-kokoro) - English voice message synthesis - Silero TTS Provider (
tool-tts-silero) - Russian voice message synthesis with SSML support
Kokoro TTS (English):
Tool: text_to_speech_en
Server setup: see KOKORO-TTS-setup guide for manual server setup.
KOKORO_TTS_URL=http://127.0.0.1:8000 # Default
KOKORO_TTS_VOICE=af_heart # Default voice
KOKORO_TTS_FORMAT=ogg # Recommended for Telegram
KOKORO_TTS_TIMEOUT_SECS=60Available voices: af_bella, af_aoede, af_alloy, af_heart (default)
Formats: ogg (recommended), mp3, wav
Silero TTS (Russian):
Tool: text_to_speech_ru
Server setup: see Oxide-Agent-TTS for containerized Kokoro + Silero TTS servers with FastAPI.
SILERO_TTS_URL=http://127.0.0.1:8001 # Default
SILERO_TTS_SPEAKER=baya # aidar | baya (default) | kseniya | xenia
SILERO_TTS_FORMAT=ogg # Recommended for Telegram
SILERO_TTS_SAMPLE_RATE=48000 # 8000 | 24000 | 48000 (default, best quality)
SILERO_TTS_TIMEOUT_SECS=60Available speakers: aidar, baya (default), kseniya, xenia
Formats: ogg (recommended), wav
SSML support: set ssml: true for SSML markup with <speak>, <break>, <prosody> tags
- Audio STT Provider (
tool-audio-stt) - audio transcription - Vision Image Provider (
tool-vision-image) - image description - Vision Video Provider (
tool-vision-video) - video description - YT-DLP Provider (
tool-ytdlp) - video and audio download from various platforms
- Todos Provider (
tool-todos) - task list management for planning - Delegation Provider (
tool-delegation) - async sub-agent spawn, wait, and cancellation - Reminder Provider (
tool-reminder) - reminder scheduling with pause/resume/retry
- Compression Provider (
tool-compression) - message compression tools - Agents MD Provider (
tool-agents-md) - topic-scoped AGENTS.md editing
- File Hoster Provider (
tool-file-delivery) - public file upload to temporary hosting (up to 4GB)
- Manager Control Plane (
manager-control-plane) - topic CRUD, bindings, contexts, RBAC, audit trail - SSH MCP Provider (
integration-ssh-mcp) - SSH infrastructure with YOLO full-permission mode - Jira MCP Provider (
integration-mcp-jira) - Jira integration - Mattermost MCP Provider (
integration-mcp-mattermost) - Mattermost integration - Stack Logs Provider (
tool-stack-logs) - Docker Compose log access
Manager Control Plane
Programmatic topic management with RBAC, audit trail, and rollback support.
- Topic CRUD:
forum_topic_create,forum_topic_edit,forum_topic_close,forum_topic_reopen,forum_topic_delete,forum_topic_list - Dynamic Bindings:
topic_binding_set,topic_binding_get,topic_binding_delete,topic_binding_rollback - Context Management:
topic_context_upsert,topic_context_get,topic_context_delete,topic_context_rollback - AGENTS.md Editing:
topic_agents_md_get,topic_agents_md_update(top-level agents only) - Infra Config:
topic_infra_upsert,topic_infra_get,topic_infra_delete,topic_infra_probe - Agent Profiles:
agent_profile_upsert,agent_profile_get,agent_profile_delete,agent_profile_rollback - Tools Management:
topic_agent_tools_enable,topic_agent_tools_disable,topic_agent_tools_get - Hooks Management:
topic_agent_hooks_enable,topic_agent_hooks_disable,topic_agent_hooks_get - Sandbox Management:
topic_sandbox_list,topic_sandbox_destroy,topic_sandbox_recreate
TELEGRAM_MANAGER_ALLOWED_USERS=123456789,987654321 # Users with manager control-plane access
MANAGER_HOME_CHAT_ID=-1001234567890 # Restrict to specific chat (optional)
MANAGER_HOME_THREAD_ID=1 # Thread ID (optional)
MANAGER_HOME_AGENT_ID=control-plane # Agent ID for manager home (optional)Note: When MANAGER_HOME_CHAT_ID is set, manager control-plane tools are only available in the designated topic.
Security
SSH, Jira, and Mattermost tools are blocked by default in private/DM chats for security.
DM_ALLOWED_TOOLS=todos_write,todos_list,spawn_sub_agents,wait_sub_agents,cancel_sub_agents # Allowlist mode
DM_BLOCKED_TOOLS=sandbox_exec # Additional blocklistSSH tools run in YOLO mode for configured topic infra bindings. Restrict access with topic-scoped allowed_tool_modes, DM tool restrictions, and secret refs.
Internal Structure, Context, Hooks, Compaction
- Topic-scoped
AGENTS.md - SQLx/Postgres-backed runtime memory
- Runtime context injections
- Enabled tools and profile instructions
Three-level loop detection system (agent/loop_detection/):
- Content Detector - analyzes repeating agent messages
- Tool Detector - tracks identical tool calls
- LLM Detector - uses LLM to analyze loop patterns
Configuration: LOOP_DETECTION_ENABLED, LOOP_TOOL_CALL_THRESHOLD (5), LOOP_LLM_CHECK_AFTER_TURNS (30), LOOP_SCOUT_MODEL
Unified session-level compaction with a single path through CompactionController:
- Detect - Pre-sampling budget check, context-limit retry, manual compact, or model-route downshift.
- Summarize - Uses a normal configured LLM route as a provider-agnostic local summary backend (
LocalLlmSummary). - Replace Atomically - Builds one
[OXIDE_COMPACTED_SUMMARY_V1]handoff, preserves pinned state and safe recent tool context, validates tool-call integrity, and replaces hot memory in one step.
Static prefix + dynamic suffix assembly maximizes provider-side prompt cache hit rate, with validated 80%+ cache hit on OpenCode Go.
Architecture:
- Assembly order:
[fallback + profile + workflow_guidance + structured_output](stable) +[wiki_context]+[date_context](dynamic) - Tool schemas: Compact sorted tool-name list in prompt text (2673->98 bytes, 27x reduction); full JSON schemas via native
tools[]payload - Budget guard:
compresstool blocked at <85% context utilization to prevent premature compaction and cache reset - Cache telemetry:
TokenUsageincludescached_tokensandcache_creation_tokens, parsed for all 9 production providers
Validated on OpenCode Go (deepseek-v4-flash):
- Peak cache hit rate: 99.7% after warmup (14 iterations, no compaction)
- Overall hit rate: 89.5% across full task
- Estimated cost: 6.4x reduction vs pre-optimization baseline ($0.014 vs $0.090 for same task)
- Premature compaction prevention (budget guard): cache hit preserved vs 93%->3.3% drop without guard
Cache telemetry parsers are deployed for all providers; live validation confirmed on OpenCode Go. Other providers return cache tokens when their upstream routes support it.
Details: docs/tips/cache-hit.md
Extensible architecture for personalizing agent behavior:
- Completion Hook - task completion handling (protected, cannot be disabled)
- Tool Access Policy - blocks tools not allowed by current profile (protected, cannot be disabled)
- Hot Context Health - monitors context health during execution
- Sub-Agent Safety - ensures safe execution environments for delegated tasks
- Search Budget - limits search tool calls (10 per session)
- Timeout Report - provides detailed timeout reporting
Manageable Hooks: search_budget, timeout_report
The agent uses a modular provider system, each offering a specialized set of tools. See Tool Providers for the full list with configuration details.
Enhanced reminder scheduling with pause/resume/retry support.
Schedules:
Once- One-time reminderInterval- Recurring every N minutes/hoursCron- Complex schedules with timezone and weekday support
Tools:
reminder_schedule- Create reminders with simplified args (date,time,every_minutes,every_hours,timezone,weekdays)reminder_list- List all remindersreminder_cancel- Cancel reminderreminder_pause/reminder_resume- Pause/resume with optional delayreminder_retry- Retry failed reminder
Statuses: scheduled, paused, completed, cancelled, failed
- Per-transport contexts live in
UserConfig.contextsviaUserContextConfig - Context-scoped storage API:
save_agent_memory_for_context,load_agent_memory_for_context,clear_agent_memory_for_context - Chat history isolated via
scoped_chat_storage_idformat:"{context_key}/{chat_uuid}"
- Flows attach/detach UX for per-session state management
- Stored by prefix:
users/{user_id}/topics/{context_key}/flows/{flow_id}/ forum_topic_listavailable for topic discovery (blocked for sub-agents)
- Storage record:
TopicAgentsMdRecord - Orchestration via storage API and
prompt/composer.rs - Limited to 300 lines for
AGENTS.md, 40 lines fortopic_context - Self-editing tools:
agents_md_get,agents_md_update(top-level agents only)
See docs/deploy.md for Docker, external services, sandbox, and operations notes.
- Send
/startto the bot. - Regular Mode: Just write messages or send files/voice notes.
- Agent Mode: Click the "Agent Mode" button. Now the bot can execute code and use advanced tools.
Agent Command Examples and Control
Agent Command Examples:
- "Write a python script that downloads the google homepage and finds the word 'Search' there"
- "Download video from YouTube via link [URL] and convert it to MP4 via FFmpeg"
- "Create a CSV file with weather data and upload it to file.io"
- "Find information about latest AI news via web search"
Topic Management Examples:
- "Create a new topic for Bug #1234"
- "Set agent profile 'developer' for this topic"
- "Enable Jira tools for this topic"
- "Get the audit trail for topic operations"
Control: Use "Clear Context", "Change Model", "Exit Agent Mode" or topic-specific controls.
File Tree (expand)
crates/
├── oxide-agent-core/ # Domain logic, LLM integrations, hooks, compaction, storage
│ └── src/
│ ├── agent/ # Agent core and execution logic
│ │ ├── compaction/ # Compaction pipeline (12 modules)
│ │ ├── hooks/ # Execution hooks (7 hooks)
│ │ ├── loop_detection/ # Loop detection (content, tool, llm)
│ │ ├── providers/ # Tool providers
│ │ │ ├── ssh_mcp.rs # SSH infrastructure
│ │ │ ├── jira_mcp/ # Jira integration
│ │ │ ├── mattermost_mcp/ # Mattermost integration
│ │ │ ├── web_search.rs # Unified indexed search
│ │ │ ├── crw/ # CRW scrape REST client
│ │ │ ├── tts/ # Kokoro TTS
│ │ │ ├── silero_tts/ # Silero TTS
│ │ │ ├── manager_control_plane/ # Topic CRUD, RBAC
│ │ │ └── ...
│ │ ├── tool_runtime/ # Typed tool registration and execution
│ │ ├── recovery/ # History repair, tool drift pruning
│ │ ├── runner/ # Execution loop, parallel tools
│ ├── llm/ # LLM provider integrations
│ │ ├── providers/ # Providers (chatgpt, zai, minimax, openrouter, opencode_go)
│ │ └── tool_correlation.rs
│ ├── sandbox/ # Sandbox facade and backends
│ │ ├── broker.rs # Unix-socket sandbox broker protocol
│ │ ├── manager.rs # Sandbox manager facade
│ │ └── traits.rs # Sandbox backend traits
│ ├── storage/ # Storage facade, SQLx backend, control-plane records
│ ├── capabilities/ # Capability module manifests
│ └── config.rs
├── oxide-agent-runtime/ # Session orchestration, execution cycle, tool providers, sandbox
│ └── src/
│ └── agent/runtime/ # Progress runtime, transport-agnostic progress
├── oxide-agent-transport-telegram/ # Telegram transport layer
│ └── src/
│ ├── bot/agent_handlers/ # Agent lifecycle, controls, callbacks, reminders
│ ├── bot/views/agent.rs # Agent Mode UI
│ ├── context.rs # Context-scoped state
│ ├── topic_route.rs # Topic binding resolution
│ ├── thread.rs # Thread-aware session isolation
│ └── session_registry.rs
├── oxide-agent-transport-web/ # Web interface server + E2E test transport
│ └── src/
│ ├── server.rs # HTTP API (axum), SSE, auth
│ ├── auth_helpers.rs # Password auth, session management
│ ├── sse.rs # Server-sent events
│ ├── static_assets.rs # Frontend serving
│ ├── task_executor.rs # Web task execution
│ ├── converters.rs # API type conversion
│ └── providers.rs # Scripted LLM provider
├── oxide-agent-web-contracts/ # Shared web API types
│ └── src/
│ ├── auth.rs # Auth types
│ ├── config.rs # Config types
│ ├── events.rs # SSE event types
│ ├── sessions.rs # Session types
│ └── tasks.rs # Task types
├── oxide-agent-web-ui/ # Web interface frontend (Leptos SPA)
│ └── src/
│ ├── components/ # UI components
│ ├── routes/ # Page routes
│ ├── sse_client.rs # SSE streaming client
│ └── styles/ # Dark theme, CSS
├── oxide-agent-sandboxd/ # Sandbox broker daemon
│ └── src/main.rs
└── oxide-agent-telegram-bot/ # Binary entry point and configuration
└── src/
├── main.rs
└── bin/chatgpt-login.rs # ChatGPT OAuth login helper
tests/ # Integration and functional tests
├── e2e/ # E2E tests for web transport
│ ├── session_tests.rs
│ ├── sse_tests.rs
│ ├── compaction_regression_tests.rs
│ ├── delegation_tests.rs
│ ├── reminder_tests.rs
│ └── tool_latency_tests.rs
docs/ # Documentation
├── silero-tts-api.md # Silero TTS integration
├── stack-logs.md # Stack logs tool
├── context-window-tracking.md # Token budget management
├── tips/cache-hit.md # Prompt cache optimization
├── hooks/ # Hooks system documentation (9 files)
├── prd/ # Product requirements documents
└── goals/ # Development goal tracking
sandbox/ # Docker configuration for sandbox
docker/ # Docker profile overlays
config/ # Configuration files (optional YAML)
.github/workflows/ # CI/CD workflows
The supported production profile composes all atomic capability features. Build with --no-default-features --features profile-full.
| Profile | Description | Key Components |
|---|---|---|
profile-full |
Full production deployment | All features |
| Category | Features |
|---|---|
| LLM Providers | llm-chatgpt, llm-minimax, llm-openai-base, llm-opencode-go, llm-openrouter |
| Search Tools | tool-web-search, tool-crw, tool-webfetch-md |
| Sandbox | tool-sandbox-exec, tool-sandbox-fileops, tool-sandbox-recreate |
| Sandbox Backend | sandbox-backend-sandboxd-client |
| Media | tool-audio-stt, tool-vision-image, tool-vision-video, tool-ytdlp |
| TTS | tool-tts-kokoro, tool-tts-silero |
| Memory | tool-compression, tool-agents-md |
| Integrations | integration-mcp-jira, integration-mcp-mattermost, integration-ssh-mcp |
| Other | tool-todos, tool-delegation, tool-reminder, tool-browser-live, tool-file-delivery, tool-stack-logs, manager-control-plane |
The browser-sidecar service runs a Chromium-based headless browser
controlled by a native Rust binary (oxide-browser-sidecar) that talks CDP
directly. It is wired into root Compose but Browser Live remains disabled until
you enable it and set a token.
Enable for the Web UI:
# .env
BROWSER_AGENT_SIDECAR_TOKEN=<set-a-long-random-token>
BROWSER_AGENT_ENABLED=true
BROWSER_AGENT_SIDECAR_BASE_URL=http://127.0.0.1:8787
BROWSER_AGENT_SIDECAR_WS_URL=ws://127.0.0.1:8787# Start the unified stack with the sidecar
docker compose -f docker-compose.yml up --build -d
# Verify the sidecar is healthy
curl -fsS http://127.0.0.1:8787/healthz -H "Authorization: Bearer ${BROWSER_AGENT_SIDECAR_TOKEN}"Browser Live is disabled by default, uses ephemeral profiles, and requires a shared token between the app and sidecar. The agent has full browser access in Yolo mode: it can type secrets, submit forms, and interact with any reachable page. Full setup details and warnings are in docs/browser-live.md.
Main Rust Libraries
- teloxide (0.17) - Telegram Bot API with macros and handlers
- tokio (1.52) - asynchronous runtime
- sqlx-postgres/sqlx-core (0.8) - PostgreSQL durable storage
- bollard (0.20) - Docker API for sandbox daemon (
sandbox-daemonfeature only) - leptos (0.8) - Web interface frontend (CSR)
- axum (0.7) - Web interface HTTP API
- reqwest (0.12/0.13) - HTTP client with multipart and streaming support
- serde_json (1.0) - JSON serialization/deserialization
- tiktoken-rs (0.9) - token counting for various models
- rmcp (1.2) - MCP client for Jira/SSH/Mattermost
- moka (0.12) - high-performance cache with TTL
- chrono (0.4) - date and time handling
- thiserror (2.0) - custom error creation
- anyhow (1.0) - simplified error handling in application
The project is distributed under the MIT License. Details in the LICENSE file.
Copyright (C) 2026 @0FL01