VisualClaw is the agent system codebase. VisualClawArena is the separate benchmark/data release containing the 200-scenario evaluation suite, scenario specs, clips, gold workspaces, and paper-facing result artifacts.
This public VisualClaw branch keeps the reusable benchmark harness code, but does not track VisualClawArena data directories or intermediate experiment outputs. In particular, benchmark/data/, benchmark/data-spec/, results/, docs/visualclawarena/, and docs/case_studies/ are intentionally excluded here.
When VisualClawArena is released, point the benchmark runners at the external scenario/spec root instead of committing those assets into the VisualClaw system repository.
The runner supports the release layout directly:
VisualClawArena/
scenarios/<scenario_id>/spec/questions.json
scenarios/<scenario_id>/spec/scripts/
scenarios/<scenario_id>/data/workspace/
scenarios/<scenario_id>/data/clip/*.mp4
Install the convenience extra:
pip install -e ".[arena]"
cp .env.example .envThen use the smoke helper to validate a local backend and dataset path.
Claude Code, using a long-lived Claude Code token:
claude setup-token
# Put these in .env:
# CLAUDE_CODE_OAUTH_TOKEN=<token printed by setup-token>
# VISUALCLAW_ARENA_ROOT=/path/to/VisualClawArena
scripts/run_visualclawarena_agent_smoke.shCodex, using either a Codex access token or OpenAI API key:
# Put these in .env:
# CODEX_ACCESS_TOKEN=<codex access token>
# or: OPENAI_API_KEY=<OpenAI API key>
# VISUALCLAW_ARENA_ROOT=/path/to/VisualClawArena
BACKEND=codex scripts/run_visualclawarena_agent_smoke.shThe helpers load .env automatically and never commit it. They delegate persistent
auth, when needed, to claude or codex login. By default the smoke helper runs one round of
mmt_q1 with 8 keyframes and medium reasoning effort; override SCENARIO,
MAX_ROUNDS, MODEL, EFFORT, or OUT_ROOT for larger runs.
Use the public self-evolve helper when you want memory-guided skill evolution instead of a plain backend smoke test:
# Put VISUALCLAW_ARENA_ROOT in .env.
# Claude Code backend and Claude Code evolver.
# Put CLAUDE_CODE_OAUTH_TOKEN in .env.
SCENARIO=mmt_q1 MAX_ROUNDS=0 scripts/run_visualclawarena_self_evolve.sh
# Codex backend with OpenAI evolver.
# Put OPENAI_API_KEY in .env.
BACKEND=codex SCENARIO=mmt_q1 MAX_ROUNDS=0 scripts/run_visualclawarena_self_evolve.shThe helper seeds from visualclaw/skills_seed/seed_universal_mc, writes evolved
skills to a run-local bank, enables runner memory, and turns on schema/file
preflight. Useful overrides include SKILLS_EVOLVE_EVERY, MEMORY_TOP_K,
MEMORY_MAX_CHARS, EVOLVER_BACKEND, and EVOLVER_MODEL.
From a local full benchmark checkout:
python scripts/prepare_visualclawarena_hf.py \
--out /path/to/VisualClawArena \
--copy-mode hardlink \
--overwriteThen upload:
huggingface-cli upload <org-or-user>/VisualClawArena /path/to/VisualClawArena --repo-type dataset