Skip to content

Latest commit

 

History

History
92 lines (67 loc) · 3.11 KB

File metadata and controls

92 lines (67 loc) · 3.11 KB

VisualClawArena Release Boundary

VisualClaw is the agent system codebase. VisualClawArena is the separate benchmark/data release containing the 200-scenario evaluation suite, scenario specs, clips, gold workspaces, and paper-facing result artifacts.

This public VisualClaw branch keeps the reusable benchmark harness code, but does not track VisualClawArena data directories or intermediate experiment outputs. In particular, benchmark/data/, benchmark/data-spec/, results/, docs/visualclawarena/, and docs/case_studies/ are intentionally excluded here.

When VisualClawArena is released, point the benchmark runners at the external scenario/spec root instead of committing those assets into the VisualClaw system repository.

Running From The HF-Format Dataset

The runner supports the release layout directly:

VisualClawArena/
  scenarios/<scenario_id>/spec/questions.json
  scenarios/<scenario_id>/spec/scripts/
  scenarios/<scenario_id>/data/workspace/
  scenarios/<scenario_id>/data/clip/*.mp4

Install the convenience extra:

pip install -e ".[arena]"
cp .env.example .env

Then use the smoke helper to validate a local backend and dataset path.

Claude Code, using a long-lived Claude Code token:

claude setup-token
# Put these in .env:
# CLAUDE_CODE_OAUTH_TOKEN=<token printed by setup-token>
# VISUALCLAW_ARENA_ROOT=/path/to/VisualClawArena
scripts/run_visualclawarena_agent_smoke.sh

Codex, using either a Codex access token or OpenAI API key:

# Put these in .env:
# CODEX_ACCESS_TOKEN=<codex access token>
# or: OPENAI_API_KEY=<OpenAI API key>
# VISUALCLAW_ARENA_ROOT=/path/to/VisualClawArena
BACKEND=codex scripts/run_visualclawarena_agent_smoke.sh

The helpers load .env automatically and never commit it. They delegate persistent auth, when needed, to claude or codex login. By default the smoke helper runs one round of mmt_q1 with 8 keyframes and medium reasoning effort; override SCENARIO, MAX_ROUNDS, MODEL, EFFORT, or OUT_ROOT for larger runs.

Self-Evolution Runs

Use the public self-evolve helper when you want memory-guided skill evolution instead of a plain backend smoke test:

# Put VISUALCLAW_ARENA_ROOT in .env.

# Claude Code backend and Claude Code evolver.
# Put CLAUDE_CODE_OAUTH_TOKEN in .env.
SCENARIO=mmt_q1 MAX_ROUNDS=0 scripts/run_visualclawarena_self_evolve.sh

# Codex backend with OpenAI evolver.
# Put OPENAI_API_KEY in .env.
BACKEND=codex SCENARIO=mmt_q1 MAX_ROUNDS=0 scripts/run_visualclawarena_self_evolve.sh

The helper seeds from visualclaw/skills_seed/seed_universal_mc, writes evolved skills to a run-local bank, enables runner memory, and turns on schema/file preflight. Useful overrides include SKILLS_EVOLVE_EVERY, MEMORY_TOP_K, MEMORY_MAX_CHARS, EVOLVER_BACKEND, and EVOLVER_MODEL.

Preparing The HF Dataset Folder

From a local full benchmark checkout:

python scripts/prepare_visualclawarena_hf.py \
  --out /path/to/VisualClawArena \
  --copy-mode hardlink \
  --overwrite

Then upload:

huggingface-cli upload <org-or-user>/VisualClawArena /path/to/VisualClawArena --repo-type dataset