Skip to content

Aafeedback changes (do not merge) - #604

Draft
ShuxinLin wants to merge 106 commits into
mainfrom
aafeedback_changes
Draft

ShuxinLin wants to merge 106 commits into
mainfrom
aafeedback_changes

Conversation

@ShuxinLin

Copy link
Copy Markdown
Collaborator

placeholder for now

DhavalRepo18 and others added 30 commits September 26, 2026 17:48
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
ruff EXE001: the file has a shebang but shipped as mode 100644.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q7W9hWwuZU8Mkwo3ZCrR7m
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
- Layout lists the real paths (src/assetops_harbor, template/, overlays/)
  and marks datasets/ as generated.
- Run steps work from the repo root: correct Dockerfile path, adapter
  defaults instead of a --template pointing at a generated task, no
  PYTHONPATH=agent, no 'push'.
- Drop the pin to main@81265cb; the branch targets aafeedback_changes.
- Replace 'Still unverified' with the Docker-host results reported in #558.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
generate_tasks.py gains three options for corpora that do not ship in the
repo, such as AssetOpsBenchScenarioGeneration/scenarios_data:

- --runtime-image: the image each task builds FROM.
- --data-dir: where the corpus lives inside that image. The per-task
  layer copies the scenario there and the healthcheck runs init_data.py
  with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy,
  as scenario_suite_runner does.
- --skip-missing: warn and skip profile scenarios absent from the corpus
  (all.yaml lists wosr-62, which the corpus lacks).

corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into
one layer over the runtime image, so tasks share it instead of each
carrying it in their build context.

Also stop copying scenario 1's description and the wosr keyword into
every generated task.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.

- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
  for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
  task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
  resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.

Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus ->
assetopsbench/runtime:suite, /opt/corpus/scenarios_data ->
/opt/suite/scenarios_data, and the generated dataset
assetopsbench-corpus -> assetopsbench-suite, with matching wording in
run.sh, the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/:
assetopsbench/runtime:corpus -> assetopsbench/runtime:suite,
/opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset
assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh,
the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
A dropped VPN made the LiteLLM proxy unreachable mid-session. Every trial
still built its containers, loaded its data, retried the model for four
minutes and exited 1, and the job recorded each as done, so a resume
would not have re-run them.

- Probe LITELLM_BASE_URL / TOKENROUTER_BASE_URL (per model prefix) before
  starting or resuming a job, and skip the model with a clear message
  when the router does not answer.
- Resume with --filter-error-type NonZeroAgentExitCodeError, so trials
  whose agent crashed run again while scored trials are kept.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Add the generate_tasks.py flags a non-open profile needs, and record that
those tasks load empty CouchDB collections: manifests reference shared/
files that exist only in the full corpus, not the runtime image, and the
loader only warns when a data file is missing.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
StirrupAgent now loads the nearest .env above the working directory before
checking router credentials, so tokenrouter/ and litellm_proxy/ models run
without exporting their variables. override=False keeps exported variables
and --ae ahead of the file, and only CREDENTIAL_ENV_VARS reach the container.

Tests chdir into tmp_path so a developer's .env cannot leak in, and the
credential fixture now restores variables that load_dotenv writes.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The earlier note said the three FMEA mini scenarios load cleanly. They do
not: their manifest lists files as an array, which the check that produced
the note skipped. All 35 mini scenarios reference corpus-only files. Also
record that the referenced files are 1.5 MB (mini) to 58 MB (all), not the
2.1 GB of the corpus shared/ directory, and how the failure shows up.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
feat(harbor): load .env for StirrupAgent credentials
Resolve the README conflict by keeping both new sections. The base's
hand-generation section now uses this branch's --runtime-image, --data-dir
and --skip-missing flags, and its known-gap note becomes what happens
without the suite image.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Add overlays/private-data.yaml, which bind-mounts the directory in
AOB_PRIVATE_DIR read-only into main at /opt/suite/scenarios_data, the path
the suite image uses. Tasks generated with --data-dir then load their
manifests' shared/ files from the mount, with no suite image rebuild when
the data changes. The variable is required with :? so an unset value fails
naming it rather than as an invalid volume spec.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
refactor: drop FMSR generation, withhold LLM keys from MCP servers
@DhavalRepo18

Copy link
Copy Markdown
Collaborator

@ShuxinLin merge already happen on main.

@ShuxinLin

Copy link
Copy Markdown
Collaborator Author

@ShuxinLin merge already happen on main.

some edit go to main and rebase to aa branch.

Restores the ten FMEA IDs dropped in 3f788d5.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
It handles the boundary case where the category field is empty, and the agent keeps trying with different options.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
ShuxinLin and others added 20 commits October 5, 2026 10:26
Standarize the System Promot and Coed Exec tool for AA Study
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Updated car list with additional values.
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
 into add-tsfm-and-car-

Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Bring TSFM and IoT server updates
run.sh gains -k (repeat each task n times, the basis for pass@k and
pass^k), -t (sampling temperature, default 0.6) and -u (agent turn
budget, default 100), with validation on each. passk.py computes
pass@k and pass^k from a leaderboard directory using the existing
"passed" reward key.

Taken from #618 without the to_reward.py change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PcbP743uYAdkJiV9UbiTjy
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Add pass@k/pass^k, temperature and turn-budget controls to harbor runs
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
A CAR answer now passes when it is a single-key object whose key matches
the gold mode (response / clarification / abstain). The value under the
key no longer affects pass or score; required-term coverage stays in the
mode_* fields as a diagnostic.

Replaces the CAR tests that targeted the removed groundtruth_eval.json
term scoring and updates docs/static-json-evaluation.md to match.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix(evaluation): score CAR scenarios on the mode key only

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants