Repository navigation
Conversation
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
ruff EXE001: the file has a shebang but shipped as mode 100644. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7W9hWwuZU8Mkwo3ZCrR7m Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
- Layout lists the real paths (src/assetops_harbor, template/, overlays/) and marks datasets/ as generated. - Run steps work from the repo root: correct Dockerfile path, adapter defaults instead of a --template pointing at a generated task, no PYTHONPATH=agent, no 'push'. - Drop the pin to main@81265cb; the branch targets aafeedback_changes. - Replace 'Still unverified' with the Docker-host results reported in #558. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
generate_tasks.py gains three options for corpora that do not ship in the repo, such as AssetOpsBenchScenarioGeneration/scenarios_data: - --runtime-image: the image each task builds FROM. - --data-dir: where the corpus lives inside that image. The per-task layer copies the scenario there and the healthcheck runs init_data.py with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy, as scenario_suite_runner does. - --skip-missing: warn and skip profile scenarios absent from the corpus (all.yaml lists wosr-62, which the corpus lacks). corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into one layer over the runtime image, so tasks share it instead of each carrying it in their build context. Also stop copying scenario 1's description and the wosr keyword into every generated task. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.
- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.
Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, and the generated dataset assetopsbench-corpus -> assetopsbench-suite, with matching wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/: assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
A dropped VPN made the LiteLLM proxy unreachable mid-session. Every trial still built its containers, loaded its data, retried the model for four minutes and exited 1, and the job recorded each as done, so a resume would not have re-run them. - Probe LITELLM_BASE_URL / TOKENROUTER_BASE_URL (per model prefix) before starting or resuming a job, and skip the model with a clear message when the router does not answer. - Resume with --filter-error-type NonZeroAgentExitCodeError, so trials whose agent crashed run again while scored trials are kept. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Add the generate_tasks.py flags a non-open profile needs, and record that those tasks load empty CouchDB collections: manifests reference shared/ files that exist only in the full corpus, not the runtime image, and the loader only warns when a data file is missing. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
StirrupAgent now loads the nearest .env above the working directory before checking router credentials, so tokenrouter/ and litellm_proxy/ models run without exporting their variables. override=False keeps exported variables and --ae ahead of the file, and only CREDENTIAL_ENV_VARS reach the container. Tests chdir into tmp_path so a developer's .env cannot leak in, and the credential fixture now restores variables that load_dotenv writes. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The earlier note said the three FMEA mini scenarios load cleanly. They do not: their manifest lists files as an array, which the check that produced the note skipped. All 35 mini scenarios reference corpus-only files. Also record that the referenced files are 1.5 MB (mini) to 58 MB (all), not the 2.1 GB of the corpus shared/ directory, and how the failure shows up. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
feat(harbor): load .env for StirrupAgent credentials
Resolve the README conflict by keeping both new sections. The base's hand-generation section now uses this branch's --runtime-image, --data-dir and --skip-missing flags, and its known-gap note becomes what happens without the suite image. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Add overlays/private-data.yaml, which bind-mounts the directory in AOB_PRIVATE_DIR read-only into main at /opt/suite/scenarios_data, the path the suite image uses. Tasks generated with --data-dir then load their manifests' shared/ files from the mount, with no suite image rebuild when the data changes. The variable is required with :? so an unset value fails naming it rather than as an invalid volume spec. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
refactor: drop FMSR generation, withhold LLM keys from MCP servers
Collaborator
|
@ShuxinLin merge already happen on main. |
Collaborator
Author
some edit go to main and rebase to aa branch. |
Restores the ten FMEA IDs dropped in 3f788d5. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
DhavalRepo18
force-pushed
the
aafeedback_changes
branch
from
October 2, 2026 22:54
04dffef to
7c4b27d
Compare
It handles the boundary case where the category field is empty, and the agent keeps trying with different options. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
DhavalRepo18
force-pushed
the
aafeedback_changes
branch
from
October 5, 2026 13:39
b4d8e41 to
7c4b27d
Compare
Standarize the System Promot and Coed Exec tool for AA Study
This PR is based on v1 run
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Updated car list with additional values.
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Add tsfm and car
Bring TSFM and IoT server updates
run.sh gains -k (repeat each task n times, the basis for pass@k and pass^k), -t (sampling temperature, default 0.6) and -u (agent turn budget, default 100), with validation on each. passk.py computes pass@k and pass^k from a leaderboard directory using the existing "passed" reward key. Taken from #618 without the to_reward.py change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PcbP743uYAdkJiV9UbiTjy Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Add pass@k/pass^k, temperature and turn-budget controls to harbor runs
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
A CAR answer now passes when it is a single-key object whose key matches the gold mode (response / clarification / abstain). The value under the key no longer affects pass or score; required-term coverage stays in the mode_* fields as a diagnostic. Replaces the CAR tests that targeted the removed groundtruth_eval.json term scoring and updates docs/static-json-evaluation.md to match. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
ShuxinLin
force-pushed
the
aafeedback_changes
branch
from
October 8, 2026 13:21
25ed13f to
8c59cd9
Compare
fix(evaluation): score CAR scenarios on the mode key only
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
placeholder for now