Skip to content

Review Workshop eval trajectory changes - #477

Open
AshishKumar4 wants to merge 1 commit into
evals/github-pr-reportfrom
evals/trajectory-review
Open

AshishKumar4 wants to merge 1 commit into
evals/github-pr-reportfrom
evals/trajectory-review

Conversation

@AshishKumar4

@AshishKumar4 AshishKumar4 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Stacked on #476.

After the deterministic baseline/candidate comparison completes, download and pretty-print both raw trajectory artifacts, then run the existing Bonk action with a fixed 3–4 sentence prompt. The report covers direction, measured changes, the strongest trajectory-level explanation, and uncertainty from model variance.

Bonk starts from the trusted base checkout, reads a precomputed baseline-to-candidate diff, and receives a no-push token. OpenCode denies every tool except file reads, downloaded inputs are excluded from Git status, and the advisory step does not affect the deterministic eval result.

@github-actions github-actions Bot added the delivery Changes to CI or release delivery label Sep 10, 2026
devin-ai-integration[bot]

This comment was marked as resolved.

@github-actions

Copy link
Copy Markdown

Preview: pr477-evals-traject-fc07c928

https://pr477-evals-traject-fc07c928-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@ask-bonk

ask-bonk Bot commented Sep 10, 2026

Copy link
Copy Markdown

Author identity unknown

*** Please tell me who you are.

Run

git config --global user.email "you@example.com"
git config --global user.name "Your Name"

to set your account's default identity.
Omit --global to set the identity only in this repository.

fatal: empty ident name (for runner@runnervmlun5p.22vpi1papmiehlbrkehlsftqdc.bx.internal.cloudapp.net) not allowed

github run

@ask-bonk

ask-bonk Bot commented Sep 10, 2026

Copy link
Copy Markdown

@AshishKumar4 Bonk workflow failed. Check the logs for details.

View workflow run · To retry, trigger Bonk again.

Comment thread .github/workflows/workshop-evals.yml Outdated
@ask-bonk

ask-bonk Bot commented Sep 10, 2026

Copy link
Copy Markdown

Posted 1 actionable inline finding.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 567e19d to 251a515 Compare September 10, 2026 22:24
@ask-bonk

ask-bonk Bot commented Sep 10, 2026

Copy link
Copy Markdown

Performance is inconclusive: all nine trials failed in both runs, with null pass-rate deltas. Candidate mean durations increased by 0.51–0.84 seconds across tasks. Both runs stopped before any model turn or tool call due to connection errors. These timing shifts likely reflect failure timing, and model variance cannot be assessed.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 251a515 to 9269040 Compare September 10, 2026 23:52
@AshishKumar4
AshishKumar4 marked this pull request as draft September 10, 2026 23:52
@ask-bonk

ask-bonk Bot commented Sep 10, 2026

Copy link
Copy Markdown

LGTM!

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 9269040 to bf4ef34 Compare September 10, 2026 23:57
Comment thread .github/workflows/workshop-evals.yml Outdated
@ask-bonk

ask-bonk Bot commented Sep 11, 2026

Copy link
Copy Markdown

Posted 1 actionable inline finding.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from bf4ef34 to 2296df5 Compare September 11, 2026 00:13
@ask-bonk

ask-bonk Bot commented Sep 11, 2026

Copy link
Copy Markdown

LGTM!

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 2296df5 to 6f4c82c Compare September 11, 2026 00:22
@ask-bonk

ask-bonk Bot commented Sep 11, 2026

Copy link
Copy Markdown

LGTM!

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 6f4c82c to 273a43d Compare September 11, 2026 00:25
Comment thread .github/workflows/workshop-evals-pr.yml Outdated
@ask-bonk

ask-bonk Bot commented Sep 11, 2026

Copy link
Copy Markdown

Posted 1 actionable inline finding.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 273a43d to 8185820 Compare September 11, 2026 00:29
@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 5167d7a to 85d5fca Compare September 14, 2026 16:46
@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 85d5fca to 94418db Compare September 14, 2026 17:04
@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 94418db to fe102f8 Compare September 14, 2026 17:32
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

LGTM!

github run

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Performance stayed materially unchanged: expense-ledger remained 0/3 and project-doc 3/3, while appointment-desk’s 1/3 to 0/3 shift was non-comparable. Expense-ledger was 40.9s faster with fewer calls and slightly lower cost, but project-doc was 34.4s slower with slightly more calls and cost. Most notably, two candidate appointment trials made zero model turns, and the third stopped after two calls, unlike the baseline’s successful 18-call build. With only three trials plus connection and timeout failures, stochastic behavior could explain these shifts.

github run

@AshishKumar4
AshishKumar4 added this pull request to stack #497 September 14, 2026 19:48
@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from fe102f8 to a5a8edd Compare September 14, 2026 19:50
@AshishKumar4
AshishKumar4 marked this pull request as ready for review September 14, 2026 19:50
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

LGTM!

github run

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

LGTM!

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 225545d to 4d3c8a9 Compare September 14, 2026 20:03
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

[P1] The earlier trust-boundary finding remains unresolved at .github/workflows/workshop-evals-pr.yml:295. opencode github run checks out the PR head before chatting, so candidate OpenCode config/plugins can still execute with supplied credentials despite the base checkout.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 4d3c8a9 to 9987974 Compare September 14, 2026 20:32
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

[P1] .github/workflows/workshop-evals-pr.yml:286: jq . still emits long JSON strings on single lines. Real artifacts contain 2–15 KB tool arguments, while OpenCode truncates lines at 2 KB, hiding trajectory details. Chunk or wrap string values during formatting.

github run

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Performance stayed materially unchanged: expense-ledger remained 0/3 and project-doc 3/3, while appointment-desk’s 1/3 to 0/3 shift was non-comparable. Expense-ledger was 244.8s faster, with mean turns dropping 9.33 to 1.67, calls 15.67 to 2, and cost $0.488 to $0.102; project-doc slowed 6.4s with slight call and cost increases. Most notably, every candidate expense-ledger run timed out after only creating the gadget, whereas one baseline reached 18 turns and 33 calls with substantial implementation and verification. With three trials per cohort plus connection and timeout failures, stochastic behavior could explain these differences.

github run

@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 9987974 to 89f593f Compare September 14, 2026 23:35
@AshishKumar4
AshishKumar4 force-pushed the evals/trajectory-review branch from 89f593f to 45f006b Compare September 14, 2026 23:37
@ask-bonk

ask-bonk Bot commented Sep 15, 2026

Copy link
Copy Markdown

Performance regressed: project-doc fell from 3/3 to 2/3, while both other cohorts remained 0/3. Project-doc averaged 12.3s faster with fewer turns, calls, and cost; appointment-desk was 148.2s faster but gained four mean tool errors. The failed candidate project-doc trial stopped after one turn and zero tool calls, unlike every baseline trial, which created and populated the document. With only three trials and anomalous stalls and malformed tool calls, stochastic behavior could explain the result.

github run

@ask-bonk

ask-bonk Bot commented Sep 15, 2026

Copy link
Copy Markdown

No additional actionable findings; the existing jq . trajectory truncation finding remains unresolved.

github run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

delivery Changes to CI or release delivery

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant