chore: add top-final benchmark run results - #68
algomaster99 wants to merge 1 commit into
Conversation
Cargo
GitHubActions
Go
Maven
npm
|
294b26f to
ebcd657
Compare
c8aafaf to
eaaa01f
Compare
ebcd657 to
009ccdf
Compare
Full top-package benchmark run (60 cases x hook/nohook x 3 reps) captured under benchmark/top-final/.
009ccdf to
e7513e9
Compare
Adds the full top-package benchmark run (60 cases × hook/nohook × 3 reps) under
benchmark/top-final/, and includes the output ofpython3 benchmark/analyze_top.pyagainst these results:Satisfied's denominator excludes 5 reps per condition where the model never declared any dependency at all (it implemented the target functionality itself - there's nothing for yul to have acted on). The newAltcolumn is the subset ofSatisfiedreached via a different-but-equivalent package or mechanism than the one the case targeted (e.g.junit-jupiterinstead ofjunit,tzdatainstead ofpytz, GitHub Actions'cache: npminstead ofactions/cache) - a legitimate solution, not a hook/tool miss. Both classifications came out of manual review (see PR discussion below) and are hardcoded inanalyze_top.pyasEXCLUDED_REPS/ALTERNATIVE_REPS, keyed per(case, condition, rep)since hook and nohook are independent runs.🤖 Generated with Claude Code