Skip to content

Add StructEval benchmark - #501

Open
reacher-z wants to merge 1 commit into
InftyAI:mainfrom
reacher-z:add-structeval-benchmark
Open

Add StructEval benchmark#501
reacher-z wants to merge 1 commit into
InftyAI:mainfrom
reacher-z:add-structeval-benchmark

Conversation

@reacher-z

Copy link
Copy Markdown

Summary

Adds StructEval to Training > Benchmark in both the README and the website data, following the repository's manual-addition path and alphabetical ordering. The website entry uses the list's existing default logo.

StructEval is a TMLR benchmark for evaluating LLM structured-output generation and conversion across 2,035 examples and 18 text and renderable visual formats.

Fit

The Training benchmark category covers standardized methods for evaluating LLM quality and capabilities. StructEval directly evaluates a cross-model capability that is not represented by the current entries.

Validation

  • Confirmed StructEval was not already present in the repository, pull requests, issues, comments, or discussions.
  • Parsed website/data.yml successfully with Ruby YAML.
  • Ran git diff --check.

Disclosure

I maintain StructEval and am submitting this project directly.

This pull request was prepared with assistance from OpenAI Codex. I reviewed and verified the final change.

@InftyAI-Agent InftyAI-Agent added needs-triage Indicates an issue or PR lacks a label and requires one. needs-priority Indicates a PR lacks a label and requires one. do-not-merge/needs-kind Indicates a PR lacks a label and requires one. labels Jul 26, 2026
@InftyAI-Agent
InftyAI-Agent requested review from cr7258 and kerthcet July 26, 2026 21:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

do-not-merge/needs-kind Indicates a PR lacks a label and requires one. needs-priority Indicates a PR lacks a label and requires one. needs-triage Indicates an issue or PR lacks a label and requires one.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants