Skip to content

docs(BENCH-QWEN38-TENSORFOLD-GAP): link measurement tools - #3361

Open
localai-org-maint-bot wants to merge 2 commits into
mudler:mainfrom
localai-org-maint-bot:docs/public-docs-audit-20260930
Open

localai-org-maint-bot wants to merge 2 commits into
mudler:mainfrom
localai-org-maint-bot:docs/public-docs-audit-20260930

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

The public Qwen3.8 TensorFold result did not link to its shipped measurement workflow. Add a README news entry and links to the endpoint harness guide, environment templates, and capture runner.

Explain the evidence limits: current comparisons provide neither token parity nor a cross-engine ratio, --draft off does not reconfigure a running vllm.cpp server, and a successful process exit does not prove a valid measurement. The comparison remains blocked by missing artifacts. No code or benchmark values change.

Tracks local issue ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ. The spec precedes the documentation commit. This is a documentation follow-up, not closure of the benchmark campaign.

Validation on CPU:

  • python3 -m pytest -q tests/scripts/test_qwen38_endpoint_bench.py tests/scripts/test_run_qwen38_tensorfold_gap.py tests/scripts/test_validate_qwen38_tensorfold_evidence.py: 58 passed.
  • python3 scripts/check-readme-structure.py and python3 scripts/check-agent-record.py: passed.
  • python3 tools/bench/validate_qwen38_tensorfold_evidence.py .agents/evidence/bench-qwen38-tensorfold-gap/latest: passed, retaining BLOCKED_MISSING_ARTIFACTS.
  • Commit style, trailer, relative-link, and whitespace checks passed.

Independent review of 700d744982537f5c5205a9a6581c10eed1c1668b passed with no findings. The reviewer reran all 58 tests and proved the link check detects a broken new target, then restored the file byte-for-byte.

The repository-wide preflight is not green. The completed checker sweeps report the same nine failures; the broader mutation sweep remains incomplete. The baseline and changed tree report unrelated release-state, benchmark-index, environment-documentation, oracle-pin, and stale-record failures. The local tool environment also lacks dependencies needed by some broad checks. These are outside this documentation-only change; no checker is weakened. No GPU, model execution, or new performance measurement was used.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]

The public result lacks links to the shipped measurement workflow. Scope a CPU-verifiable documentation follow-up for ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ without changing benchmark claims.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]
Readers could not reach the shipped endpoint harness from the public result.
Link its instructions and templates, and explain the evidence limits.
The benchmark stays blocked and publishes no new measurement.

The three focused CPU suites pass all 58 tests. README structure,
retained evidence, relative links, and whitespace checks pass.
Full preflight has existing unrelated failures; the operator verifies
its baseline comparison before handoff. No executable behavior changes.

Issue: ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:gpt-6 [exec_command]

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant