diff --git a/.agents/issues/BENCH-QWEN38-TENSORFOLD-GAP/ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ.md b/.agents/issues/BENCH-QWEN38-TENSORFOLD-GAP/ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ.md new file mode 100644 index 000000000..29818f3fe --- /dev/null +++ b/.agents/issues/BENCH-QWEN38-TENSORFOLD-GAP/ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ.md @@ -0,0 +1,23 @@ +ID: ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ +Title: Link the Qwen3.8 measurement workflow from public docs +Row: BENCH-QWEN38-TENSORFOLD-GAP +State: OPEN +Kind: documentation +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-30 +Updated: 2026-09-30 +Closed: - + +## Problem + +The public TensorFold benchmark detail records the blocked result but does not link the shipped endpoint harness instructions or runner environment templates. Readers cannot find the reproduction prerequisites from the result page. + +## Resolution + +30 September 2026: Implemented README news and public benchmark links to the +existing harness guide and environment templates. The three focused CPU suites +passed all 58 tests. README structure, retained evidence, relative links, and +whitespace checks passed. Independent review and upstream landing remain +pending, so this issue stays open. diff --git a/.agents/specs/qwen38-measurement-docs.md b/.agents/specs/qwen38-measurement-docs.md new file mode 100644 index 000000000..e47305cc9 --- /dev/null +++ b/.agents/specs/qwen38-measurement-docs.md @@ -0,0 +1,74 @@ +# Make the Qwen3.8 measurement workflow discoverable + +## Scope + +Documentation follow-up for `BENCH-QWEN38-TENSORFOLD-GAP`, tracked by +`ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ`. The benchmark lifecycle remains +`BLOCKED_MISSING_ARTIFACTS`. This change adds no implementation or measurement. + +| Surface | Source | Change | Verification | +|---|---|---|---| +| README news | `tools/bench/qwen38_endpoint_bench.py:466` | Announce the endpoint measurement tool with the blocked comparison caveat | Read parser and public result | +| Public benchmark detail | `benchmarks/manifests/qwen38_tensorfold/README.md:1` | Link the harness instructions and environment templates; explain comparison limits | Resolve links and inspect verdict logic | +| Capture prerequisites | `tools/bench/run_qwen38_tensorfold_gap.sh` | Link existing runner instructions where available, without inventing a runnable GPU recipe | Inspect parser and templates | + +## Design and sources + +Add one short news item to `README.md` and a measurement-tools section to +`docs/benchmarks/qwen38-tensorfold-gap.md`. Reuse the committed manifest guide +instead of duplicating its command. Explain that prompt fingerprints and reply +text evidence are distinct. `PROFILE_COMPARISON` is not token parity and +supplies no cross-engine ratio. Preserve the bounded discovery scope, absent +MTP weights, existing evidence, and benchmark values. + +The parser and comparison contract live in +`tools/bench/qwen38_endpoint_bench.py`. The original benchmark spec is +`bench-qwen38-tensorfold-gap.md`. This documentation change does not port vLLM +behavior. Upstream source execution and performance comparisons are not +applicable because no executable code changes. + +## Tests and gates + +No tests to port. Use existing CPU-only suites: + +```sh +python3 -m pytest -q tests/scripts/test_qwen38_endpoint_bench.py \ + tests/scripts/test_run_qwen38_tensorfold_gap.py \ + tests/scripts/test_validate_qwen38_tensorfold_evidence.py +python3 scripts/check-readme-structure.py +python3 tools/bench/validate_qwen38_tensorfold_evidence.py .agents/evidence/bench-qwen38-tensorfold-gap/latest +``` + +Check each added relative link and run `git diff --check`. Run repository +preflight and distinguish pre-existing failures from changes introduced here. +A fresh reviewer checks an immutable commit against source. No runtime +behavior or test guarantee changes, so code mutation testing is not applicable. + +## Risks and stop conditions + +Do not imply a successful endpoint capture, a performance improvement, model +parity, or usable MTP weights. Do not change historical EXL3 measurements or +benchmark-index work already covered by another pull request. Use no GPU, +external compute, model downloads, or service management. Open one fork pull +request; do not merge upstream. + +## Now + +Documentation implemented and CPU checks passed. Independent source review and +upstream landing remain pending. The benchmark remains `BLOCKED_MISSING_ARTIFACTS`. + +## Outcome + +README news links to the public result, which now links the existing harness +guide and both environment templates. The text explains comparison limits, +draft configuration, and refused measurements without adding benchmark values. + +On 30 September 2026, the three focused pytest suites passed all 58 tests. +The README structure check, evidence validator, relative-link checks, and +`git diff --check` passed. The evidence verdict remains +`BLOCKED_MISSING_ARTIFACTS`. Repository preflight encountered existing record, +index, environment-documentation, and release-state failures outside this scope. +The operator compares its full result with the baseline before handoff. + +No executable behavior or test guarantee changed. Red-first code tests, +mutation tests, GPU execution, and new performance measurements do not apply. diff --git a/README.md b/README.md index 741faa9aa..776050777 100644 --- a/README.md +++ b/README.md @@ -37,6 +37,9 @@ ## News +- **2026-09** **Qwen3.8 gains an endpoint measurement tool.** The TensorFold comparison remains + blocked by missing artifacts, with no speed result. See the + [measurement instructions and limits](docs/benchmarks/qwen38-tensorfold-gap.md#measurement-tools). - **2026-09** **Multimodal chat reaches the model through HTTP.** Qwen3-VL accepts images; dots3-note also accepts audio and multiple media items. CPU tests use synthetic weights; real-checkpoint token parity remains unverified. See the [input guide](docs/guides/multimodal-input.md) diff --git a/docs/benchmarks/qwen38-tensorfold-gap.md b/docs/benchmarks/qwen38-tensorfold-gap.md index 22fc84253..0db435278 100644 --- a/docs/benchmarks/qwen38-tensorfold-gap.md +++ b/docs/benchmarks/qwen38-tensorfold-gap.md @@ -60,6 +60,31 @@ The campaign can resume measurement only after the runner prerequisites are available in the documented/configured staging scope. It must then rerun correctness before timing rather than reusing this blocker as benchmark evidence. +## Measurement tools + +The [endpoint harness guide](../../benchmarks/manifests/qwen38_tensorfold/README.md) +includes a request example, required tokenizer identity, and the evidence contract. +The harness measures an already-running OpenAI-compatible endpoint. For the +`vllm-cpp` adapter, `--draft off` records the declared policy but does not +reconfigure speculative decoding on the running server. + +Current TensorFold/vllm.cpp comparisons report `PROFILE_COMPARISON`: they prove +neither token parity nor a cross-engine ratio. Matched prompt-token fingerprints +are separate from reply-text token evidence. TensorFold's generated token IDs +cannot be compared with vllm.cpp's retokenized reply text. Check the output +JSON's `refusal_reason` and sample evidence. A zero exit status alone does not +prove a valid measurement. + +For serial capture, start with the +[TensorFold environment template](../../benchmarks/manifests/qwen38_tensorfold/tensorfold.env.example) +and [vllm.cpp environment template](../../benchmarks/manifests/qwen38_tensorfold/vllm_cpp.env.example). +They list the source revisions, artifact hashes, launch commands, and server +lifecycle commands required by the +[capture runner](../../tools/bench/run_qwen38_tensorfold_gap.sh). +Fill copies outside the repository with measured local values. The runner +requires an active `dgx:gpu0` lease. These templates do not resolve the missing +artifacts or establish a successful capture. + ## Evidence and validation Versioned evidence: