Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
ID: ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ
Title: Link the Qwen3.8 measurement workflow from public docs
Row: BENCH-QWEN38-TENSORFOLD-GAP
State: OPEN
Kind: documentation
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-30
Updated: 2026-09-30
Closed: -

## Problem

The public TensorFold benchmark detail records the blocked result but does not link the shipped endpoint harness instructions or runner environment templates. Readers cannot find the reproduction prerequisites from the result page.

## Resolution

30 September 2026: Implemented README news and public benchmark links to the
existing harness guide and environment templates. The three focused CPU suites
passed all 58 tests. README structure, retained evidence, relative links, and
whitespace checks passed. Independent review and upstream landing remain
pending, so this issue stays open.
74 changes: 74 additions & 0 deletions .agents/specs/qwen38-measurement-docs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Make the Qwen3.8 measurement workflow discoverable

## Scope

Documentation follow-up for `BENCH-QWEN38-TENSORFOLD-GAP`, tracked by
`ISSUE-LOCAL-01M3R485ZSPG7VHK1Y8B2207DZ`. The benchmark lifecycle remains
`BLOCKED_MISSING_ARTIFACTS`. This change adds no implementation or measurement.

| Surface | Source | Change | Verification |
|---|---|---|---|
| README news | `tools/bench/qwen38_endpoint_bench.py:466` | Announce the endpoint measurement tool with the blocked comparison caveat | Read parser and public result |
| Public benchmark detail | `benchmarks/manifests/qwen38_tensorfold/README.md:1` | Link the harness instructions and environment templates; explain comparison limits | Resolve links and inspect verdict logic |
| Capture prerequisites | `tools/bench/run_qwen38_tensorfold_gap.sh` | Link existing runner instructions where available, without inventing a runnable GPU recipe | Inspect parser and templates |

## Design and sources

Add one short news item to `README.md` and a measurement-tools section to
`docs/benchmarks/qwen38-tensorfold-gap.md`. Reuse the committed manifest guide
instead of duplicating its command. Explain that prompt fingerprints and reply
text evidence are distinct. `PROFILE_COMPARISON` is not token parity and
supplies no cross-engine ratio. Preserve the bounded discovery scope, absent
MTP weights, existing evidence, and benchmark values.

The parser and comparison contract live in
`tools/bench/qwen38_endpoint_bench.py`. The original benchmark spec is
`bench-qwen38-tensorfold-gap.md`. This documentation change does not port vLLM
behavior. Upstream source execution and performance comparisons are not
applicable because no executable code changes.

## Tests and gates

No tests to port. Use existing CPU-only suites:

```sh
python3 -m pytest -q tests/scripts/test_qwen38_endpoint_bench.py \
tests/scripts/test_run_qwen38_tensorfold_gap.py \
tests/scripts/test_validate_qwen38_tensorfold_evidence.py
python3 scripts/check-readme-structure.py
python3 tools/bench/validate_qwen38_tensorfold_evidence.py .agents/evidence/bench-qwen38-tensorfold-gap/latest
```

Check each added relative link and run `git diff --check`. Run repository
preflight and distinguish pre-existing failures from changes introduced here.
A fresh reviewer checks an immutable commit against source. No runtime
behavior or test guarantee changes, so code mutation testing is not applicable.

## Risks and stop conditions

Do not imply a successful endpoint capture, a performance improvement, model
parity, or usable MTP weights. Do not change historical EXL3 measurements or
benchmark-index work already covered by another pull request. Use no GPU,
external compute, model downloads, or service management. Open one fork pull
request; do not merge upstream.

## Now

Documentation implemented and CPU checks passed. Independent source review and
upstream landing remain pending. The benchmark remains `BLOCKED_MISSING_ARTIFACTS`.

## Outcome

README news links to the public result, which now links the existing harness
guide and both environment templates. The text explains comparison limits,
draft configuration, and refused measurements without adding benchmark values.

On 30 September 2026, the three focused pytest suites passed all 58 tests.
The README structure check, evidence validator, relative-link checks, and
`git diff --check` passed. The evidence verdict remains
`BLOCKED_MISSING_ARTIFACTS`. Repository preflight encountered existing record,
index, environment-documentation, and release-state failures outside this scope.
The operator compares its full result with the baseline before handoff.

No executable behavior or test guarantee changed. Red-first code tests,
mutation tests, GPU execution, and new performance measurements do not apply.
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,9 @@

## News

- **2026-09** **Qwen3.8 gains an endpoint measurement tool.** The TensorFold comparison remains
blocked by missing artifacts, with no speed result. See the
[measurement instructions and limits](docs/benchmarks/qwen38-tensorfold-gap.md#measurement-tools).
- **2026-09** **Multimodal chat reaches the model through HTTP.** Qwen3-VL accepts images;
dots3-note also accepts audio and multiple media items. CPU tests use synthetic weights;
real-checkpoint token parity remains unverified. See the [input guide](docs/guides/multimodal-input.md)
Expand Down
25 changes: 25 additions & 0 deletions docs/benchmarks/qwen38-tensorfold-gap.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,31 @@ The campaign can resume measurement only after the runner prerequisites are
available in the documented/configured staging scope. It must then rerun correctness
before timing rather than reusing this blocker as benchmark evidence.

## Measurement tools

The [endpoint harness guide](../../benchmarks/manifests/qwen38_tensorfold/README.md)
includes a request example, required tokenizer identity, and the evidence contract.
The harness measures an already-running OpenAI-compatible endpoint. For the
`vllm-cpp` adapter, `--draft off` records the declared policy but does not
reconfigure speculative decoding on the running server.

Current TensorFold/vllm.cpp comparisons report `PROFILE_COMPARISON`: they prove
neither token parity nor a cross-engine ratio. Matched prompt-token fingerprints
are separate from reply-text token evidence. TensorFold's generated token IDs
cannot be compared with vllm.cpp's retokenized reply text. Check the output
JSON's `refusal_reason` and sample evidence. A zero exit status alone does not
prove a valid measurement.

For serial capture, start with the
[TensorFold environment template](../../benchmarks/manifests/qwen38_tensorfold/tensorfold.env.example)
and [vllm.cpp environment template](../../benchmarks/manifests/qwen38_tensorfold/vllm_cpp.env.example).
They list the source revisions, artifact hashes, launch commands, and server
lifecycle commands required by the
[capture runner](../../tools/bench/run_qwen38_tensorfold_gap.sh).
Fill copies outside the repository with measured local values. The runner
requires an active `dgx:gpu0` lease. These templates do not resolve the missing
artifacts or establish a successful capture.

## Evidence and validation

Versioned evidence:
Expand Down
Loading