Skip to content

Measure vllm-tt-plugin gateability on the row's P150 (qwen35 served, tokens emitted) #3261

Description

@lu-zero

Row: BACKEND-TENSTORRENT

Local-Issue: ISSUE-LOCAL-01M32FVCE8XBW2PDTE48ZD8TWN
Kind: bug

The vllm-tt-plugin oracle record (.agents/oracles/vllm-tt-plugin.md, pin 7250ddfaa, 2026-09-21) is filed with gateable = no: nothing from the plugin has executed in this repository. The oracle-registry checker requires an owing issue named on a gateable=no record; this is it. Owed measurement: install vLLM 0.26.0 (VLLM_TARGET_DEVICE=empty) plus the plugin at the pinned HEAD inside a tt-metal environment on the row's P150, serve one registered qwen35 checkpoint, and record tokens emitted. Establishing gateability also unlocks the primary-rank denominators the Tenstorrent rows want: on-device qwen35 graph correctness at bf16 (vs our GGUF-decoded bf16 arm, per-layer like the llama.cpp bisect) and the native ~50 tok/s qwen35 rate measured side by side on the same card. Source: vllm.ai blog 2026-09-07; supported architectures include TTQwen3_5ForConditionalGeneration. NOT established by this measurement: GGUF-decode parity (the plugin serves HF weights; llama.cpp stays the GGUF decode oracle).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions