diff --git a/.agents/issues/_owed/ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX.md b/.agents/issues/_owed/ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX.md new file mode 100644 index 000000000..f102d217d --- /dev/null +++ b/.agents/issues/_owed/ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX.md @@ -0,0 +1,19 @@ +ID: ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX +Title: Align the server reference with shipped decision and prompt-scoring routes +Row: - +State: OPEN +Kind: docs +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-10-01 +Updated: 2026-10-01 +Closed: - + +## Problem + +The server reference omits decision routes, says HTTP omits prompt_logprobs, and omits the tokenizer EOS fallback. README news omits the Tev1 decision lane. Correct the prose against source without new runtime or benchmark claims. + +## Resolution + +- diff --git a/.agents/specs/docs-server-surface-refresh.md b/.agents/specs/docs-server-surface-refresh.md new file mode 100644 index 000000000..56edc1cd2 --- /dev/null +++ b/.agents/specs/docs-server-surface-refresh.md @@ -0,0 +1,161 @@ +# Refresh the public server guide + +## Scope + +Audit base: `fce36733b`, fetched from upstream `main` on 1 October 2026. +This documentation repair describes shipped behavior. It changes no runtime, +model lifecycle, oracle pin, or benchmark result. No GPU is required. + +| Surface | Implementation anchors | Documentation repair | +|---|---|---| +| Decision routes | `src/vllm/entrypoints/openai/api_server.cpp` route registration; `src/vllm/entrypoints/openai/server_main.cpp` callback wiring; `docs/models/tev1.md` | Add the conditional decision routes to the server reference and Tev1 to README news | +| Prompt log probabilities | `serving_completion.cpp` choice builder; `serving_chat.cpp` response builder; `protocol.cpp` validation and serialization, all under `src/vllm/entrypoints/openai/` | Replace the stale HTTP-unavailable claim in the server reference and usage guide | +| EOS fallback | `src/vllm/v1/engine/input_processor.cpp`; `src/vllm/transformers_utils/hf_config.cpp`; `.agents/specs/tev1-eos-fallback.md` | Explain tokenizer fallback and config precedence in the server reference | + +## Design + +Keep reference prose short. Describe request conditions, response locations, +and registration conditions with links to existing examples and model recipes. +Give prompt log probabilities one authoritative explanation in the server +reference and link it from the usage guide. Correct the usage introduction +that calls every decision model non-generative. Tev1 samples one answer token +per question through the same engine as chat. + +README news describes the newly shipped Tev1 decision path without claiming +GPU correctness, full vLLM parity, or speed. Preserve existing benchmark values. +Avoid unrelated ABI-version edits covered by open pull request 3343. + +## Upstream chain and tests to port + +No implementation is ported. The checked-in request validators, serializers, +route registration, engine input processing, and their existing tests define +this repair. Existing model specs retain their pinned reference evidence. +No new oracle execution, performance measurement, or runtime test is applicable. + +## Gates + +- Trace every changed behavior claim to code and existing tests. +- Run `scripts/check-readme-structure.py`, `scripts/check-site.py`, and + `scripts/check-benchmark-index.py` with Python. +- Check changed local links and heading anchors. Parse any added JSON example. +- Run `git diff --check`, commit-style, and commit-trailer checks against the + pinned base. No source, test, or build file may change. +- Run the CPU-only preflight. Compare failures with the untouched baseline. + Do not claim full success if unrelated checks fail or prerequisites are absent. +- A fresh reviewer checks the immutable commit. Scratch mutations of links, + anchors, and any added JSON must fail the scoped checks. Runtime reachability + mutations do not apply to a documentation-only change. +- The operator independently reruns the focused gates before the fork push. + +## Work breakdown + +1. Commit this scope and its canonical local issue. +2. Delegate the public prose to a fresh implementer in a separate worktree. +3. Review the immutable implementation independently and repair any findings. +4. Record results here and open a reviewed pull request from the user's fork. + +## Constraints and stop conditions + +Only `README.md`, `docs/reference/server.md`, and the relevant server sections +of `docs/USAGE.md` are implementation scope. The spec and its issue carry the +work record. No lifecycle transition requires a shared matrix edit. +Use local CPU tools only. Do not download weights, use an accelerator, change +services, or merge upstream. Stop on an unresolved source contradiction. +Unspecified external environment values remain unavailable. This task needs +none. The user authorizes autonomous fork work and an upstream pull request. + +## Owed + +- ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX: repair the public server descriptions. + + +## Coordination + +The implementer uses the `SERVE-OAI-BASIC` helper role for this serving-document +repair. The canonical issue remains rowless and belongs to this spec. No +implementation or lifecycle transition of that matrix row is claimed. +The operator owns the record and the fork pull request. The implementer owns +only the three public files listed in scope. + + +## Outcome + +The public edits are committed in `88e00bf3a` and the independently requested +EOS repair in `6bf6e37a8`. The operator integrated those commits as `299cb5302` +and `ed8d708a4`. The three public files are byte-identical to the reviewed +repair head. + +The server reference now documents conditional decision routes, prompt log +probability response fields, request limits, and EOS precedence. The usage +page links to those explanations. README news links Tev1's setup and measured +limits. No runtime, test, build file, benchmark value, or lifecycle changes. +Benchmark disposition: NOT APPLICABLE to this documentation-only repair. +The existing Tev1 entry in `docs/FEATURES.md` already describes the shipped +decision lane and remains unchanged. + +Source verification uses these anchors at the audited base: + +| Claim | Code and existing tests | +|---|---| +| Conditional decision routes and Tev1 chat coexistence | `src/vllm/entrypoints/openai/api_server.cpp:1797`; `src/vllm/entrypoints/openai/server_main.cpp:1414,1610,2305`; `src/vllm/model_executor/models/tev1_inference.cpp:120` | +| Prompt log probability validation and response fields | `src/vllm/entrypoints/openai/protocol.cpp:63,845`; `serving_completion.cpp:435` and `serving_chat.cpp:1114` in that directory; `tests/vllm/entrypoints/openai/test_api_server.cpp:5308` | +| EOS precedence and missing-generation-config fallback | `src/vllm/v1/engine/input_processor.cpp:39`; `src/vllm/transformers_utils/hf_config.cpp:399,679`; `tests/vllm/v1/test_input_processor.cpp:598,611` | + +The independent reviewer found one remaining contradiction in the usage page's +old EOS explanation. A fresh implementer replaced the duplicate with a link. +The reviewer checked the repair independently and reported no remaining static +findings. Scratch mutations of a local link and the EOS heading each fail the +link check. Each restoration is byte-exact and the restored checks pass. +No JSON example was added. Runtime reachability mutations do not apply. + +Operator checks on the repaired public files, using local Python 3.14.7: + +- `python3 scripts/check-readme-structure.py`: PASS. +- `python3 scripts/check-site.py`: PASS, 12 published documents. +- `python3 -m unittest tests.scripts.test_check_readme_structure`: 19 PASS. +- Changed local links and heading anchors: 7 PASS against `fce36733b`. +- `python3 scripts/check-agent-record.py`: PASS. +- `python3 scripts/check-tree-compiles.py --base fce36733b`: no C++ source, + header, or build file in scope. +- `git diff --check`, commit style, and commit trailers: PASS. +- `python3 scripts/check-benchmark-index.py`: FAIL, with the same 16 orphan + detail files as the base. This repair changes none of those files or the index. + +No model was loaded and no accelerator was used. Source inspection and the +existing tests establish what the public descriptions mean. This session +makes no new runtime correctness or performance claim. + +The operator's full `scripts/agent-preflight.sh --quiet` exits 1 with 51 +failed checks and 14 skips. The run starts before public edits. Its later +README and site checks read the repaired files. Logs remain at +`/tmp/vllm-docs-baseline.txt` in this session's environment. + +Failures include existing benchmark-index, release-state, environment-document, +test-registration, and oracle-pin drift. The Alpine environment lacks PyYAML, +NumPy, CMake, Ninja, and binary-inspection tools. Process-management fixture +suites also fail here. The wrapper fixes its comparison base to the stale +`origin/main` at `1de097c46`, which includes unrelated upstream commits in its +style and trailer checks. Checks of this contribution instead use the fetched +base `fce36733b` and pass. The diff-scoped compilation check finds no compilation +work for this change. No full-preflight or integration-readiness success is +claimed. + +The implementation's full preflight also exits 1 with 51 failures and 14 +skips. The operator independently compares both sets with the initial run: +no failure or skip is added or removed. Its log is +`/tmp/vllm-docs-implement-preflight.txt`. The final EOS repair changes only a +usage paragraph to a checked local link. Focused checks were rerun on that +repair by its implementer, the reviewer, and the operator. + +The independent review's full preflight on `88e00bf3a` exits 1 with the same +51 failures and 14 skips. The operator compares those sets independently and +finds no additions or removals. The reviewer then checks the final repair at +`6bf6e37a8`, including the seventh link and fresh scratch mutations, and returns +PASS for the documentation. Its log is `/tmp/vllm-docs-review-preflight.txt`. + +## Now + +Documentation review is complete. The local issue stays open until the fork +pull request lands. No upstream merge is authorized by this task. Resume with +`git log --oneline -- .agents/specs/docs-server-surface-refresh.md` and the +fork branch `docs/source-alignment-20261001`. diff --git a/README.md b/README.md index 741faa9aa..f102d2347 100644 --- a/README.md +++ b/README.md @@ -37,6 +37,9 @@ ## News +- **2026-09** **Tev1 serves typed decisions and chat from one engine.** + The decision route returns probabilities for each question. See the + [Tev1 recipe](docs/models/tev1.md) for activation, CPU comparisons, and remaining validation gaps. - **2026-09** **Multimodal chat reaches the model through HTTP.** Qwen3-VL accepts images; dots3-note also accepts audio and multiple media items. CPU tests use synthetic weights; real-checkpoint token parity remains unverified. See the [input guide](docs/guides/multimodal-input.md) diff --git a/docs/USAGE.md b/docs/USAGE.md index 33fc004d4..d7d7f140f 100644 --- a/docs/USAGE.md +++ b/docs/USAGE.md @@ -696,24 +696,9 @@ Registered in | GET | `/v1/videos/{id}` | Job status | | GET | `/v1/videos/{id}/content` | The finished MP4 (`video/mp4`) | -`prompt_logprobs` is accepted on `/v1/completions` and `/v1/chat/completions` -and the engine computes it — every prompt position is scored against the token -that followed it, accumulated across chunked prefill — but the **response body -does not carry it yet**: emitting it needs the OpenAI `echo` wiring, which is -not done. Until then it is reachable through the library -(`RequestOutput.prompt_logprobs`), not over HTTP. `logprobs`/`top_logprobs` on -GENERATED tokens are emitted normally. - -That computation is gated on the **CPU** backend only. A step that owes prompt -logits takes the full-logits route, and on that route the sampler is handed a -host-resident logits buffer carrying the accelerator's device label — sound on -unified memory, and **not yet verified on CUDA at all, discrete or otherwise**. -Treat `prompt_logprobs` on a GPU build as unverified until that gate runs; the -mechanism and the exact owed invocation are in -[`.agents/specs/prompt-logprobs.md`](../.agents/specs/prompt-logprobs.md) -(risk 4 and the `PENDING` CUDA smoke gate). Requests that do NOT set it are -unaffected on every backend — the route is only taken for a step where some -request asked. +Both completion endpoints return prompt log probabilities in non-streaming +responses. See [Prompt log probabilities](reference/server.md#prompt-log-probabilities) +for request values, response fields, streaming restrictions, and verification limits. The four `/v1/videos` routes are registered **only** when the server was started with `--video-dit`; without it they are absent (404) and the server is identical @@ -736,35 +721,10 @@ ceiling. ### Which token ids stop a generation -Stop ids come from two files in the checkpoint, not one. `config.json`'s -`eos_token_id` supplies the **primary** eos id, and the sibling -`generation_config.json` supplies **secondary** stop ids that are usually a -superset of it. Gemma-4-26B is the clearest case: - -``` -config.json eos_token_id: [1, 106] -generation_config.json eos_token_id: [1, 106, 50] -``` - -Both are read, mirroring vLLM's default `--generation-config auto`. The -secondary ids are merged into the request's `stop_token_ids`, so a chat model -stops on its turn-level token rather than running to the length cap. A missing -or malformed `generation_config.json` is a silent no-op. - -When neither `config.json` nor the `tokenizer.json` post-processor names an -eos, the engine takes the tokenizer's own `eos_token` from the sibling -`tokenizer_config.json` as the primary eos id, which is vLLM's primary source. -When no `generation_config.json` exists, it also adds the text config's -`eos_token_id` as a secondary id, as vLLM's `from_model_config` fallback does. -Tev1 is such a checkpoint: it stops on `<|im_end|>` without a -`stop_token_ids` field. A checkpoint that already names its eos keeps it; the -engine does not yet move the primary id to the tokenizer's for every model, as -vLLM does ([spec](../.agents/specs/tev1-eos-fallback.md)). - -`ignore_eos: true` suppresses **all** of them, primary and secondary alike, and -generation then runs to the token budget. The ids still count toward -`min_tokens` masking either way, so `min_tokens` cannot be satisfied by emitting -a stop token early. +The engine resolves a primary end-of-sequence (EOS) ID and merges secondary +EOS IDs into `stop_token_ids`. `ignore_eos: true` suppresses both primary and +secondary EOS IDs. See the [server reference's EOS rules](reference/server.md#which-token-ids-stop-a-generation) +for config precedence, tokenizer fallback, and `min_tokens` masking. ### Server flags @@ -826,10 +786,11 @@ API surface, auth, and metrics. ## System 1 decisions with `/v1/systemone` -A decision model answers typed questions about a `state` without generating -text. The server registers `POST /v1/systemone` when the model directory -resolves to a decision architecture, and `vllm_decide` takes the same body -through the C ABI ([C API reference](reference/c-api.md#decisions-and-option-scoring)). +A decision model answers typed questions about a `state`. Depending on the +model, it scores options directly or generates an answer token. +The server registers `POST /v1/systemone` when the model directory resolves +to a decision architecture. The `vllm_decide` C API accepts the same body. +See the [C API reference](reference/c-api.md#decisions-and-option-scoring). ```sh curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{ diff --git a/docs/reference/server.md b/docs/reference/server.md index 080a4e112..a23477414 100644 --- a/docs/reference/server.md +++ b/docs/reference/server.md @@ -54,6 +54,11 @@ Registered in | POST | `/detokenize` | Detokenize token ids back to text | | GET | `/server_info` | Server info (`vllm_config`, `vllm_env`, `system_env`) | | POST | `/reset_prefix_cache` | Reset the prefix cache; returns `{"success": bool}` | +| POST | `/v1/ner` | Extract named entities. Requires an NER callback | +| POST | `/v1/systemone` | Answer typed questions. Requires an NER, decision, or request-level decision callback | +| POST | `/v1/systemone/permute` | Repeat a choice question with reordered options. Requires an NER or per-question decision callback | +| POST | `/v1/systemone/separate` | Answer each question in a separate call. Requires an NER or per-question decision callback | +| POST | `/v1/score` | Score candidate options. Requires a scoring callback | | POST | `/v1/embeddings` | Embeddings. Registered **only** when an embedder is attached, so a text server answers 404 at the route table | | POST | `/v1/audio/transcriptions` | Speech to text (multipart: audio as `file`, `response_format` as a form field). Registered **only** when a transcriber is attached | | POST | `/v1/videos` | Start a video generation job, returns `{id, status}` (MiniMax-H3) | @@ -62,6 +67,14 @@ Registered in | GET | `/v1/videos/{id}/content` | The finished MP4 (`video/mp4`) | | POST | `/v1/audio/speech` | Text (or lyrics + a music description) to audio; responds with `audio/wav` bytes. Registered **only** when a synthesizer is attached (`--speech-model`) | +The server registers decision routes according to the loaded architecture. +Nimble, CLM, and Tev1 use request-level callbacks, so their `/permute` and +`/separate` routes are absent and return HTTP 404. Tev1 requires +`architectures: ["Tev1Model"]` in `config.json` and keeps the chat routes. +See [System 1 requests](../USAGE.md#system-1-decisions-with-v1systemone), +[the decision and scoring C API](c-api.md#decisions-and-option-scoring), and the +[Tev1 recipe](../models/tev1.md). + `/v1/audio/speech` is registered only when you start the server with `--speech-model`. Without that flag, the route returns 404. MiniMax-Music3 returns a 44.1 kHz stereo WAV. See the @@ -117,31 +130,40 @@ A chat template can refuse the request itself, through an unknown message role or a kwarg value the template rejects. That answers **HTTP 400**, not 500, on both `/v1/chat/completions` and `/tokenize`. -`prompt_logprobs` is accepted on `/v1/completions` and `/v1/chat/completions` -and the engine computes it, every prompt position is scored against the token -that followed it, accumulated across chunked prefill, but the **response body -does not carry it yet**: emitting it needs the OpenAI `echo` wiring, which is -not done. Until then it is reachable through the library -(`RequestOutput.prompt_logprobs`), not over HTTP. `logprobs`/`top_logprobs` on -GENERATED tokens are emitted normally. - -That computation is gated on the **CPU** backend only. A step that owes prompt -logits takes the full-logits route, and on that route the sampler is handed a -host-resident logits buffer carrying the accelerator's device label, sound on -unified memory, and **not yet verified on CUDA at all, discrete or otherwise**. -Treat `prompt_logprobs` on a GPU build as unverified until that gate runs; the -mechanism and the exact owed invocation are in -[`.agents/specs/prompt-logprobs.md`](../../.agents/specs/prompt-logprobs.md) -(risk 4 and the `PENDING` CUDA smoke gate). Requests that do NOT set it are -unaffected on every backend, the route is only taken for a step where some -request asked. - The four `/v1/videos` routes are registered **only** when the server was started with `--video-dit`; without it they are absent (404) and the server is identical to one built without video support. Use the [MiniMax-H3 recipe](../models/minimax-h3.md) for the combined video and audio workflow. +## Prompt log probabilities + +For non-streaming requests, both completion endpoints return `prompt_logprobs`: + +| Endpoint | Response field | +|---|---| +| `/v1/completions` | `choices[i].prompt_logprobs` on every choice | +| `/v1/chat/completions` | Top-level `prompt_logprobs` | + +Set `prompt_logprobs` to a positive integer for that many top alternatives plus +the actual prompt token. Use `0` for the actual token alone, or `-1` for the +full vocabulary. Counts above the model's vocabulary size are refused. + +The response array follows prompt-token order. The first entry is `null` +because the first token has no predecessor. Later entries map token IDs to +objects with `logprob`, `rank`, and `decoded_token`. Without requested prompt +log probabilities, the response field is `null`. + +With `stream: true`, positive values and `-1` return HTTP 400. The validator +accepts `0` with streaming, but streaming responses do not carry this payload. +Other negative values return HTTP 400. The separate `echo` behavior, which +prepends prompt text and its token log probabilities to generated output, +remains unavailable. + +The recorded computation and HTTP tests use CPU fixtures. GPU correctness +remains unverified. See the [prompt-logprobs specification](../../.agents/specs/prompt-logprobs.md) +for the pending accelerator gate. + ## `max_tokens`: what a non-positive value means Some clients (Hermes among them) send `max_tokens: -1` to mean "no client-side @@ -158,20 +180,30 @@ ceiling. ## Which token ids stop a generation -Stop ids come from two files in the checkpoint, not one. `config.json`'s -`eos_token_id` supplies the **primary** eos id, and the sibling -`generation_config.json` supplies **secondary** stop ids that are usually a -superset of it. Gemma-4-26B is the clearest case: +The engine resolves the primary end-of-sequence (EOS) token in this order: -``` +1. The top-level `eos_token_id` in `config.json`, using the first integer when + the value is a list. +2. The tokenizer's EOS ID, if available. +3. `eos_token` from `tokenizer_config.json`, only if it encodes to exactly one token. + +The last fallback accepts a string or an object with a `content` string. It +does not replace an EOS ID resolved earlier. This lets checkpoints such as +[Tev1](../models/tev1.md#chat-completions) stop on their tokenizer's turn-ending token. + +Additional EOS IDs come from `config.json` and the sibling +`generation_config.json`. For example, Gemma-4-26B uses: + +```text config.json eos_token_id: [1, 106] generation_config.json eos_token_id: [1, 106, 50] ``` -Both are read, mirroring vLLM's default `--generation-config auto`. The -secondary ids are merged into the request's `stop_token_ids`, so a chat model -stops on its turn-level token rather than running to the length cap. A missing -or malformed `generation_config.json` is a silent no-op. +The engine merges secondary IDs into the request's `stop_token_ids`. When +`generation_config.json` is absent, model-config EOS IDs supply the fallback. +The top-level value wins over a nested text-config value. The nested value +applies only when the top-level value is absent or `null`. +A present but malformed `generation_config.json` supplies no additional IDs. `ignore_eos: true` suppresses **all** of them, primary and secondary alike, and generation then runs to the token budget. The ids still count toward