Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions .agents/issues/_owed/ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
ID: ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX
Title: Align the server reference with shipped decision and prompt-scoring routes
Row: -
State: OPEN
Kind: docs
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-01
Updated: 2026-10-01
Closed: -

## Problem

The server reference omits decision routes, says HTTP omits prompt_logprobs, and omits the tokenizer EOS fallback. README news omits the Tev1 decision lane. Correct the prose against source without new runtime or benchmark claims.

## Resolution

-
161 changes: 161 additions & 0 deletions .agents/specs/docs-server-surface-refresh.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,161 @@
# Refresh the public server guide

## Scope

Audit base: `fce36733b`, fetched from upstream `main` on 1 October 2026.
This documentation repair describes shipped behavior. It changes no runtime,
model lifecycle, oracle pin, or benchmark result. No GPU is required.

| Surface | Implementation anchors | Documentation repair |
|---|---|---|
| Decision routes | `src/vllm/entrypoints/openai/api_server.cpp` route registration; `src/vllm/entrypoints/openai/server_main.cpp` callback wiring; `docs/models/tev1.md` | Add the conditional decision routes to the server reference and Tev1 to README news |
| Prompt log probabilities | `serving_completion.cpp` choice builder; `serving_chat.cpp` response builder; `protocol.cpp` validation and serialization, all under `src/vllm/entrypoints/openai/` | Replace the stale HTTP-unavailable claim in the server reference and usage guide |
| EOS fallback | `src/vllm/v1/engine/input_processor.cpp`; `src/vllm/transformers_utils/hf_config.cpp`; `.agents/specs/tev1-eos-fallback.md` | Explain tokenizer fallback and config precedence in the server reference |

## Design

Keep reference prose short. Describe request conditions, response locations,
and registration conditions with links to existing examples and model recipes.
Give prompt log probabilities one authoritative explanation in the server
reference and link it from the usage guide. Correct the usage introduction
that calls every decision model non-generative. Tev1 samples one answer token
per question through the same engine as chat.

README news describes the newly shipped Tev1 decision path without claiming
GPU correctness, full vLLM parity, or speed. Preserve existing benchmark values.
Avoid unrelated ABI-version edits covered by open pull request 3343.

## Upstream chain and tests to port

No implementation is ported. The checked-in request validators, serializers,
route registration, engine input processing, and their existing tests define
this repair. Existing model specs retain their pinned reference evidence.
No new oracle execution, performance measurement, or runtime test is applicable.

## Gates

- Trace every changed behavior claim to code and existing tests.
- Run `scripts/check-readme-structure.py`, `scripts/check-site.py`, and
`scripts/check-benchmark-index.py` with Python.
- Check changed local links and heading anchors. Parse any added JSON example.
- Run `git diff --check`, commit-style, and commit-trailer checks against the
pinned base. No source, test, or build file may change.
- Run the CPU-only preflight. Compare failures with the untouched baseline.
Do not claim full success if unrelated checks fail or prerequisites are absent.
- A fresh reviewer checks the immutable commit. Scratch mutations of links,
anchors, and any added JSON must fail the scoped checks. Runtime reachability
mutations do not apply to a documentation-only change.
- The operator independently reruns the focused gates before the fork push.

## Work breakdown

1. Commit this scope and its canonical local issue.
2. Delegate the public prose to a fresh implementer in a separate worktree.
3. Review the immutable implementation independently and repair any findings.
4. Record results here and open a reviewed pull request from the user's fork.

## Constraints and stop conditions

Only `README.md`, `docs/reference/server.md`, and the relevant server sections
of `docs/USAGE.md` are implementation scope. The spec and its issue carry the
work record. No lifecycle transition requires a shared matrix edit.
Use local CPU tools only. Do not download weights, use an accelerator, change
services, or merge upstream. Stop on an unresolved source contradiction.
Unspecified external environment values remain unavailable. This task needs
none. The user authorizes autonomous fork work and an upstream pull request.

## Owed

- ISSUE-LOCAL-01M3TPN0QZ8856GY1MHN2TKNHX: repair the public server descriptions.


## Coordination

The implementer uses the `SERVE-OAI-BASIC` helper role for this serving-document
repair. The canonical issue remains rowless and belongs to this spec. No
implementation or lifecycle transition of that matrix row is claimed.
The operator owns the record and the fork pull request. The implementer owns
only the three public files listed in scope.


## Outcome

The public edits are committed in `88e00bf3a` and the independently requested
EOS repair in `6bf6e37a8`. The operator integrated those commits as `299cb5302`
and `ed8d708a4`. The three public files are byte-identical to the reviewed
repair head.

The server reference now documents conditional decision routes, prompt log
probability response fields, request limits, and EOS precedence. The usage
page links to those explanations. README news links Tev1's setup and measured
limits. No runtime, test, build file, benchmark value, or lifecycle changes.
Benchmark disposition: NOT APPLICABLE to this documentation-only repair.
The existing Tev1 entry in `docs/FEATURES.md` already describes the shipped
decision lane and remains unchanged.

Source verification uses these anchors at the audited base:

| Claim | Code and existing tests |
|---|---|
| Conditional decision routes and Tev1 chat coexistence | `src/vllm/entrypoints/openai/api_server.cpp:1797`; `src/vllm/entrypoints/openai/server_main.cpp:1414,1610,2305`; `src/vllm/model_executor/models/tev1_inference.cpp:120` |
| Prompt log probability validation and response fields | `src/vllm/entrypoints/openai/protocol.cpp:63,845`; `serving_completion.cpp:435` and `serving_chat.cpp:1114` in that directory; `tests/vllm/entrypoints/openai/test_api_server.cpp:5308` |
| EOS precedence and missing-generation-config fallback | `src/vllm/v1/engine/input_processor.cpp:39`; `src/vllm/transformers_utils/hf_config.cpp:399,679`; `tests/vllm/v1/test_input_processor.cpp:598,611` |

The independent reviewer found one remaining contradiction in the usage page's
old EOS explanation. A fresh implementer replaced the duplicate with a link.
The reviewer checked the repair independently and reported no remaining static
findings. Scratch mutations of a local link and the EOS heading each fail the
link check. Each restoration is byte-exact and the restored checks pass.
No JSON example was added. Runtime reachability mutations do not apply.

Operator checks on the repaired public files, using local Python 3.14.7:

- `python3 scripts/check-readme-structure.py`: PASS.
- `python3 scripts/check-site.py`: PASS, 12 published documents.
- `python3 -m unittest tests.scripts.test_check_readme_structure`: 19 PASS.
- Changed local links and heading anchors: 7 PASS against `fce36733b`.
- `python3 scripts/check-agent-record.py`: PASS.
- `python3 scripts/check-tree-compiles.py --base fce36733b`: no C++ source,
header, or build file in scope.
- `git diff --check`, commit style, and commit trailers: PASS.
- `python3 scripts/check-benchmark-index.py`: FAIL, with the same 16 orphan
detail files as the base. This repair changes none of those files or the index.

No model was loaded and no accelerator was used. Source inspection and the
existing tests establish what the public descriptions mean. This session
makes no new runtime correctness or performance claim.

The operator's full `scripts/agent-preflight.sh --quiet` exits 1 with 51
failed checks and 14 skips. The run starts before public edits. Its later
README and site checks read the repaired files. Logs remain at
`/tmp/vllm-docs-baseline.txt` in this session's environment.

Failures include existing benchmark-index, release-state, environment-document,
test-registration, and oracle-pin drift. The Alpine environment lacks PyYAML,
NumPy, CMake, Ninja, and binary-inspection tools. Process-management fixture
suites also fail here. The wrapper fixes its comparison base to the stale
`origin/main` at `1de097c46`, which includes unrelated upstream commits in its
style and trailer checks. Checks of this contribution instead use the fetched
base `fce36733b` and pass. The diff-scoped compilation check finds no compilation
work for this change. No full-preflight or integration-readiness success is
claimed.

The implementation's full preflight also exits 1 with 51 failures and 14
skips. The operator independently compares both sets with the initial run:
no failure or skip is added or removed. Its log is
`/tmp/vllm-docs-implement-preflight.txt`. The final EOS repair changes only a
usage paragraph to a checked local link. Focused checks were rerun on that
repair by its implementer, the reviewer, and the operator.

The independent review's full preflight on `88e00bf3a` exits 1 with the same
51 failures and 14 skips. The operator compares those sets independently and
finds no additions or removals. The reviewer then checks the final repair at
`6bf6e37a8`, including the seventh link and fresh scratch mutations, and returns
PASS for the documentation. Its log is `/tmp/vllm-docs-review-preflight.txt`.

## Now

Documentation review is complete. The local issue stays open until the fork
pull request lands. No upstream merge is authorized by this task. Resume with
`git log --oneline -- .agents/specs/docs-server-surface-refresh.md` and the
fork branch `docs/source-alignment-20261001`.
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,9 @@

## News

- **2026-09** **Tev1 serves typed decisions and chat from one engine.**
The decision route returns probabilities for each question. See the
[Tev1 recipe](docs/models/tev1.md) for activation, CPU comparisons, and remaining validation gaps.
- **2026-09** **Multimodal chat reaches the model through HTTP.** Qwen3-VL accepts images;
dots3-note also accepts audio and multiple media items. CPU tests use synthetic weights;
real-checkpoint token parity remains unverified. See the [input guide](docs/guides/multimodal-input.md)
Expand Down
63 changes: 12 additions & 51 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -696,24 +696,9 @@ Registered in
| GET | `/v1/videos/{id}` | Job status |
| GET | `/v1/videos/{id}/content` | The finished MP4 (`video/mp4`) |

`prompt_logprobs` is accepted on `/v1/completions` and `/v1/chat/completions`
and the engine computes it — every prompt position is scored against the token
that followed it, accumulated across chunked prefill — but the **response body
does not carry it yet**: emitting it needs the OpenAI `echo` wiring, which is
not done. Until then it is reachable through the library
(`RequestOutput.prompt_logprobs`), not over HTTP. `logprobs`/`top_logprobs` on
GENERATED tokens are emitted normally.

That computation is gated on the **CPU** backend only. A step that owes prompt
logits takes the full-logits route, and on that route the sampler is handed a
host-resident logits buffer carrying the accelerator's device label — sound on
unified memory, and **not yet verified on CUDA at all, discrete or otherwise**.
Treat `prompt_logprobs` on a GPU build as unverified until that gate runs; the
mechanism and the exact owed invocation are in
[`.agents/specs/prompt-logprobs.md`](../.agents/specs/prompt-logprobs.md)
(risk 4 and the `PENDING` CUDA smoke gate). Requests that do NOT set it are
unaffected on every backend — the route is only taken for a step where some
request asked.
Both completion endpoints return prompt log probabilities in non-streaming
responses. See [Prompt log probabilities](reference/server.md#prompt-log-probabilities)
for request values, response fields, streaming restrictions, and verification limits.

The four `/v1/videos` routes are registered **only** when the server was started
with `--video-dit`; without it they are absent (404) and the server is identical
Expand All @@ -736,35 +721,10 @@ ceiling.

### Which token ids stop a generation

Stop ids come from two files in the checkpoint, not one. `config.json`'s
`eos_token_id` supplies the **primary** eos id, and the sibling
`generation_config.json` supplies **secondary** stop ids that are usually a
superset of it. Gemma-4-26B is the clearest case:

```
config.json eos_token_id: [1, 106]
generation_config.json eos_token_id: [1, 106, 50]
```

Both are read, mirroring vLLM's default `--generation-config auto`. The
secondary ids are merged into the request's `stop_token_ids`, so a chat model
stops on its turn-level token rather than running to the length cap. A missing
or malformed `generation_config.json` is a silent no-op.

When neither `config.json` nor the `tokenizer.json` post-processor names an
eos, the engine takes the tokenizer's own `eos_token` from the sibling
`tokenizer_config.json` as the primary eos id, which is vLLM's primary source.
When no `generation_config.json` exists, it also adds the text config's
`eos_token_id` as a secondary id, as vLLM's `from_model_config` fallback does.
Tev1 is such a checkpoint: it stops on `<|im_end|>` without a
`stop_token_ids` field. A checkpoint that already names its eos keeps it; the
engine does not yet move the primary id to the tokenizer's for every model, as
vLLM does ([spec](../.agents/specs/tev1-eos-fallback.md)).

`ignore_eos: true` suppresses **all** of them, primary and secondary alike, and
generation then runs to the token budget. The ids still count toward
`min_tokens` masking either way, so `min_tokens` cannot be satisfied by emitting
a stop token early.
The engine resolves a primary end-of-sequence (EOS) ID and merges secondary
EOS IDs into `stop_token_ids`. `ignore_eos: true` suppresses both primary and
secondary EOS IDs. See the [server reference's EOS rules](reference/server.md#which-token-ids-stop-a-generation)
for config precedence, tokenizer fallback, and `min_tokens` masking.

### Server flags

Expand Down Expand Up @@ -826,10 +786,11 @@ API surface, auth, and metrics.

## System 1 decisions with `/v1/systemone`

A decision model answers typed questions about a `state` without generating
text. The server registers `POST /v1/systemone` when the model directory
resolves to a decision architecture, and `vllm_decide` takes the same body
through the C ABI ([C API reference](reference/c-api.md#decisions-and-option-scoring)).
A decision model answers typed questions about a `state`. Depending on the
model, it scores options directly or generates an answer token.
The server registers `POST /v1/systemone` when the model directory resolves
to a decision architecture. The `vllm_decide` C API accepts the same body.
See the [C API reference](reference/c-api.md#decisions-and-option-scoring).

```sh
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
Expand Down
Loading
Loading