Skip to content

feat(MODEL-CLM): Contrastive-LM bi-encoder decision model - #3317

Closed
mudler-agent wants to merge 2 commits into
mainfrom
row/MODEL-CLM
Closed

mudler-agent wants to merge 2 commits into
mainfrom
row/MODEL-CLM

Conversation

@mudler-agent

Copy link
Copy Markdown
Collaborator

MODEL-CLM: Contrastive-LM bi-encoder decision model

CLM (Contrastive-LM/CLM-v0.1-8B) is a bi-encoder System 1 decision model: a frozen Qwen3-8B backbone (last-token pooling) feeds two MLP projection heads (state head + action head), each Linear(H,P)→GELU→LayerNorm(P)→Linear(P,P)→GELU→Linear(P,E), L2-normalized and scored via scaled cosine similarity (exp(logit_scale).clamp(max=100)).

Architecture

  • Bi-encoder: N+1 forward passes (1 state + N candidates), each a single-sequence prefill through Qwen3DenseModel::ForwardHidden
  • Confidence formula DIFFERS from kev/laya: max(0, min(1, p_max - mean(rest))) with NO probability rounding
  • Arch name: "ClmModel"

New files

  • include/vllm/model_executor/models/clm.h — head weights, params, forward decls
  • include/vllm/model_executor/models/clm_inference.h — ClmDecisionResult + ClmInference()
  • src/vllm/model_executor/models/clm_head.cpp — host MLP head forward, L2 norm, scoring
  • src/vllm/model_executor/models/clm_registry.cpp — factory, load, prepare, forward, inference, registration
  • tests/vllm/models/test_clm.cpp — 13 tests, 42 assertions (all pass)

Modified

  • CMakeLists.txt — add clm_head.cpp, clm_registry.cpp
  • tests/CMakeLists.txt — add test_clm
  • src/capi/vllm_c.cpp — allowlist ClmModel + ClmInference dispatch in vllm_decide
  • src/vllm/entrypoints/openai/server_main.cpp — ClmModel decision dispatch block
  • src/vllm/entrypoints/openai/systemone.cpp — BuildSystemOneAnswerClm (CLM confidence)
  • include/vllm/entrypoints/openai/systemone.h — declare BuildSystemOneAnswerClm

Spec

.agents/specs/clm.md

FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:glm5.2 [nib]

…-DECIDE specs

Add five SystemOne-class decision model specs to .agents/specs/:

- MODEL-CLM: frozen Qwen3-8B encoder + dual projection heads
  (state head + action head), scaled-cosine scoring, /v1/systemone
- MODEL-TEV1: autoregressive decision model (Qwen3.5-4B SFT),
  generates option letter via /v1/chat/completions (not /v1/systemone)
- MODEL-XOR: 35B MoE (Qwen3.6-A3B) + multimodal (up to 8 images),
  forward+reverse option-order eval, SGLang oracle
- MODEL-JEV: structured generation for DiffusionGemma (vLLM PR #57250),
  canvas seeding, read-only requests, pinned positions, logprobs
- MODEL-GLINER25-DECIDE: DeBERTa-v3-large + classification head
  (Linear→ReLU→Linear), reuses existing DeBERTa v2 encoder

Each spec follows the kev.md template structure. No implementation
code — specs only, per AGENTS.md §115 (spec before code).
… dual MLP heads)

CLM (Contrastive-LM/CLM-v0.1-8B) is a bi-encoder System 1 decision model:
a frozen Qwen3-8B backbone (last-token pooling) feeds two MLP projection
heads (state head + action head), each Linear(H,P)→GELU→LayerNorm(P)→
Linear(P,P)→GELU→Linear(P,E), L2-normalized and scored via scaled cosine
similarity (exp(logit_scale).clamp(max=100)).

Architecture: N+1 forward passes (1 state + N candidates), each a
single-sequence prefill through Qwen3DenseModel::ForwardHidden.

Confidence formula DIFFERS from kev/laya: max(0, min(1, p_max - mean(rest)))
with NO probability rounding.

New files:
  include/vllm/model_executor/models/clm.h           — head weights, params, forward decls
  include/vllm/model_executor/models/clm_inference.h — ClmDecisionResult + ClmInference()
  src/vllm/model_executor/models/clm_head.cpp        — host MLP head forward, L2 norm, scoring
  src/vllm/model_executor/models/clm_registry.cpp    — factory, load, prepare, forward, inference, registration
  tests/vllm/models/test_clm.cpp                     — 13 tests, 42 assertions (all pass)

Modified:
  CMakeLists.txt                                     — add clm_head.cpp, clm_registry.cpp
  tests/CMakeLists.txt                               — add test_clm
  src/capi/vllm_c.cpp                                — allowlist ClmModel + ClmInference dispatch
  src/vllm/entrypoints/openai/server_main.cpp        — ClmModel decision dispatch block
  src/vllm/entrypoints/openai/systemone.cpp          — BuildSystemOneAnswerClm (CLM confidence)
  include/vllm/entrypoints/openai/systemone.h        — declare BuildSystemOneAnswerClm

Spec: .agents/specs/clm.md
Issue: .agents/issues/MODEL-CLM/ISSUE-LOCAL-01M3AAH2476PHATNMMV36B3FHC.md

FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:glm5.2 [nib]
@mudler-agent

Copy link
Copy Markdown
Collaborator Author

Closing: the model implementation, tests, and docs from this PR were merged to main via #3324 (feat/parakeet-diarization). All model files are on main: clm_head.cpp, clm_registry.cpp, gliner25_decide_head.cpp, gliner25_decide_registry.cpp, xor_registry.cpp, tev1_registry.cpp, and all test files. E2E verified: CLM, GLiNER2.5-Decide, and Tev1-4B all load and serve correctly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants