fix(GATE-ISSUE-INDEX-TABLE-SHAPE): declare the three rows their own specs name, so the issue index can validate a record again - #3342
Conversation
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
Source review at d8b6fc2c87acb3943043693c23fbf76a591d12ee: declaring these rows addresses the issue-index lookup, but the change introduces two failures in the broader record contract.
scripts/check-agent-record.py:49still sets the MODEL matrix count to 384, andcheck_matricesat lines 670-673 enforces it. The PR adds three recognized rows (384 -> 387) but only updates the Markdown rollup. Update the checker inventory count for the three new entries and verify the full record checker, not onlycheck-model-checklist.py/agent-issue-index.py.- The new
MODEL-GLINER25-DECIDErow isSPIKEwith ownerunassigned.check_row_contractsat lines 1015-1025 requires aCLAIM-*owner backed by a claim for every SPIKE/ACTIVE row. Assign its real owner and matching claim, or use a lifecycle state that truthfully reflects the work's ownership; do not weaken the gate.
I fetched the matrix and checker at this exact head and checked the row-selection logic with an independent JavaScript transcription: current main selects 384 model rows and this head selects 387. This is source-backed evidence, NOT an executed Python gate; Python is unavailable here.
@mudler these record-contract issues should be resolved before merge.
d8b6fc2 to
2e1ec0e
Compare
|
Both findings were already fixed before this comment landed — the review was written against 1. I had run 2. Also fixed while in there: the two checkpoint rows were written one cell short — the architecture-class cell was missing, so the table had 5 pipes where Gates at this head, all rc=0: Your JavaScript transcription of the row-selection logic was the right tool to reach for with no Python available. Worth noting the two numbers agreed exactly: 384 on |
mudler-agent
left a comment
There was a problem hiding this comment.
Changes requested: the PR says it closes its local issue and records a completed resolution, but the canonical issue remains State: OPEN with Closed: -. Close it through the repository issue workflow or remove the closure claim.
…pecs name, and the three defects that declaration was hiding
Twelve canonical records were invalid. Nine were the _intake
frozen-archive byte comparison, which is the CRLF class
ENG-EOL-BYTE-EXACT addresses. The other three were rows that three
specs and three open records already named and that no matrix
declared, so canonical_rows() could not see them and each record died
with "row is not canonical and claimable":
- MODEL-GLINER25-DECIDE, named by the heading of
specs/gliner2.5-decide.md, assigned to a new model-matrix row
under MODEL-TOKCLS by that spec's own line 22
- MODEL-DSV41-EXL3 and MODEL-DSV41-GGUF-Q1_0, the two V4.1
checkpoint campaigns named by the specs/deepseek-v4-1-flash.md
header
The gap is invisible from the tool's own output. agent-issue-index.py
--check raises on the FIRST invalid record and prints one line, and the
first record it reaches is an _intake one, so the orphan rows were
never named by anything the maintainer would run. Enumerating every
record directly finds all twelve.
THE THREE ROWS. MODEL-GLINER25-DECIDE goes in the MODEL-TOKCLS table
as SPIKE, not as its spec's "Implementation not started", because the
tree contradicts that: src/vllm/model_executor/models/
gliner25_decide_registry.cpp and gliner25_decide_head.cpp, a suite
registered at tests/CMakeLists.txt:785, and a live C API path at
src/capi/vllm_c.cpp:2106-2160. The other two go in beside
MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model, their own spec's
sibling, and both are BLOCKED, matching it. The checklist gains one
line per row with the mark its state allows, and the rollup moves
SPIKE 10 to 11, BLOCKED 5 to 7 and Total 384 to 387 in the same
commit, which is what that file's own same-commit rule requires.
THREE DEFECTS THE DECLARATION WAS HIDING, all found by replaying this
work onto d15b1cc and all fixed here.
The two checkpoint rows were one cell short. They were written with
four cells where the table carries five, and check-agent-record.py
reported "table has 5 pipes; expected 6" at model-matrix.md:166 and
:167. The architecture-class cell was the missing one, not the model
name: they opened | blocked | DeepSeek-V4.1-Flash EXL3 ... |
description | row-id while every neighbour opens | mark | class |
name | description | row-id. Both are V4.1-Flash checkpoints of one
architecture, so the cell reads DeepseekV41ForCausalLM, the class the
base row already carries. The checker reports one failure at a time,
so this was masking the next two.
The row-count constant was never bumped. check-agent-record.py pins
each matrix's expected count in MATRICES as a hardcoded literal, not
from the rollup, so the commit that moved the rollup to 387 red the
record gate with "387 MODEL rows; expected 384". Both read 387 now,
and the comment block above the constant records why in the form every
previous bump in that file uses: a new row EXISTS, never a transition
made to pass.
MODEL-GLINER25-DECIDE had no owner, and SPIKE rows must have one.
check-agent-record.py:1016 requires a CLAIM-* that actually claims the
row. The owner cell read "unassigned", which is what all 333 other
unassigned rows carry -- and every one of those is INVENTORIED, a
state that needs no claim, so this row alone failed. It is now
CLAIM-MODEL-GLINER25-DECIDE, a NEW claim file rather than a second
line on CLAIM-MODEL-GLINER25, because adding it to the sibling claim
produces "duplicate active claim CLAIM-MODEL-GLINER25": one claim owns
exactly one active row. The split matches the work, which shares the
DeBERTa v2 tower with MODEL-GLINER25 and differs only in the head
(Linear -> ReLU -> Linear to one output in place of the NER boundary
pooler), plus the /v1/systemone route and vllm_decide ABI it already
shares with kev and Laya.
TWO SPEC CORRECTIONS ship with it. gliner2.5-decide.md's ## Now said
Implementation not started and is corrected to SPIKE with the code
anchors that contradict it. deepseek-v4-1-flash.md's Matrices line
assigned both campaigns to kernel-matrix.md, which is wrong -- they
are model rows -- and is corrected to model-matrix.md with the rows
that decide it named.
MEASURED: invalid canonical records 12 to 9, the residue being exactly
the nine _intake comparisons. In a worktree carrying ENG-EOL-BYTE-EXACT
as well, agent-issue-index.py --check returns rc=0, so the two compose
rather than compete.
GATES, all rc=0: check-agent-record (MODEL=387), check-model-checklist,
check-gate-commands (143 gated rows, up from 141), check-surface-
coverage, check-conflict-markers, check-commit-trailers --range,
check-commit-style --range.
ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]
2e1ec0e to
15d48b0
Compare
…ded fix resolves The fix commit 15d48b0 declared the three orphan rows, corrected the two specs, and wrote this record's ## Resolution section, but left the header at State: OPEN with Closed: -. The record contract rejects that combination once a record is closed (scripts/issue_records.py:536-551: CLOSED requires a real Closed date and Resolution evidence), and the stale header contradicted the closure the PR itself claims, which the second review flagged. This commit flips State to CLOSED and dates Updated and Closed to 2026-10-01; the Resolution section already carries the evidence and matches the landed tree (SPIKE 11, BLOCKED 7, Total 387, and one checklist entry per new row). Gates re-run on the new head: check-agent-record.py and check-model-checklist.py both exit 0 (MODEL=387, checklist matches the row states); check-commit-style and check-commit-trailers pass on the range 15d48b0..HEAD. ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff]
Closes the local record added in this branch,
.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md.The gate cannot see the defect it is failing on
agent-issue-index.py --checkprints exactly one line:load_local_filesraises on the first invalid record, and the first record it reaches is an_intakeone, so the real residue is never named. Enumerating every record directly finds 12 invalid, in three distinct causes:_intakeProof that the split is 9-and-3 rather than a guess — the same sweep in a worktree at
row/ENG-EOL-BYTE-EXACT(97397f60b), where the working copy carries 0 CR:What the specs already said
The specs and the issues agree with each other. Only the matrices are missing.
deepseek-v4-1-flash.md:3-4namesMODEL-DSV41-EXL3andMODEL-DSV41-GGUF-Q1_0as rows of that campaign, and:9-10names their issue IDs.gliner2.5-decide.md:1carries the ID in its own heading,:22says "New model-matrix row underMODEL-TOKCLS", and:23names the issue.The rows carry the state their evidence supports
MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model— the same spec's sibling row, same wave, same 2026-09-11 scoping pass — is already declared atmodel-matrix.md:611with stateBLOCKED. The two new campaign rows follow it, and each records what this tree's own answer is:deepseek_v4_1_registry.cpp:109-135throws"the safetensors weight loader is not ported"for safetensors andDeepseekV41GgufRefusal()for any GGUF, so no V4.1 weights load in this tree in any format and W1 makes the architecture RESOLVE only.MODEL-GLINER25-DECIDEisSPIKE, not absent, because all four anchors are in this tree:src/vllm/model_executor/models/gliner25_decide_registry.cpp,gliner25_decide_head.cpptests/vllm/models/test_gliner25_decide.cpp:1-181, registered attests/CMakeLists.txt:785src/capi/vllm_c.cpp:2106-2160,Gliner25DecideInferenceat ABI v29Two spec corrections travel with the rows
gliner2.5-decide.md## NowsaidImplementation not startedwhile the same tree carried all four anchors above. The checklist row forMODEL-GLINER25already said "gliner25_decide code and test suites are in tree", so the file contradicted its own matrix too. Corrected toSPIKEwith the anchors, and the earlier claim withdrawn in the text rather than quietly overwritten.deepseek-v4-1-flash.mdassigned the two campaigns tokernel-matrix.md. That is wrong, and two rows decide it: the DSpark sibling andMODEL-DSV4-EXL3are both inmodel-matrix.md.kernel-matrix.mdadditionally pins its row count in three places (## Count invariants: "exactly 52 practical kernel-family rows", the lifecycle tally, andscripts/check-agent-record.pyholding the total), so neither row can land there without a second change to that contract — a cost, not a reason to file a model campaign as a kernel family.The gate that caught me
check-model-checklist.pywent rc=1 → rc=0, and it went red first:The checklist gains one line per row, each with the only mark its state allows (
🚫forBLOCKED,📋forSPIKE), and the rollup movesSPIKE 10→11,BLOCKED 5→7,Total 384→387in the same commit, which is what the file's own rule demands.Evidence
_intakeCRLF, owned by #3333)check-model-checklist.pyagent-issue-index.py --check(with #3333)OK: .agents\issue-index.generated.md matches canonical issue filesagent-issue-index.py --refreshThe end-state rows are measured in a worktree carrying both this branch and #3333, which is the order they will merge in; this branch alone takes the residue from 12 to 9 and cannot reach 0 on a
core.autocrlf=truecheckout, and the PR says so rather than claiming the green.check-symbol-anchors.pyoutput is byte-identical to main (5 pre-existingcheck-agent-record.pyanchors from mergec12b376b2, #3338).check-oracle-pins.pyis byte-identical to main (11 pre-existing errors, #3335).check-surface-coverage,check-gate-commands,check-readme-structure,check-build-runtime-deps,check-fusion-consistency,check-device-leakage,check-oracle-denominator-flagsare rc=0.check-commit-styleandcheck-commit-trailersover--range 1de097c46..HEADare green.No C++ build gate was run: no source file is touched here, only four markdown records.
🤖 Generated with Codebuff
Co-Authored-By: Codebuff noreply@codebuff.com
Update: rebased onto
d15b1cc09, which found three more defects in this branchReplaying this work onto current
mainsurfaced three problems in the row declarations themselves. All three are fixed in the same commit. None of them is a new row.The two checkpoint rows were one cell short.
MODEL-DSV41-EXL3andMODEL-DSV41-GGUF-Q1_0were written with four cells where the table carries five, andcheck-agent-record.pyreportedtable has 5 pipes; expected 6atmodel-matrix.md:166and:167. The missing cell was the architecture class, not the model name — they opened| blocked | DeepSeek-V4.1-Flash EXL3 ... | description | row-idwhile every neighbour opens| mark | class | name | description | row-id. Both are V4.1-Flash checkpoints of one architecture, so the cell now readsDeepseekV41ForCausalLM, the class the base row already carries. The checker reports one failure at a time, so this was masking the next two.The row-count constant was never bumped.
check-agent-record.pypins each matrix's expected count inMATRICESas a hardcoded literal, not from the rollup. The rollup moved 384 to 387, the literal stayed 384, and the same commit that fixed the index red the record gate with387 MODEL rows; expected 384. Both read 387 now, with the rationale recorded in the comment block above the constant in the form every previous bump in that file uses.MODEL-GLINER25-DECIDEhad no owner, andSPIKErows must have one.check-agent-record.py:1016requires aCLAIM-*that actually claims the row. The owner cell readunassigned— which is what all 333 other unassigned rows carry, and every one of those isINVENTORIED, a state that needs no claim. This row alone failed.It is now
CLAIM-MODEL-GLINER25-DECIDE, a new claim file rather than a second line onCLAIM-MODEL-GLINER25, because adding it to the sibling claim producesduplicate active claim CLAIM-MODEL-GLINER25— one claim owns exactly one active row. The split also matches the work: it shares the DeBERTa v2 tower withMODEL-GLINER25and differs only in the head (Linear -> ReLU -> Linearto one output, in place of the NER boundary pooler), plus the/v1/systemoneroute andvllm_decideABI it already shares with kev and Laya.check-agent-record.pyreports one failure at a time, which is why the original finding needed a full enumeration of every record to see past the first_intakeone. The same property hid these three behind each other.Gates, all rc=0:
check-agent-record(MODEL=387),check-model-checklist,check-gate-commands(143 gated rows, up from 141),check-surface-coverage,check-conflict-markers,check-commit-trailers --range,check-commit-style --range.agent-issue-index.py --checkis still rc=2 here, on the_intakefrozen-archive byte comparison — the CRLF class #3333 addresses. In a worktree carrying both, that check reaches these rows and validates them, so the two compose rather than compete.FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:codebuff/buffy [freebuff]