Conversation
server : add images to the /decision endpoint Add an optional "images" field to POST /decision. Each entry maps 1:1 to a context and holds base64 image data, or an array of base64 strings when one context carries several images. A data URL prefix is accepted and stripped. The bitmaps are passed to the decision engine through options::context_bitmaps. Media markers are only prepended when the context text does not already contain them, so a caller that interleaves markers with its own labels keeps control of where each image lands. parallel-decision : score media chunks in the same batch A prompt part is now either text tokens or a media chunk. tokenize_mm expands media markers into chunks and reports text tokens with LLAMA_TOKEN_NULL at the image positions; decide_batch turns that into interleaved parts and decode_parts encodes chunks through the mtmd batch API before the surrounding text. Models that reject a second clip chunk with "batch too large" fall back to encoding one chunk at a time, which is what SmolVLM2 needs. Set LLAMA_DECISION_DEBUG in the environment to trace this path. examples : add vision decision harness A Flask UI that drives /decision over images and text: folder upload, batch runs with SSE progress, image selection over a folder, a text classification suite, and snippet export.
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
server : accept images on /decision, add harness example
What
POST /decisiongains an optionalimagesfield. Each entry maps 1:1 to acontext and carries base64 image data, or an array of base64 strings when one
context should see several images. A data URL prefix is accepted and stripped.
The decoded bitmaps reach the decision engine through
options::context_bitmaps.Media markers are prepended only when the context text does not already contain
them. A caller that interleaves markers with its own labels keeps control of
where each image lands; a caller that just wants "this image, this question"
keeps the old behaviour unchanged.
Engine
A
prompt_partis now either a list of text tokens or a media chunk.tokenize_mmexpands media markers into chunks and reports text tokens withLLAMA_TOKEN_NULLat the image positions.decide_batchturns that intointerleaved parts, and
decode_partsencodes chunks through the mtmd batch API(
mtmd_batch_init/add_chunk/encode/get_output_embd) before thesurrounding text, which is the path the completion endpoint already uses.
Models that reject a second clip chunk with "batch too large" fall back to
encoding one chunk at a time. SmolVLM2 needs that fallback; Qwen2.5-VL takes
the batch path.
LLAMA_DECISION_DEBUGin the environment traces this path.Example
examples/vision-decision-harness/is a Flask UI that drives the endpoint overimages and text: folder upload, batch runs with SSE progress, image selection
across a folder, a text classification suite, and snippet export. It is an
example, not a dependency, and nothing in the server links against it.
Notes
group. There is no autoregressive loop.
when a context actually has a chunk.
12-image folder picks the intended image; the text suite runs 20 files at
~80% on the 7B.