Skip to content

server : accept images on /decision, add harness example - #17

Open
graydini wants to merge 1 commit into
thecodacus:parallel-decisionfrom
graydini:multimodal-decision
Open

graydini wants to merge 1 commit into
thecodacus:parallel-decisionfrom
graydini:multimodal-decision

Conversation

@graydini

@graydini graydini commented Sep 28, 2026 •

Copy link
Copy Markdown

server : accept images on /decision, add harness example

What

POST /decision gains an optional images field. Each entry maps 1:1 to a
context and carries base64 image data, or an array of base64 strings when one
context should see several images. A data URL prefix is accepted and stripped.
The decoded bitmaps reach the decision engine through
options::context_bitmaps.

Media markers are prepended only when the context text does not already contain
them. A caller that interleaves markers with its own labels keeps control of
where each image lands; a caller that just wants "this image, this question"
keeps the old behaviour unchanged.

Engine

A prompt_part is now either a list of text tokens or a media chunk.
tokenize_mm expands media markers into chunks and reports text tokens with
LLAMA_TOKEN_NULL at the image positions. decide_batch turns that into
interleaved parts, and decode_parts encodes chunks through the mtmd batch API
(mtmd_batch_init / add_chunk / encode / get_output_embd) before the
surrounding text, which is the path the completion endpoint already uses.

Models that reject a second clip chunk with "batch too large" fall back to
encoding one chunk at a time. SmolVLM2 needs that fallback; Qwen2.5-VL takes
the batch path.

LLAMA_DECISION_DEBUG in the environment traces this path.

Example

examples/vision-decision-harness/ is a Flask UI that drives the endpoint over
images and text: folder upload, batch runs with SSE progress, image selection
across a folder, a text classification suite, and snippet export. It is an
example, not a dependency, and nothing in the server links against it.

Notes

  • Vision encoding and branch scoring still complete in one batched pass per
    group. There is no autoregressive loop.
  • Text-only decisions never enter the mtmd path. The batch is only constructed
    when a context actually has a chunk.
  • Verified on Qwen2.5-VL-7B and SmolVLM2-500M. Multi-image selection over a
    12-image folder picks the intended image; the text suite runs 20 files at
    ~80% on the 7B.

server : add images to the /decision endpoint

Add an optional "images" field to POST /decision. Each entry maps 1:1
to a context and holds base64 image data, or an array of base64 strings
when one context carries several images. A data URL prefix is accepted
and stripped. The bitmaps are passed to the decision engine through
options::context_bitmaps.

Media markers are only prepended when the context text does not already
contain them, so a caller that interleaves markers with its own labels
keeps control of where each image lands.

parallel-decision : score media chunks in the same batch

A prompt part is now either text tokens or a media chunk. tokenize_mm
expands media markers into chunks and reports text tokens with
LLAMA_TOKEN_NULL at the image positions; decide_batch turns that into
interleaved parts and decode_parts encodes chunks through the mtmd batch
API before the surrounding text.

Models that reject a second clip chunk with "batch too large" fall back
to encoding one chunk at a time, which is what SmolVLM2 needs.

Set LLAMA_DECISION_DEBUG in the environment to trace this path.

examples : add vision decision harness

A Flask UI that drives /decision over images and text: folder upload,
batch runs with SSE progress, image selection over a folder, a text
classification suite, and snippet export.
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 1917877a-2469-4903-86dc-6859801fa513

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added documentation Improvements or additions to documentation server mtmd labels Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation mtmd server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant