From cf95ee37ad4948b06e8a014ae8da3a7e2b5841fe Mon Sep 17 00:00:00 2001 From: graydini Date: Sun, 27 Sep 2026 21:16:54 -0700 Subject: [PATCH 1/3] server : accept images on /decision, add harness example server : add images to the /decision endpoint Add an optional "images" field to POST /decision. Each entry maps 1:1 to a context and holds base64 image data, or an array of base64 strings when one context carries several images. A data URL prefix is accepted and stripped. The bitmaps are passed to the decision engine through options::context_bitmaps. Media markers are only prepended when the context text does not already contain them, so a caller that interleaves markers with its own labels keeps control of where each image lands. parallel-decision : score media chunks in the same batch A prompt part is now either text tokens or a media chunk. tokenize_mm expands media markers into chunks and reports text tokens with LLAMA_TOKEN_NULL at the image positions; decide_batch turns that into interleaved parts and decode_parts encodes chunks through the mtmd batch API before the surrounding text. Models that reject a second clip chunk with "batch too large" fall back to encoding one chunk at a time, which is what SmolVLM2 needs. Set LLAMA_DECISION_DEBUG in the environment to trace this path. examples : add vision decision harness A Flask UI that drives /decision over images and text: folder upload, batch runs with SSE progress, image selection over a folder, a text classification suite, and snippet export. --- tools/mtmd/mtmd.h | 5 + tools/parallel-decision/CMakeLists.txt | 2 +- tools/parallel-decision/README.md | 33 + tools/parallel-decision/decision-engine.cpp | 249 +++++- tools/parallel-decision/decision-engine.h | 46 +- .../vision-decision-harness/.gitignore | 7 + .../vision-decision-harness/README.md | 152 ++++ .../examples/vision-decision-harness/app.py | 507 ++++++++++++ .../vision-decision-harness/requirements.txt | 3 + .../templates/index.html | 749 ++++++++++++++++++ .../tests/answer_key.json | 190 +++++ .../tests/evaluate_text_tests.py | 195 +++++ .../tests/test_decision_scoring.py | 633 +++++++++++++++ .../tests/text_samples/answer_key.json | 159 ++++ .../tests/text_samples/schema.json | 12 + .../tests/text_samples/t01.txt | 1 + .../tests/text_samples/t02.txt | 1 + .../tests/text_samples/t03.txt | 1 + .../tests/text_samples/t04.txt | 1 + .../tests/text_samples/t05.txt | 1 + .../tests/text_samples/t06.txt | 1 + .../tests/text_samples/t07.txt | 1 + .../tests/text_samples/t08.txt | 1 + .../tests/text_samples/t09.txt | 1 + .../tests/text_samples/t10.txt | 1 + .../tests/text_samples/t11.txt | 1 + .../tests/text_samples/t12.txt | 1 + .../tests/text_samples/t13.txt | 1 + .../tests/text_samples/t14.txt | 1 + .../tests/text_samples/t15.txt | 1 + .../tests/text_samples/t16.txt | 1 + .../tests/text_samples/t17.txt | 1 + .../tests/text_samples/t18.txt | 1 + .../tests/text_samples/t19.txt | 1 + .../tests/text_samples/t20.txt | 1 + tools/server/server-context.cpp | 91 ++- 36 files changed, 3031 insertions(+), 22 deletions(-) create mode 100644 tools/parallel-decision/examples/vision-decision-harness/.gitignore create mode 100644 tools/parallel-decision/examples/vision-decision-harness/README.md create mode 100644 tools/parallel-decision/examples/vision-decision-harness/app.py create mode 100644 tools/parallel-decision/examples/vision-decision-harness/requirements.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/templates/index.html create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt create mode 100644 tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt diff --git a/tools/mtmd/mtmd.h b/tools/mtmd/mtmd.h index c2de26eeee2c..85ca905afaff 100644 --- a/tools/mtmd/mtmd.h +++ b/tools/mtmd/mtmd.h @@ -507,6 +507,11 @@ struct bitmap { struct bitmaps { std::vector entries; ~bitmaps() = default; + bitmaps() = default; + bitmaps(bitmaps && other) noexcept = default; + bitmaps & operator=(bitmaps && other) noexcept = default; + bitmaps(const bitmaps &) = delete; + bitmaps & operator=(const bitmaps &) = delete; // return list of pointers to mtmd_bitmap // example: // auto bitmaps_c_ptr = bitmaps.c_ptr(); diff --git a/tools/parallel-decision/CMakeLists.txt b/tools/parallel-decision/CMakeLists.txt index f4d41d2f27e9..2c7ca8d68e02 100644 --- a/tools/parallel-decision/CMakeLists.txt +++ b/tools/parallel-decision/CMakeLists.txt @@ -1,7 +1,7 @@ # shared engine: used by llama-parallel-decision and by llama-server's /decision endpoint add_library(llama-decision STATIC decision-engine.cpp decision-engine.h) target_include_directories(llama-decision PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}) -target_link_libraries(llama-decision PUBLIC llama-common llama) +target_link_libraries(llama-decision PUBLIC llama-common llama mtmd) target_compile_features(llama-decision PUBLIC cxx_std_17) # llama-server links it into a shared library (libllama-server-impl) set_target_properties(llama-decision PROPERTIES POSITION_INDEPENDENT_CODE ON) diff --git a/tools/parallel-decision/README.md b/tools/parallel-decision/README.md index 76e86b90e773..eb372d9e794e 100644 --- a/tools/parallel-decision/README.md +++ b/tools/parallel-decision/README.md @@ -132,3 +132,36 @@ line). Environment: `DECIDE_TREE`, `DECIDE_TREE_MAX`, `DECIDE_NSEQ`, `DECIDE_SPL [decision-playground](https://github.com/thecodacus/decision-playground) is a browser-only playground: it talks straight to your llama-server, runs a decision and the same question as a chat completion side by side with live timers, and has a small game whose agents decide through the endpoint. + +### Vision Decision Harness (Example) + +For multimodal decision testing with image support, this directory ships a runnable example under +`examples/vision-decision-harness/`. It is a small Flask web UI that: + +- Accepts image folder uploads and runs them through `/decision` in a batch +- Adds an image selection mode: every image in a folder is scored in one decision pass and the + UI reports the single best match for a question +- Streams results back over SSE as each file is scored +- Ships a text test suite for measuring classification accuracy and calibration +- Ships `tests/scan_for_secrets.py`, which uses the decision endpoint itself to flag files that + look like they contain private data before you commit + +Build and run the server first: + +```bash +./build/bin/llama-server --host 0.0.0.0 --port 8081 \ + -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \ + --decision-seqs 8 +``` + +Then the harness: + +```bash +cd tools/parallel-decision/examples/vision-decision-harness +pip install -r requirements.txt +python3 app.py +``` + +The harness listens on port 5786 and proxies to the server on port 8081. Set `LLAMA_SERVER_URL` +to point it somewhere else. + diff --git a/tools/parallel-decision/decision-engine.cpp b/tools/parallel-decision/decision-engine.cpp index c0f1d65000c5..cf630b942368 100644 --- a/tools/parallel-decision/decision-engine.cpp +++ b/tools/parallel-decision/decision-engine.cpp @@ -2,6 +2,7 @@ #include "chat.h" #include "common.h" +#include "mtmd-helper.h" #include #include @@ -10,6 +11,10 @@ #include #include +// Toggle debug output with LLAMA_DECISION_DEBUG env var +namespace { bool debug_enabled() { static bool v = std::getenv("LLAMA_DECISION_DEBUG") != nullptr; return v; } } +#define DECISION_DEBUG(fmt, ...) do { if (debug_enabled()) { fprintf(stderr, "[decision-debug] " fmt "\n", ##__VA_ARGS__); } } while(0) + namespace llama_decision { namespace { @@ -173,10 +178,11 @@ struct decision_field { // ---------------------------------------------------------------- engine -engine::engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs) +engine::engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs, mtmd_context * mctx) : ctx(ctx), vocab(llama_model_get_vocab(llama_get_model(ctx))), mem(llama_get_memory(ctx)), seq_snap(seq_base), seq_pool(seq_base + 1), n_pool(n_seqs - 1), - pad_branches(llama_model_is_recurrent(llama_get_model(ctx)) || llama_model_is_hybrid(llama_get_model(ctx))) { + pad_branches(llama_model_is_recurrent(llama_get_model(ctx)) || llama_model_is_hybrid(llama_get_model(ctx))), + mctx(mctx) { if (n_seqs < 3) { throw std::invalid_argument("a decision engine needs at least 3 sequences"); } @@ -192,11 +198,94 @@ tokens_t engine::tokenize(const std::string & text, bool add_special) const { return toks; } +// Tokenize text that may contain media markers. When mctx is available and the text +// contains a media marker, uses mtmd_tokenize to split the text into text chunks and +// image chunks. Text tokens with LLAMA_TOKEN_NULL at image positions are returned, +// and the image chunks are collected for separate encoding in decode_parts. +engine::multimodal_tokens engine::tokenize_mm(const std::string & text, bool add_special, + const mtmd::bitmaps * bitmaps) const { + multimodal_tokens result; + + if (!mctx) { + // no multimodal context, fall back to regular tokenization + result.toks = tokenize(text, add_special); + return result; + } + + const char * marker = mctx ? mtmd_get_marker(mctx) : nullptr; + if (marker == nullptr || text.find(marker) == std::string::npos) { + // no media marker in text, use regular tokenization + result.toks = tokenize(text, add_special); + return result; + } + + // Use mtmd_tokenize to properly split text and image chunks, passing bitmaps + mtmd::input_chunks chunks(mtmd_input_chunks_init()); // MUST be initialized! + mtmd_input_text input_text; + input_text.text = text.c_str(); + input_text.text_len = text.size(); + input_text.add_special = add_special; + input_text.parse_special = true; + + auto bmp_ptr = [bitmaps]() { + std::vector res; + if (bitmaps) { + res.reserve(bitmaps->entries.size()); + for (const auto & b : bitmaps->entries) { + res.push_back(b.ptr.get()); + } + } + return res; + }(); + int32_t rc = mtmd_tokenize(mctx, chunks.ptr.get(), &input_text, bmp_ptr.data(), (int32_t) bmp_ptr.size()); + DECISION_DEBUG("tokenize_mm: mtmd_tokenize rc=%d chunks.size=%zu", (int)rc, chunks.size()); + if (rc != 0) { + // fall back to regular tokenization on error + result.toks = tokenize(text, add_special); + return result; + } + + // Extract tokens from chunks, interleaving text tokens and LLAMA_TOKEN_NULL for images + for (size_t i = 0; i < chunks.size(); ++i) { + const mtmd_input_chunk * chunk = chunks[i]; + enum mtmd_input_chunk_type type = mtmd_input_chunk_get_type(chunk); + if (type == MTMD_INPUT_CHUNK_TYPE_TEXT) { + size_t n_tokens = 0; + const llama_token * toks = mtmd_input_chunk_get_tokens_text(chunk, &n_tokens); + for (size_t j = 0; j < n_tokens; ++j) { + result.toks.push_back(toks[j]); + } + } else if (type == MTMD_INPUT_CHUNK_TYPE_IMAGE) { + // image chunk: insert LLAMA_TOKEN_NULL placeholder and record the chunk copy + size_t n_tokens = mtmd_input_chunk_get_n_tokens(chunk); + result.toks.insert(result.toks.end(), n_tokens, LLAMA_TOKEN_NULL); + // copy the chunk so it stays valid after chunks object is destroyed + result.chunks.push_back(mtmd_input_chunk_copy(chunk)); + } else if (type == MTMD_INPUT_CHUNK_TYPE_AUDIO) { + // audio chunk: insert LLAMA_TOKEN_NULL placeholder and record the chunk copy + size_t n_tokens = mtmd_input_chunk_get_n_tokens(chunk); + result.toks.insert(result.toks.end(), n_tokens, LLAMA_TOKEN_NULL); + result.chunks.push_back(mtmd_input_chunk_copy(chunk)); + } + } + + // deduplicate BOS if needed (consistent with tokenize()) + const llama_token bos = llama_vocab_bos(vocab); + if (result.toks.size() >= 2 && result.toks[0] == bos && result.toks[1] == bos) { + result.toks.erase(result.toks.begin()); + } + + return result; +} + // Decode several prompts, each on its own sequence, packed into as few batches as n_batch allows. +// Image chunks are encoded via the mtmd batch API (mtmd_batch_init/add_chunk/encode/get_output_embd) +// followed by mtmd_helper_decode_image_chunk, the same path the server's completion endpoint uses. +// Text tokens use a plain llama_batch for batched llama_decode. void engine::decode_parts(const std::vector & parts) { const int n_batch = (int) llama_n_batch(ctx); llama_batch batch = llama_batch_init(n_batch, 0, 1); - auto flush = [&]() { + auto flush_text = [&]() { const int rc = batch.n_tokens > 0 ? llama_decode(ctx, batch) : 0; common_batch_clear(batch); if (rc != 0) { @@ -205,15 +294,93 @@ void engine::decode_parts(const std::vector & parts) { : "llama_decode failed on the decision prompt (" + std::to_string(rc) + ")"); } }; + // Collect image chunks that need encoding + std::vector img_chunks; + for (const auto & p : parts) { + if (p.kind == prompt_part::CHUNK) { + img_chunks.push_back(p.chunk); + } + } + + // Encode image chunks via the mtmd batch API, the same path the server's completion + // endpoint uses. Some models (e.g. SmolVLM2) do not support batching in the CLIP context, + // so mtmd_batch_add_chunk rejects the second chunk with "batch too large" (rc=2). In that + // case each chunk is encoded on its own with mtmd_encode_chunk instead. + // Text-only decisions never reach this and skip the whole block. + mtmd::batch_ptr mbatch; + bool use_batch = false; + if (!img_chunks.empty() && mctx) { + mbatch.reset(mtmd_batch_init(mctx)); + use_batch = true; + for (auto * chunk : img_chunks) { + DECISION_DEBUG("decode_parts: adding image chunk with n_tokens=%d", (int) mtmd_input_chunk_get_n_tokens(chunk)); + int32_t add_rc = mtmd_batch_add_chunk(mbatch.get(), chunk); + DECISION_DEBUG("decode_parts: mtmd_batch_add_chunk rc=%d", (int) add_rc); + if (add_rc == 2) { + // batch too large: this model does not support batched clip encoding + use_batch = false; + mbatch.reset(); + break; + } + if (add_rc != 0) { + throw std::runtime_error("mtmd_batch_add_chunk failed (" + std::to_string(add_rc) + ")"); + } + } + if (use_batch) { + int32_t enc_rc = mtmd_batch_encode(mbatch.get()); + DECISION_DEBUG("decode_parts: mtmd_batch_encode rc=%d", (int) enc_rc); + if (enc_rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_batch_encode failed on the decision prompt (" + std::to_string(enc_rc) + ")"); + } + } + } + + // Now iterate parts: decode text via llama_batch, decode image via mtmd_helper_decode_image_chunk for (const auto & p : parts) { - for (size_t i = 0; i < p.toks->size(); ++i) { - if (batch.n_tokens == n_batch) { - flush(); + if (p.kind == prompt_part::CHUNK) { + // Flush pending text tokens first so image decoding starts at the right position + flush_text(); + DECISION_DEBUG("decode_parts: decoding CHUNK pos0=%d seq=%d", (int) p.pos0, (int) p.seq); + float * embd = nullptr; + if (use_batch && mbatch) { + embd = mtmd_batch_get_output_embd(mbatch.get(), p.chunk); + DECISION_DEBUG("decode_parts: embd ptr from batch=%p", (void*)embd); + } else { + // Non-batch fallback: encode this chunk individually + DECISION_DEBUG("decode_parts: encoding chunk individually via mtmd_encode_chunk"); + int32_t enc_rc = mtmd_encode_chunk(mctx, p.chunk); + DECISION_DEBUG("decode_parts: mtmd_encode_chunk rc=%d", (int) enc_rc); + if (enc_rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_encode_chunk failed on the decision prompt (" + std::to_string(enc_rc) + ")"); + } + embd = mtmd_get_output_embd(mctx); + DECISION_DEBUG("decode_parts: embd ptr from mtmd=%p", (void*)embd); + } + if (!embd) { + llama_batch_free(batch); + throw std::runtime_error("failed to get image embedding for chunk"); + } + DECISION_DEBUG("decode_parts: mtmd_helper_decode_image_chunk n_batch=%d", n_batch); + llama_pos new_n_past = p.pos0; + int32_t rc = mtmd_helper_decode_image_chunk(mctx, ctx, p.chunk, embd, p.pos0, p.seq, n_batch, &new_n_past, nullptr, nullptr); + DECISION_DEBUG("decode_parts: mtmd_helper_decode_image_chunk rc=%d new_n_past=%d", (int) rc, (int) new_n_past); + if (rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_helper_decode_image_chunk failed on the decision prompt (" + std::to_string(rc) + ")"); + } + } else { + DECISION_DEBUG("decode_parts: TOKS pos0=%d seq=%d n_toks=%zu", (int) p.pos0, (int) p.seq, p.toks->size()); + for (size_t i = 0; i < p.toks->size(); ++i) { + if (batch.n_tokens == n_batch) { + flush_text(); + } + common_batch_add(batch, (*p.toks)[i], p.pos0 + (llama_pos) i, { p.seq }, false); } - common_batch_add(batch, (*p.toks)[i], p.pos0 + (llama_pos) i, { p.seq }, false); } } - flush(); + flush_text(); llama_batch_free(batch); } @@ -229,7 +396,9 @@ bool engine::prepare_prefix(const tokens_t & shared, bool allow_cache) { } cached.clear(); if (!shared.empty()) { - decode_parts({ { &shared, 0, seq_snap } }); + DECISION_DEBUG("prepare_prefix: calling decode_parts with shared.size=%zu", shared.size()); + decode_parts({ { prompt_part::TOKS, &shared, nullptr, 0, seq_snap } }); + DECISION_DEBUG("prepare_prefix: decode_parts returned"); cached = shared; } return false; @@ -326,12 +495,15 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve throw std::invalid_argument("a decision needs at least one context"); } const tokens_t shared = tokenize(shared_text, true); - std::vector prefixes; - for (const auto & text : contexts) { - prefixes.push_back(tokenize(text, shared.empty())); - if (prefixes.back().empty()) { + std::vector prefixes; + for (size_t i = 0; i < contexts.size(); ++i) { + const mtmd::bitmaps * ctx_bitmaps = (i < opt.context_bitmaps.size()) ? &opt.context_bitmaps[i] : nullptr; + auto mtoks = tokenize_mm(contexts[i], shared.empty(), ctx_bitmaps); + DECISION_DEBUG("decide_batch: tokenize_mm returned toks=%zu chunks=%zu", mtoks.toks.size(), mtoks.chunks.size()); + if (mtoks.toks.empty()) { throw std::invalid_argument("the decision context must not be empty"); } + prefixes.push_back(std::move(mtoks)); } std::vector fields; @@ -406,6 +578,16 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve const size_t n_group = std::min(per_group, contexts.size() - g0); const auto tp = std::chrono::steady_clock::now(); + // We need to keep segment token vectors alive while parts reference them + // Reserve enough capacity to prevent reallocation (which would invalidate pointers) + // With N image chunks interleaved with text, there can be up to N+1 text segments. + // Use a generous reserve based on the max chunks in any prefix. + std::vector seg_storage; + size_t max_chunks = 0; + for (const auto & mtoks : prefixes) { + max_chunks = std::max(max_chunks, mtoks.chunks.size()); + } + seg_storage.reserve((max_chunks + 1) * n_group); // at most (chunks+1) text segments per context std::vector parts; for (size_t i = 0; i < n_group; ++i) { const llama_seq_id trunk = seq_pool + (llama_seq_id) i; @@ -413,9 +595,44 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve if (!shared.empty()) { llama_memory_seq_cp(mem, seq_snap, trunk, -1, -1); } - parts.push_back({ &prefixes[g0 + i], (llama_pos) shared.size(), trunk }); + // Build prompt_parts from multimodal_tokens: text tokens become TOKS parts, + // image chunks become CHUNK parts with proper position tracking + const auto & mtoks = prefixes[g0 + i]; + llama_pos pos = (llama_pos) shared.size(); + size_t chunk_idx = 0; + size_t seg_start = 0; + for (size_t t_idx = 0; t_idx < mtoks.toks.size(); ++t_idx) { + if (mtoks.toks[t_idx] == LLAMA_TOKEN_NULL) { + // Flush preceding text tokens as a TOKS part + if (t_idx > seg_start) { + seg_storage.push_back(tokens_t(mtoks.toks.begin() + seg_start, mtoks.toks.begin() + t_idx)); + parts.push_back({ prompt_part::TOKS, &seg_storage.back(), nullptr, pos, trunk }); + pos += (llama_pos) seg_storage.back().size(); + seg_start = t_idx + 1; + } + // Add image chunk if available + if (chunk_idx < mtoks.chunks.size()) { + const size_t n_img_tokens = mtmd_input_chunk_get_n_tokens(mtoks.chunks[chunk_idx]); + parts.push_back({ prompt_part::CHUNK, nullptr, mtoks.chunks[chunk_idx], pos, trunk }); + pos += (llama_pos) n_img_tokens; + chunk_idx++; + seg_start = t_idx + 1; + } else { + // No image chunk is left for this NULL token, so skip it + // (it's a placeholder that has no corresponding chunk) + seg_start = t_idx + 1; + } + } + } + // Flush remaining text tokens + if (seg_start < mtoks.toks.size()) { + seg_storage.push_back(tokens_t(mtoks.toks.begin() + seg_start, mtoks.toks.end())); + parts.push_back({ prompt_part::TOKS, &seg_storage.back(), nullptr, pos, trunk }); + } } + DECISION_DEBUG("decide_batch: calling decode_parts with %zu parts", parts.size()); decode_parts(parts); + DECISION_DEBUG("decide_batch: decode_parts returned"); llama_synchronize(ctx); // llama_decode is asynchronous: wait for the prefill so its time isn't billed to scoring out.prefill_ms += ms_since(tp); @@ -429,7 +646,7 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve std::vector> owner; // (context in group, field) for (size_t i = 0; i < n_group; ++i) { const llama_seq_id trunk = seq_pool + (llama_seq_id) i; - const llama_pos pos0 = (llama_pos) (shared.size() + prefixes[g0 + i].size()); + const llama_pos pos0 = (llama_pos) (shared.size() + prefixes[g0 + i].toks.size()); for (size_t f = 0; f < state[i].size(); ++f) { auto & fd = state[i][f]; if (fd.use_tree) { @@ -487,7 +704,7 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve for (size_t i = 0; i < n_group; ++i) { llama_memory_seq_rm(mem, seq_pool + (llama_seq_id) i, -1, -1); result & r = out.items[g0 + i]; - r.context_tokens = prefixes[g0 + i].size(); + r.context_tokens = prefixes[g0 + i].toks.size(); r.rows = total; for (auto & fd : state[i]) { if (fd.use_tree && fd.probs.empty()) { diff --git a/tools/parallel-decision/decision-engine.h b/tools/parallel-decision/decision-engine.h index 49e7f104fb44..40ddfc6dbf54 100644 --- a/tools/parallel-decision/decision-engine.h +++ b/tools/parallel-decision/decision-engine.h @@ -11,6 +11,8 @@ // larger fields walk the trie greedily. #include "llama.h" +#include "mtmd.h" +#include "mtmd-helper.h" #include "json.h" #include @@ -34,6 +36,10 @@ struct options { size_t tree_max = 128; bool split_boundary = false; // legacy: tokenise suffix and values separately bool allow_cache = true; // reuse the cached static prefix when it matches + // Optional per-context bitmaps for multimodal decision. If non-empty, must have the same + // size as the contexts vector in decide_batch. Each entry holds bitmaps to prepend to + // that context (the context text should contain media markers at the corresponding positions). + std::vector context_bitmaps; }; struct field_result { @@ -72,7 +78,7 @@ struct batch_result { // flight, then branches. The context needs a unified KV cache so branches share the trunk's cells. class engine { public: - engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs); + engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs, mtmd_context * mctx = nullptr); result decide(const std::string & shared_text, const std::string & context_text, const std::vector & fields, const options & opt); @@ -83,8 +89,11 @@ class engine { const std::vector & fields, const options & opt); private: + // A prompt part can be either a list of text tokens or a media chunk (image/audio). struct prompt_part { - const tokens_t * toks; + enum type { TOKS, CHUNK } kind; + const tokens_t * toks = nullptr; // when kind == TOKS + const mtmd_input_chunk * chunk = nullptr; // when kind == CHUNK llama_pos pos0; llama_seq_id seq; }; @@ -102,7 +111,40 @@ class engine { int n_pool; bool pad_branches; // recurrent/hybrid model: branches in a decode need equal lengths tokens_t cached; + mtmd_context * mctx; // optional multimodal context for vision input + + // Tokenize text that may contain media markers, expanding them into chunks via mtmd. + // Returns text tokens with LLAMA_TOKEN_NULL at image positions, and the image chunks + // that need to be encoded separately in decode_parts. + struct multimodal_tokens { + tokens_t toks; // text tokens, LLAMA_TOKEN_NULL at image positions + std::vector chunks; // image/audio chunks to decode at those positions + + ~multimodal_tokens() { + for (const auto * chunk : chunks) { + mtmd_input_chunk_free(const_cast(chunk)); + } + } + multimodal_tokens() = default; + multimodal_tokens(multimodal_tokens && other) noexcept + : toks(std::move(other.toks)), chunks(std::move(other.chunks)) {} + multimodal_tokens & operator=(multimodal_tokens && other) noexcept { + if (this != &other) { + for (const auto * chunk : chunks) { + mtmd_input_chunk_free(const_cast(chunk)); + } + toks = std::move(other.toks); + chunks = std::move(other.chunks); + } + return *this; + } + // non-copyable (chunks are owned) + multimodal_tokens(const multimodal_tokens &) = delete; + multimodal_tokens & operator=(const multimodal_tokens &) = delete; + }; + multimodal_tokens tokenize_mm(const std::string & text, bool add_special, + const mtmd::bitmaps * bitmaps = nullptr) const; tokens_t tokenize(const std::string & text, bool add_special) const; void decode_parts(const std::vector & parts); bool prepare_prefix(const tokens_t & shared, bool allow_cache); diff --git a/tools/parallel-decision/examples/vision-decision-harness/.gitignore b/tools/parallel-decision/examples/vision-decision-harness/.gitignore new file mode 100644 index 000000000000..ebbebb5db2a6 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/.gitignore @@ -0,0 +1,7 @@ +__pycache__/ +*.pyc +# written by the test runners on each run +test_results.json +evaluation_results.json +original_jev_script.txt +tests/scan_for_secrets.py diff --git a/tools/parallel-decision/examples/vision-decision-harness/README.md b/tools/parallel-decision/examples/vision-decision-harness/README.md new file mode 100644 index 000000000000..261e2fa57618 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/README.md @@ -0,0 +1,152 @@ +# Vision Decision Harness + +A web-based UI for testing the multimodal `/decision` endpoint of llama-server, built on top of the [parallel-decision branch](https://github.com/thecodacus/llama.cpp) of llama.cpp. + +## Overview + +The harness provides a browser-based interface for evaluating multimodal decision-making capabilities of llama-server. It supports: + +- **Image batch processing**: Upload folders of images and run classification questions +- **Image Selection mode**: Scan a folder of images and find which one matches a question (e.g. "which image contains a rubber duck?") +- **Text file batch processing**: Run text files through the decision endpoint for scoring/classification evaluation +- **Dynamic decision tags**: Add/remove decision options with a + button (tag-list style) +- **Prompt autocompletion**: Get suggestions from the llama-server `/completion` endpoint +- **Benchmarking**: Per-file timing, total batch timing, progress bar +- **Streaming results**: Results appear as each file completes (SSE) +- **Image preview toggle**: Show/hide thumbnails in results +- **Export snippet**: Generate a Python API request snippet with docstring +- **Secret scanning**: Automated scan for sensitive information before committing + +## Prerequisites + +1. **llama-server with decision + multimodal support** — built from the `multimodal-decision` branch of the [thecodacus/llama.cpp](https://github.com/thecodacus/llama.cpp) fork. The server must be started with `--decision-seqs N` (N >= 3) and `--mmproj` for vision support. + + ```bash + ./build/bin/llama-server \ + --host 0.0.0.0 \ + --model model-Q4_K_M.gguf \ + --mmproj mmproj-model.gguf \ + --decision-seqs 8 \ + --device Vulkan1 \ + --port 8081 + ``` + +2. **Python 3** with Flask and requests: + ```bash + pip install flask flask-cors requests + ``` + +## Running + +```bash +cd vision-decision-harness +python3 app.py +``` + +The UI is available at `http://0.0.0.0:5786`. + +The server URL defaults to `http://0.0.0.0:8081`. Override with: +```bash +LLAMA_SERVER_URL=http://localhost:8081 python3 app.py +``` + +## API Endpoints + +| Endpoint | Method | Description | +|---|---|---| +| `/` | GET | Main UI | +| `/api/health` | GET | Health check (proxies to server) | +| `/api/decision` | POST | Proxy to llama-server `/decision` | +| `/api/completion` | POST | Proxy to llama-server `/completion` | +| `/api/upload` | POST | Upload files/folders from browser | +| `/api/batch` | POST | Batch process files (non-streaming) | +| `/api/batch-stream` | POST | Batch process files with SSE streaming | +| `/api/image-selection/stream` | POST | Image selection mode with SSE streaming | +| `/api/export-snippet` | POST | Generate Python API request snippet | +| `/api/tag-suggestions` | POST | Get tag suggestions from `/completion` | + +## Image Selection Mode + +The image selection mode sends all images in a folder to the `/decision` endpoint in a single batched request with a yes/no schema. Each image is evaluated as a separate context, and results are aggregated to identify the single best match. + +**Request:** +```json +{ + "source": "/path/to/image/folder", + "question": "Which image contains a yellow rubber duck?", + "instructions": "Look at this image and determine if it contains the described object. Answer YES if it does, NO if it does not.", + "seed": 42 +} +``` + +**SSE Events:** +- `start`: Found N images, question +- `progress`: Per-image file loaded +- `result`: Per-image yes/no evaluation +- `decision`: Single winning image with file name and probability +- `done`: Total time + +## Text Test Suite + +The `tests/` directory contains a comprehensive text-based decision scoring test suite: + +- `text_samples/` — 20 text files with sentiment classification questions +- `answer_key.json` — Expected results for all test files +- `schema.json` — Shared JSON Schema for the tests +- `evaluate_text_tests.py` — Runner that sends each file through `/decision` and compares to the answer key +- `test_decision_scoring.py` — 41 built-in decision scoring tests (text-only) + +Run with: +```bash +python3 tests/evaluate_text_tests.py +python3 tests/test_decision_scoring.py +``` + +## Secret Scanning + +Before committing, scan all harness files and modified llama.cpp files for sensitive information: + +```bash +python3 tests/scan_for_secrets.py +``` + +This script uses the llama-server `/decision` endpoint to classify each file as containing or not containing: +- Passwords +- API keys +- Email addresses +- Phone numbers +- Home directory paths +- IP addresses +- Credentials/tokens + +Any files flagged as sensitive are reported for manual review before committing. + +## Multimodal Decision Support + +### Server-side Changes + +The fork adds the following to `tools/server/server-context.cpp`: + +1. **Multi-image contexts**: The `images` field can be an array of arrays — one sub-array per context, enabling multiple images per decision context. + +2. **Media marker handling**: Context text can contain inline media markers (fetched from `/props`) to position images at specific points in the text. If no markers are present, they are prepended for backward compatibility. + +3. **Batch vision encoding path**: Uses `mtmd_batch_init` → `mtmd_batch_add_chunk` → `mtmd_batch_encode` → `mtmd_batch_get_output_embd` → `mtmd_helper_decode_image_chunk` for vision encoding, matching the server's working completion path. Falls back to per-chene encoding for models that don't support batch encoding (e.g., SmolVLM2). + +### Decision Engine Changes + +- `decision-engine.h/cpp` — Added `multimodal_tokens` struct, `tokenize_mm()` method, `options::context_bitmaps` field, and batch vision encoding in `decode_parts()` +- `tools/mtmd/mtmd.h` — Added explicit move constructors for `bitmaps` and `bitmap` to support vector storage +- `tools/parallel-decision/CMakeLists.txt` — Links `mtmd` target + +### Debug Toggle + +All debug prints in the decision engine are controlled by the `LLAMA_DECISION_DEBUG` environment variable: + +```bash +LLAMA_DECISION_DEBUG=1 ./build/bin/llama-server --model ... +``` + +## License + +This harness is provided as an example for the parallel-decision multimodal extensions. See the main [llama.cpp README](https://github.com/ggml-org/llama.cpp) for the underlying project license. diff --git a/tools/parallel-decision/examples/vision-decision-harness/app.py b/tools/parallel-decision/examples/vision-decision-harness/app.py new file mode 100644 index 000000000000..3c435eab1803 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/app.py @@ -0,0 +1,507 @@ +#!/usr/bin/env python3 +""" +Vision Decision Harness + +A web UI for testing the multimodal /decision endpoint of llama-server. + +Features: + - Select an image or a folder containing images for batch processing + - Also supports text files in batch (processes .txt and .md files) + - Add decision options dynamically with a + button (tag-list style) + - Prompt autocompletion: type a prompt, and suggestions are fetched + from the llama-server /completion endpoint as you type + - When adding a tag, the harness sends the current prompt to /completion + and uses the model's output to suggest tag names + - Benchmarking: per-file timing, total batch timing, progress bar + - Image Selection mode: scan a folder of images and find which ONE + matches a question (e.g. "which image contains a rubber duck?"). + All images are processed in a SINGLE batched /decision request as + parallel contexts with a yes/no schema. The winning image is + selected from the yes responses, giving ONE result with ONE + probability. + +Runs on 0.0.0.0:5786. Proxies to llama-server at 0.0.0.0:8081 +(override with LLAMA_SERVER_URL env var). +""" + +from flask import Flask, render_template, request, jsonify, Response +from flask_cors import CORS +import requests +import os +import base64 +import json +import time +import uuid +import string + +app = Flask(__name__, static_folder="static", template_folder="templates") +CORS(app) + +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") + +# Temporary upload directory for browser-side file/folder uploads +UPLOAD_DIR = os.environ.get("UPLOAD_DIR", "/tmp/vision-harness-uploads") +os.makedirs(UPLOAD_DIR, exist_ok=True) + + +@app.route("/") +def index(): + return render_template("index.html") + + +@app.route("/api/health") +def health(): + try: + resp = requests.get(f"{SERVER_URL}/v1/models", timeout=5) + return jsonify({"status": "ok", "server": resp.status_code}), 200 + except Exception as e: + return jsonify({"status": "error", "server": str(e)}), 502 + + +@app.route("/api/decision", methods=["POST"]) +def decision(): + data = request.get_json(force=True, silent=True) or {} + try: + resp = requests.post(f"{SERVER_URL}/decision", json=data, timeout=120) + return jsonify(resp.json()), resp.status_code + except Exception as e: + return jsonify({"error": str(e)}), 502 + + +@app.route("/api/completion", methods=["POST"]) +def completion(): + data = request.get_json(force=True, silent=True) or {} + payload = { + "prompt": data.get("prompt", ""), + "n_predict": data.get("n_predict", 16), + "temperature": data.get("temperature", 0.2), + "top_p": data.get("top_p", 0.9), + "seed": data.get("seed", 42), + "stream": False, + } + try: + resp = requests.post(f"{SERVER_URL}/completion", json=payload, timeout=60) + return jsonify(resp.json()), resp.status_code + except Exception as e: + return jsonify({"error": str(e)}), 502 + + +@app.route("/api/upload", methods=["POST"]) +def upload_files(): + """Upload files from the browser (supports folder uploads via webkitdirectory).""" + if "files" not in request.files: + return jsonify({"error": "No files provided"}), 400 + + upload_id = str(uuid.uuid4()) + upload_dir = os.path.join(UPLOAD_DIR, upload_id) + os.makedirs(upload_dir, exist_ok=True) + + uploaded = [] + for storage in request.files.getlist("files"): + rel_path = storage.filename + if not rel_path: + continue + filename = os.path.basename(rel_path) + + # Create subdirectories if this is a folder upload + rel_dir = os.path.dirname(rel_path) + dest_dir = os.path.join(upload_dir, rel_dir) + os.makedirs(dest_dir, exist_ok=True) + dest = os.path.join(dest_dir, filename) + storage.save(dest) + uploaded.append({"path": dest, "name": filename}) + + return jsonify({"upload_id": upload_id, "dir": upload_dir, "files": uploaded, "count": len(uploaded)}) + + +@app.route("/api/batch", methods=["POST"]) +def batch(): + """Batch process all supported files in a folder (or a single file).""" + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + schema = data.get("schema", {}) + instructions = data.get("instructions", "") + contexts = data.get("contexts", [""]) + seed = data.get("seed", 42) + + # Gather files + files = [] + if os.path.isfile(source): + files = [(source, os.path.basename(source))] + elif os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif", ".txt", ".md")): + files.append((fpath, fname)) + + if not files: + return jsonify({"error": "No supported files found", "results": [], "total_files": 0, "total_ms": 0}) + + results = [] + t_start = time.time() + for idx, (fpath, fname) in enumerate(files): + entry = {"file": fname, "index": idx, "total": len(files)} + is_image = fname.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")) + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": seed, + } + if is_image: + try: + with open(fpath, "rb") as f: + img_b64 = base64.b64encode(f.read()).decode() + body["images"] = [img_b64] + except Exception as e: + entry["error"] = f"Failed to read image: {e}" + results.append(entry) + continue + else: + # Text file: use file content as the context + try: + with open(fpath, "r") as f: + body["contexts"] = [f.read()] + except Exception as e: + entry["error"] = f"Failed to read text file: {e}" + results.append(entry) + continue + + t0 = time.time() + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + elapsed = time.time() - t0 + entry["elapsed_ms"] = round(elapsed * 1000, 1) + if resp.status_code == 200: + entry["result"] = resp.json() + else: + entry["error"] = f"HTTP {resp.status_code}: {resp.text[:200]}" + except Exception as e: + entry["error"] = str(e) + results.append(entry) + + total_ms = round((time.time() - t_start) * 1000, 1) + return jsonify({"results": results, "total_files": len(files), "total_ms": total_ms}) + + +@app.route("/api/batch-stream", methods=["POST"]) +def batch_stream(): + """Batch process files and stream results as Server-Sent Events.""" + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + schema = data.get("schema", {}) + instructions = data.get("instructions", "") + contexts = data.get("contexts", [""]) + seed = data.get("seed", 42) + + # Gather files (same logic as /api/batch) + files = [] + if os.path.isfile(source): + files = [(source, os.path.basename(source))] + elif os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif", ".txt", ".md")): + files.append((fpath, fname)) + + def event_stream(): + if not files: + yield "data: " + json.dumps({"error": "No supported files found", "total_files": 0}) + "\n\n" + return + + t_start = time.time() + yield "data: " + json.dumps({"event": "start", "total_files": len(files), "files": [f[1] for f in files]}) + "\n\n" + + for idx, (fpath, fname) in enumerate(files): + t0 = time.time() + entry = {"file": fname, "index": idx, "total": len(files)} + is_image = fname.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")) + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": seed, + } + if is_image: + try: + with open(fpath, "rb") as f: + img_b64 = base64.b64encode(f.read()).decode() + body["images"] = [img_b64] + except Exception as e: + entry["error"] = f"Failed to read image: {e}" + entry["elapsed_ms"] = 0 + yield "data: " + json.dumps(entry) + "\n\n" + continue + else: + try: + with open(fpath, "r") as f: + body["contexts"] = [f.read()] + except Exception as e: + entry["error"] = f"Failed to read text file: {e}" + entry["elapsed_ms"] = 0 + yield "data: " + json.dumps(entry) + "\n\n" + continue + + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + elapsed = time.time() - t0 + entry["elapsed_ms"] = round(elapsed * 1000, 1) + if resp.status_code == 200: + entry["result"] = resp.json() + entry["status"] = "ok" + else: + entry["error"] = f"HTTP {resp.status_code}: {resp.text[:200]}" + entry["status"] = "error" + except Exception as e: + entry["error"] = str(e) + entry["elapsed_ms"] = round((time.time() - t0) * 1000, 1) + entry["status"] = "error" + + # Include image data for display if this is an image file + if is_image: + entry["is_image"] = True + with open(fpath, "rb") as f: + entry["image_data"] = base64.b64encode(f.read()).decode() + else: + entry["is_image"] = False + + yield "data: " + json.dumps(entry) + "\n\n" + + total_ms = round((time.time() - t_start) * 1000, 1) + yield "data: " + json.dumps({"event": "done", "total_files": len(files), "total_ms": total_ms}) + "\n\n" + + return Response(event_stream(), mimetype="text/event-stream") + + +@app.route("/api/image-selection/stream", methods=["POST"]) +def image_selection_stream(): + """Scan a folder of images and select which ONE image matches a question. + + Sends ALL images in a SINGLE /decision request as parallel contexts + (one context per image). The schema uses binary yes/no choices, so the + model evaluates each image against the question. Results are aggregated + to identify the single best match — only the winning image gets a + probability shown, satisfying the constraint that the model picks ONE + image from all options presented together. + + This approach is used because Qwen2.5-VL's vision architecture cannot + reliably associate text labels with specific images when multiple images + are interleaved in a single multimodal context. The yes/no approach + evaluates each image in its own dedicated context while all images are + still processed together in one batched decision pass. + + SSE events: + start: total_images, image_files, question + progress: per-image file loaded + result: per-image processing update + decision: single winning image with file name and probability + done: total_ms + """ + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + question = data.get("question", "Which image contains a yellow rubber duck?") + instructions = data.get("instructions", + "Look at this image and determine if it contains the described object. Answer YES if it does, NO if it does not.") + seed = data.get("seed", 42) + + # Gather image files (only images, sorted) + files = [] + if os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + if not os.path.isfile(fpath): + continue + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")): + files.append((fpath, fname)) + elif os.path.isfile(source) and source.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")): + files = [(source, os.path.basename(source))] + + def event_stream(): + if not files: + yield "data: " + json.dumps({"error": "No image files found", "total_images": 0}) + "\n\n" + return + + image_names = [fname for (_, fname) in files] + + # All images processed in a single decision request with yes/no schema. + # Each image is a separate context so the model can clearly associate + # the question with that one image. + schema = {"properties": {"match": {"type": "string", "enum": ["yes", "no"]}}} + contexts = [question for _ in files] + + images_b64 = [] + for idx, (fpath, fname) in enumerate(files): + try: + with open(fpath, "rb") as f: + images_b64.append(base64.b64encode(f.read()).decode()) + yield "data: " + json.dumps({"event": "progress", "file": fname, "loaded": True, "index": idx, "total": len(files)}) + "\n\n" + except Exception as e: + yield "data: " + json.dumps({"error": f"Failed to read {fname}: {e}", "file": fname}) + "\n\n" + return + + yield "data: " + json.dumps({ + "event": "start", + "total_images": len(files), + "image_files": image_names, + "question": question + }) + "\n\n" + + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "images": images_b64, + "seed": seed, + } + + t0 = time.time() + elapsed = 0.0 + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=300) + elapsed = time.time() - t0 + + if resp.status_code == 200: + data = resp.json() + results = data.get("results", []) + + # Aggregate: find images with "yes" answer, sorted by confidence + yes_matches = [] + for i, res in enumerate(results): + field = res.get("fields", {}).get("match", {}) + value = field.get("value", "no") + probability = field.get("probability", 0.0) + tokens = res.get("usage", {}).get("context_tokens", 0) + + entry = { + "file": image_names[i], + "match": value == "yes", + "probability": probability, + "value": value, + "context_tokens": tokens, + "index": i, + "total": len(results), + "is_image": True, + } + yield "data: " + json.dumps({"event": "result", "data": entry}) + "\n\n" + + if value == "yes": + yes_matches.append((i, probability, image_names[i])) + + # Select the single best match (highest yes confidence) + # Only one image gets reported as the decision winner + if yes_matches: + yes_matches.sort(key=lambda x: -x[1]) # Sort by probability descending + best = yes_matches[0] + idx, prob, fname = best + + decision = { + "selected_file": fname, + "selected_index": idx, + "probability": prob, + "question": question, + "total_matches": len(yes_matches), + "all_matches": [{"file": f, "probability": p} for (_, p, f) in yes_matches], + "image_data": images_b64[idx] if 0 <= idx < len(images_b64) else None, + "is_image": True, + "timings": data.get("timings", {}), + } + yield "data: " + json.dumps({"event": "decision", "data": decision}) + "\n\n" + else: + # No matches found — report the highest "no" confidence as near-miss + all_probs = [] + for i, res in enumerate(results): + field = res.get("fields", {}).get("match", {}) + prob = field.get("probability", 0.0) + all_probs.append((i, prob, image_names[i])) + + # Find the one with highest "yes" probability even if it answered "no" + # This gives useful info about confidence + decision = { + "selected_file": None, + "selected_index": -1, + "probability": 0.0, + "question": question, + "total_matches": 0, + "timings": data.get("timings", {}), + } + yield "data: " + json.dumps({"event": "decision", "data": decision}) + "\n\n" + + else: + yield "data: " + json.dumps({"error": f"HTTP {resp.status_code}: {resp.text[:300]}"}) + "\n\n" + except Exception as e: + yield "data: " + json.dumps({"error": str(e)}) + "\n\n" + + yield "data: " + json.dumps({"event": "done", "total_ms": round(elapsed * 1000, 1)}) + "\n\n" + + return Response(event_stream(), mimetype="text/event-stream") + + +@app.route("/api/export-snippet", methods=["POST"]) +def export_snippet(): + data = request.get_json(force=True, silent=True) or {} + schema_json = json.dumps(data.get("schema", {}), indent=2) if data.get("schema") else "{}" + contexts_json = json.dumps(data.get("contexts", []), indent=2) + instructions = data.get("instructions", "") + + snippet = '''"""Vision Decision API request snippet. + +Endpoint: POST ''' + SERVER_URL + '''/decision + +Request body: + contexts: list[str] -- one string per decision context + schema : dict -- JSON Schema with "properties" (required) where each + property defines an enum/integer/number/boolean field + images : list[str] -- optional, base64-encoded image data (one per media marker in context) + instructions : str -- appended to the system prompt + seed : int -- optional, for reproducibility + +Response format: + { + "object": "decision", + "results": [ + { + "decision": { "action": "CLICK_OK" }, + "fields": { "action": { "value": "CLICK_OK", "probability": 0.99, "scored_nodes": 2, "tree": true }}, + "usage": { "context_tokens": 18, "scored_rows": 11 } + } + ], + "timings": { + "prefill_ms": 340.5, + "scoring_ms": 12.3, + "total_ms": 352.8, + "rounds": 1 + } + } +""" + +import base64, requests, json + +schema = ''' + schema_json + ''' +contexts = ''' + contexts_json + ''' +instructions = ''' + json.dumps(instructions) + ''' + +payload = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": 42, +} + +# Optional: add images to payload if needed: +# with open("path/to/image.jpg", "rb") as f: +# img_b64 = base64.b64encode(f.read()).decode() +# payload["images"] = [img_b64] + +resp = requests.post("''' + SERVER_URL + '''/decision", json=payload, timeout=120) +print(json.dumps(resp.json(), indent=2)) +''' + return jsonify({"snippet": snippet}) + + +if __name__ == "__main__": + port = int(os.environ.get("PORT", 5786)) + app.run(host="0.0.0.0", port=port, debug=False) + diff --git a/tools/parallel-decision/examples/vision-decision-harness/requirements.txt b/tools/parallel-decision/examples/vision-decision-harness/requirements.txt new file mode 100644 index 000000000000..690a3cefbe89 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/requirements.txt @@ -0,0 +1,3 @@ +flask +flask-cors +requests diff --git a/tools/parallel-decision/examples/vision-decision-harness/templates/index.html b/tools/parallel-decision/examples/vision-decision-harness/templates/index.html new file mode 100644 index 000000000000..be2256c9b6da --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/templates/index.html @@ -0,0 +1,749 @@ + + + + + + Vision Decision Harness + + + +
+

Vision Decision Harness

+
Harness on 0.0.0.0:5786 | Proxies to llama-server at 0.0.0.0:8081
+ +
+ +
+
+

Source

+
+ +
+ + +
+ + Drag-drop or browse a folder of images (.jpg/.png/etc) and text files (.txt/.md) +
+
+ + +
+
+ +
+

Prompt & Tags

+
+ + +
+
+
+ + +
+
+ + +
+
+ + +
+
+ +
+ + + +
+
+
+
+
+ + +
+
+ +
+

Actions

+
+ +
+
+ +
+
+ +
+
+ +
+
+ +
+
+
+ + +
+
+

Results

+
+
+ + + + + + + +
FileImageDecisionProbTokensTime (ms)Status
+
+ +
+

Log

+
+
+
+
+
+ + + + diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json b/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json new file mode 100644 index 000000000000..f002a3165b36 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json @@ -0,0 +1,190 @@ +{ + "description": "Answer key for decision scoring test suite", + "total_tests": 23, + "tests": [ + { + "id": "t01_basic_color", + "context": "What is the dominant color in this image?", + "expected": "brown", + "difficulty": "easy", + "category": "color", + "schema_enum": ["red", "green", "blue", "brown", "orange"] + }, + { + "id": "t02_animal_type", + "context": "What type of animal is in this image?", + "expected": "dog", + "difficulty": "medium", + "category": "animal", + "schema_enum": ["dog", "cat", "bird", "horse", "fish"] + }, + { + "id": "t03_ui_element", + "context": "What type of UI element should you click to submit this form?", + "expected": "button", + "difficulty": "easy", + "category": "ui", + "schema_enum": ["button", "link", "checkbox", "text_field", "dropdown"] + }, + { + "id": "t04_action_select", + "context": "Based on this content, what is the best action to take?", + "expected": "CLICK_OK", + "difficulty": "easy", + "category": "action", + "schema_enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"] + }, + { + "id": "t05_sentiment", + "context": "What is the sentiment expressed in this content?", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "schema_enum": ["positive", "negative", "neutral"] + }, + { + "id": "t06_navigation", + "context": "Where should you navigate to complete this task?", + "expected": "settings", + "difficulty": "hard", + "category": "navigation", + "schema_enum": ["home", "settings", "profile", "logout", "help"] + }, + { + "id": "t07_content_type", + "context": "What type of content is shown in this image?", + "expected": "image", + "difficulty": "easy", + "category": "content", + "schema_enum": ["text", "image", "video", "audio", "interactive"] + }, + { + "id": "t08_priority", + "context": "What priority level should this task be assigned?", + "expected": "medium", + "difficulty": "hard", + "category": "priority", + "schema_enum": ["low", "medium", "high", "urgent"] + }, + { + "id": "t09_safety", + "context": "Is it safe to proceed with this action?", + "expected": "yes", + "difficulty": "medium", + "category": "safety", + "schema_enum": ["yes", "no", "caution"] + }, + { + "id": "t10_layout", + "context": "Where should new items be added in this layout?", + "expected": "bottom_right", + "difficulty": "hard", + "category": "layout", + "schema_enum": ["top_left", "top_right", "bottom_left", "bottom_right", "center"] + }, + { + "id": "t11_text_format", + "context": "What format is this text in?", + "expected": "plain", + "difficulty": "medium", + "category": "format", + "schema_enum": ["plain", "markdown", "html", "json", "xml"] + }, + { + "id": "t12_object_count", + "context": "How many main objects are visible in this image?", + "expected": "one", + "difficulty": "medium", + "category": "counting", + "schema_enum": ["one", "two", "three", "four", "many"] + }, + { + "id": "t13_file_op", + "context": "What file operation should be performed?", + "expected": "save", + "difficulty": "easy", + "category": "fileop", + "schema_enum": ["save", "delete", "rename", "copy", "move"] + }, + { + "id": "t14_error_type", + "context": "What type of error is indicated by this message?", + "expected": "not_found", + "difficulty": "hard", + "category": "error", + "schema_enum": ["syntax", "runtime", "network", "permission", "not_found"] + }, + { + "id": "t15_time_urgency", + "context": "How urgent is this action?", + "expected": "soon", + "difficulty": "hard", + "category": "time", + "schema_enum": ["immediate", "soon", "later", "anytime"] + }, + { + "id": "t16_language", + "context": "What language is this text written in?", + "expected": "english", + "difficulty": "hard", + "category": "language", + "schema_enum": ["english", "spanish", "french", "german", "chinese"] + }, + { + "id": "t17_component", + "context": "What type of UI component is shown here?", + "expected": "card", + "difficulty": "medium", + "category": "ui", + "schema_enum": ["modal", "sidebar", "header", "footer", "card"] + }, + { + "id": "t18_data_source", + "context": "Where should this data be fetched from?", + "expected": "api", + "difficulty": "hard", + "category": "data", + "schema_enum": ["database", "api", "file", "cache", "user_input"] + }, + { + "id": "t19_interaction", + "context": "What type of user interaction does this represent?", + "expected": "click", + "difficulty": "medium", + "category": "interaction", + "schema_enum": ["click", "hover", "drag", "scroll", "type"] + }, + { + "id": "t20_state", + "context": "What is the correct state for this toggle?", + "expected": "on", + "difficulty": "hard", + "category": "state", + "schema_enum": ["on", "off", "indeterminate", "disabled"] + }, + { + "id": "t21_direction", + "context": "Which direction should this carousel move?", + "expected": "right", + "difficulty": "hard", + "category": "direction", + "schema_enum": ["left", "right", "up", "down", "none"] + }, + { + "id": "t22_file_type", + "context": "What type of file is this?", + "expected": "image", + "difficulty": "easy", + "category": "filetype", + "schema_enum": ["image", "document", "spreadsheet", "presentation", "archive"] + }, + { + "id": "t23_security", + "context": "What security level is required for this resource?", + "expected": "internal", + "difficulty": "hard", + "category": "security", + "schema_enum": ["public", "internal", "confidential", "restricted"] + } + ] +} diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py b/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py new file mode 100644 index 000000000000..98ef30836fdb --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py @@ -0,0 +1,195 @@ +#!/usr/bin/env python3 +""" +Evaluation runner for the text-based decision test suite. + +Processes all .txt files in tests/text_samples/ through the llama-server +/decision endpoint using the shared sentiment schema, then compares results +against answer_key.json. + +Usage: + python3 evaluate_text_tests.py + +Environment: + LLAMA_SERVER_URL - defaults to http://0.0.0.0:8081 +""" + +import json +import os +import sys +import time +import requests + +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") +SAMPLES_DIR = os.path.join(os.path.dirname(__file__), "text_samples") + + +def load_answer_key(): + path = os.path.join(SAMPLES_DIR, "answer_key.json") + with open(path) as f: + return json.load(f) + + +def run_single_decision(context_text, schema, instructions, seed=42): + """Send a single decision request to the llama-server.""" + body = { + "contexts": [context_text], + "schema": schema, + "instructions": instructions, + "seed": seed, + } + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + return resp + + +def main(): + key_data = load_answer_key() + schema = key_data["schema"] + question = key_data["question"] + tests = key_data["tests"] + expected_field = list(schema["properties"].keys())[0] # "sentiment" + + print(f"Text Decision Test Suite") + print(f"Server: {SERVER_URL}") + print(f"Question: {question}") + print(f"Schema: {json.dumps(schema['properties'])}") + print(f"Tests: {len(tests)}") + print() + + # Check server is reachable + try: + requests.get(f"{SERVER_URL}/v1/models", timeout=5) + except Exception: + print(f"ERROR: Cannot reach llama-server at {SERVER_URL}") + print("The server must be running for tests to execute.") + return 1 + + instructions = "Classify the sentiment expressed in the following text." + results = [] + t_total = time.time() + + for i, test in enumerate(tests): + filepath = os.path.join(SAMPLES_DIR, test["file"]) + try: + with open(filepath, "r") as f: + text = f.read().strip() + except Exception as e: + results.append({"file": test["file"], "error": str(e), "correct": False}) + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} ERROR: {e}") + continue + + t0 = time.time() + resp = run_single_decision(text, schema, instructions) + elapsed = time.time() - t0 + + if resp.status_code == 200: + data = resp.json() + result = data.get("results", [{}])[0] + fields = result.get("fields", {}) + field_data = fields.get(expected_field, {}) + predicted = field_data.get("value", "UNKNOWN") + probability = field_data.get("probability", 0.0) + tokens = result.get("usage", {}).get("context_tokens", 0) + timings = data.get("timings", {}) + + expected = test["expected"] + correct = predicted == expected + status = "PASS" if correct else "FAIL" + prob_str = f"{probability*100:.1f}%" + + mark = "" if correct else f" [expected: {expected}]" + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} {status} -> {predicted:<10} ({prob_str}) [{elapsed*1000:.0f}ms, {tokens} toks]{mark}") + + results.append({ + "file": test["file"], + "predicted": predicted, + "expected": expected, + "correct": correct, + "probability": probability, + "elapsed_ms": round(elapsed * 1000, 1), + "context_tokens": tokens, + "difficulty": test["difficulty"], + "category": test["category"], + "timings": timings, + }) + else: + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} HTTP {resp.status_code}: {resp.text[:100]}") + results.append({ + "file": test["file"], + "error": f"HTTP {resp.status_code}", + "correct": False, + "difficulty": test["difficulty"], + "category": test["category"], + }) + + total_elapsed = time.time() - t_total + + # Print summary + print() + print("=" * 80) + print("SUMMARY") + print("=" * 80) + + correct_count = sum(1 for r in results if r.get("correct")) + total_count = len(results) + print(f"Accuracy: {correct_count}/{total_count} ({correct_count/total_count*100:.1f}%)") + print(f"Total time: {total_elapsed:.1f}s ({total_elapsed/total_count*1000:.0f}ms per test avg)") + + # By difficulty + print("\nBy Difficulty:") + for diff in ["easy", "medium", "hard"]: + diff_results = [r for r in results if r.get("difficulty") == diff] + if diff_results: + diff_correct = sum(1 for r in diff_results if r.get("correct")) + correct_probs = [r.get("probability", 0) for r in diff_results if r.get("correct")] + wrong_probs = [r.get("probability", 0) for r in diff_results if not r.get("correct") and "probability" in r] + avg_prob = sum(correct_probs) / len(correct_probs) if correct_probs else 0 + avg_wrong = sum(wrong_probs) / len(wrong_probs) if wrong_probs else 0 + print(f" {diff:<8}: {diff_correct}/{len(diff_results)} ({diff_correct/len(diff_results)*100:.0f}%) avg_prob={avg_prob:.1%} avg_wrong_prob={avg_wrong:.1%}") + + # By category + print("\nBy Category:") + categories = sorted(set(r.get("category", "") for r in results)) + for cat in categories: + cat_results = [r for r in results if r.get("category") == cat] + cat_correct = sum(1 for r in cat_results if r.get("correct")) + print(f" {cat:<14}: {cat_correct}/{len(cat_results)} ({cat_correct/len(cat_results)*100:.0f}%)") + + # Calibration check + correct_probs = [r["probability"] for r in results if r.get("correct") and "probability" in r] + wrong_probs = [r["probability"] for r in results if not r.get("correct") and "probability" in r] + if correct_probs and wrong_probs: + avg_correct = sum(correct_probs) / len(correct_probs) + avg_wrong = sum(wrong_probs) / len(wrong_probs) + print(f"\nCalibration:") + print(f" Avg confidence (correct answers): {avg_correct:.1%}") + print(f" Avg confidence (wrong answers): {avg_wrong:.1%}") + print(f" Well calibrated: {'YES' if avg_correct > avg_wrong else 'NO'}") + + # Wrong answers + wrong = [r for r in results if not r.get("correct")] + if wrong: + print(f"\nWrong answers ({len(wrong)}):") + for r in wrong: + print(f" {r['file']}: expected '{r.get('expected', '?')}', got '{r.get('predicted', '?')}' ({r.get('probability', 0):.1%})") + + # Save results + results_path = os.path.join(SAMPLES_DIR, "evaluation_results.json") + with open(results_path, "w") as f: + json.dump({ + "question": question, + "schema": schema, + "total_tests": total_count, + "passed": correct_count, + "failed": total_count - correct_count, + "accuracy": f"{correct_count/total_count*100:.1f}%", + "total_time_ms": round(total_elapsed * 1000, 1), + "avg_time_ms": round(total_elapsed / total_count * 1000, 1), + "results": results, + }, f, indent=2) + print(f"\nDetailed results saved to {results_path}") + + return 0 if correct_count == total_count else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py b/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py new file mode 100644 index 000000000000..ccc7141de4cc --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py @@ -0,0 +1,633 @@ +#!/usr/bin/env python3 +""" +Test suite for evaluating the llama-server /decision endpoint's +scoring and classification abilities. + +These are TEXT-ONLY tests — no image references. Each test presents +a text context and asks the model to classify or choose among options +based on the textual content alone. +""" + +import json +import os +import base64 +import time +import requests +from pathlib import Path + +# ---- Configuration ---- +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") +TEST_IMAGE_DIR = os.environ.get("TEST_IMAGE_DIR", + os.path.join(os.path.dirname(os.path.dirname(__file__)), "..", "llama.cpp-thecodacus", "tools", "mtmd")) + + +# ---- Test Cases (TEXT-ONLY, no image references) ---- +TESTS = [ + # 1. Sentiment classification + { + "id": "t01_sentiment_pos", + "context": "The product arrived early and exceeded all expectations. The packaging was perfect and the quality is outstanding.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + }, + # 2. Sentiment - negative + { + "id": "t02_sentiment_neg", + "context": "This service was terrible. The staff was rude, the food was cold, and I had to wait an hour. Never coming back.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + }, + # 3. Sentiment - neutral + { + "id": "t03_sentiment_neu", + "context": "The meeting is scheduled for 3 PM in conference room B. Please bring the quarterly report.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + }, + # 4. Urgency classification + { + "id": "t04_urgency_high", + "context": "The production server is down and customers cannot complete purchases. Immediate action required.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "critical", + "difficulty": "easy", + "category": "urgency", + }, + # 5. Urgency - medium + { + "id": "t05_urgency_med", + "context": "Please review the design documents and provide feedback by end of week.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "medium", + "difficulty": "medium", + "category": "urgency", + }, + # 6. Urgency - low + { + "id": "t06_urgency_low", + "context": "When you have time, could you update the project wiki with the new API endpoints?", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "low", + "difficulty": "hard", + "category": "urgency", + }, + # 7. Action selection + { + "id": "t07_action_ok", + "context": "The user has confirmed their email address and completed the registration form. All validation checks passed.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "CLICK_OK", + "difficulty": "easy", + "category": "action", + }, + # 8. Action - cancel + { + "id": "t08_action_cancel", + "context": "The payment was declined due to insufficient funds. The user needs to provide alternative payment.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "CLICK_CANCEL", + "difficulty": "medium", + "category": "action", + }, + # 9. Action - no action + { + "id": "t09_action_none", + "context": "All system checks passed. The service is running normally with no issues detected. No further action needed.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "NO_ACTION", + "difficulty": "hard", + "category": "action", + }, + # 10. Content type classification + { + "id": "t10_content_news", + "context": "Breaking: The city council voted today to approve the new budget proposal. The measure passed with a 7-2 majority after three hours of debate.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "news", + "difficulty": "easy", + "category": "content", + }, + # 11. Content type - instruction + { + "id": "t11_content_instruction", + "context": "To reset your password, first navigate to the login page. Click the 'Forgot Password' link. Enter your email address and submit.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "instruction", + "difficulty": "medium", + "category": "content", + }, + # 12. Content type - summary + { + "id": "t12_content_summary", + "context": "In summary, the experiment demonstrated a significant improvement in performance. Key findings include a 15% increase in throughput and 20% reduction in latency compared to baseline.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "summary", + "difficulty": "medium", + "category": "content", + }, + # 13. Priority classification + { + "id": "t13_priority_high", + "context": "Security vulnerability detected in production. SQL injection risk in the user authentication endpoint. Fix immediately.", + "schema": {"properties": {"priority": {"type": "string", + "enum": ["low", "medium", "high", "urgent"]}}}, + "expected": "urgent", + "difficulty": "medium", + "category": "priority", + }, + # 14. Priority - low + { + "id": "t14_priority_low", + "context": "Consider updating the style guide documentation when time permits next quarter.", + "schema": {"properties": {"priority": {"type": "string", + "enum": ["low", "medium", "high", "urgent"]}}}, + "expected": "low", + "difficulty": "hard", + "category": "priority", + }, + # 15. File operation + { + "id": "t15_file_save", + "context": "The user has finished editing the document and clicked the save button. The changes should be persisted to disk.", + "schema": {"properties": {"operation": {"type": "string", + "enum": ["save", "delete", "rename", "copy", "move"]}}}, + "expected": "save", + "difficulty": "easy", + "category": "fileop", + }, + # 16. File operation - delete + { + "id": "t16_file_delete", + "context": "The temporary cache files from last week's processing run are no longer needed and should be removed to free up space.", + "schema": {"properties": {"operation": {"type": "string", + "enum": ["save", "delete", "rename", "copy", "move"]}}}, + "expected": "delete", + "difficulty": "medium", + "category": "fileop", + }, + # 17. File type + { + "id": "t17_filetype_image", + "context": "File: photo_2024_09_25_143022.jpg — JPEG image, 1920x1080 pixels, ICC profile: sRGB, EXIF data present, camera: iPhone 14 Pro.", + "schema": {"properties": {"filetype": {"type": "string", + "enum": ["image", "document", "spreadsheet", "presentation", "archive"]}}}, + "expected": "image", + "difficulty": "easy", + "category": "filetype", + }, + # 18. File type - document + { + "id": "t18_filetype_doc", + "context": "File: quarterly_report.pdf — PDF document, 42 pages, contains tables, charts, and formatted text. Last modified: 2024-09-20.", + "schema": {"properties": {"filetype": {"type": "string", + "enum": ["image", "document", "spreadsheet", "presentation", "archive"]}}}, + "expected": "document", + "difficulty": "medium", + "category": "filetype", + }, + # 19. Error type + { + "id": "t19_error_notfound", + "context": "Error: Resource not found at /api/users/12345. The requested user ID does not exist in the database.", + "schema": {"properties": {"error": {"type": "string", + "enum": ["syntax", "runtime", "network", "permission", "not_found"]}}}, + "expected": "not_found", + "difficulty": "hard", + "category": "error", + }, + # 20. Error type - permission + { + "id": "t20_error_permission", + "context": "Access denied: User does not have permission to read /etc/shadow. Required role: root, current role: standard_user.", + "schema": {"properties": {"error": {"type": "string", + "enum": ["syntax", "runtime", "network", "permission", "not_found"]}}}, + "expected": "permission", + "difficulty": "hard", + "category": "error", + }, + # 21. UI element type + { + "id": "t21_ui_button", + "context": "The form contains a blue rectangular button labeled 'Submit Order'. Clicking it sends the form data to the server.", + "schema": {"properties": {"element": {"type": "string", + "enum": ["button", "link", "checkbox", "text_field", "dropdown"]}}}, + "expected": "button", + "difficulty": "easy", + "category": "ui", + }, + # 22. UI element - dropdown + { + "id": "t22_ui_dropdown", + "context": "The settings panel shows a dropdown menu with options: Light Mode, Dark Mode, Auto. The current selection is Dark Mode.", + "schema": {"properties": {"element": {"type": "string", + "enum": ["button", "link", "checkbox", "text_field", "dropdown"]}}}, + "expected": "dropdown", + "difficulty": "medium", + "category": "ui", + }, + # 23. Language detection + { + "id": "t23_lang_english", + "context": "The quick brown fox jumps over the lazy dog. Pack my box with five dozen liquor jugs. How vexingly quick daft zebras jump!", + "schema": {"properties": {"lang": {"type": "string", + "enum": ["english", "spanish", "french", "german", "chinese"]}}}, + "expected": "english", + "difficulty": "easy", + "category": "language", + }, + # 24. Language - Spanish + { + "id": "t24_lang_spanish", + "context": "Hola, como estas? Me encantaria visitar Espana algun dia. La comida espanola es deliciosa, especialmente la paella.", + "schema": {"properties": {"lang": {"type": "string", + "enum": ["english", "spanish", "french", "german", "chinese"]}}}, + "expected": "spanish", + "difficulty": "medium", + "category": "language", + }, + # 25. Navigation + { + "id": "t25_navigation_settings", + "context": "To change your notification preferences, go to the account section and look for the bell icon. Click it to open the settings panel.", + "schema": {"properties": {"destination": {"type": "string", + "enum": ["home", "settings", "profile", "logout", "help"]}}}, + "expected": "settings", + "difficulty": "hard", + "category": "navigation", + }, + # 26. Format detection + { + "id": "t26_format_json", + "context": '{"users": [{"name": "Alice", "id": 1}, {"name": "Bob", "id": 2}], "count": 2}', + "schema": {"properties": {"format": {"type": "string", + "enum": ["plain", "markdown", "html", "json", "xml"]}}}, + "expected": "json", + "difficulty": "easy", + "category": "format", + }, + # 27. Format - markdown + { + "id": "t27_format_markdown", + "context": "# Project README\n\nThis project does **important things**. See the [documentation](docs.md) for details.\n\n```python\nprint('hello')\n```", + "schema": {"properties": {"format": {"type": "string", + "enum": ["plain", "markdown", "html", "json", "xml"]}}}, + "expected": "markdown", + "difficulty": "medium", + "category": "format", + }, + # 28. Interaction type + { + "id": "t28_interaction_click", + "context": "The user pressed the red submit button with their mouse. The button highlighted blue briefly and then the form was submitted.", + "schema": {"properties": {"interaction": {"type": "string", + "enum": ["click", "hover", "drag", "scroll", "type"]}}}, + "expected": "click", + "difficulty": "easy", + "category": "interaction", + }, + # 29. Interaction - scroll + { + "id": "t29_interaction_scroll", + "context": "The user moved the scrollbar down using the mouse wheel to see more content below the fold of the webpage.", + "schema": {"properties": {"interaction": {"type": "string", + "enum": ["click", "hover", "drag", "scroll", "type"]}}}, + "expected": "scroll", + "difficulty": "hard", + "category": "interaction", + }, + # 30. State classification + { + "id": "t30_state_on", + "context": "The switch is in the active position. The LED indicator is lit green. Power is flowing to the connected device.", + "schema": {"properties": {"state": {"type": "string", + "enum": ["on", "off", "indeterminate", "disabled"]}}}, + "expected": "on", + "difficulty": "easy", + "category": "state", + }, + # 31. State - off + { + "id": "t31_state_off", + "context": "The device is powered down. The LED is unlit. No power is being drawn from the battery. The switch is in the inactive position.", + "schema": {"properties": {"state": {"type": "string", + "enum": ["on", "off", "indeterminate", "disabled"]}}}, + "expected": "off", + "difficulty": "medium", + "category": "state", + }, + # 32. Data source + { + "id": "t32_data_api", + "context": "The frontend fetches user data by making a GET request to /api/v2/users with authentication headers. Results are returned as JSON.", + "schema": {"properties": {"source": {"type": "string", + "enum": ["database", "api", "file", "cache", "user_input"]}}}, + "expected": "api", + "difficulty": "medium", + "category": "data", + }, + # 33. Data source - user_input + { + "id": "t33_data_user", + "context": "The search query was typed directly into the search box by the user. No predefined data source was queried.", + "schema": {"properties": {"source": {"type": "string", + "enum": ["database", "api", "file", "cache", "user_input"]}}}, + "expected": "user_input", + "difficulty": "hard", + "category": "data", + }, + # 34. Security level + { + "id": "t34_security_internal", + "context": "This document contains company-wide policies for internal use only. It is accessible to all employees but not to external parties.", + "schema": {"properties": {"level": {"type": "string", + "enum": ["public", "internal", "confidential", "restricted"]}}}, + "expected": "internal", + "difficulty": "medium", + "category": "security", + }, + # 35. Security - public + { + "id": "t35_security_public", + "context": "The marketing brochure is available on the company website for anyone to download and share freely.", + "schema": {"properties": {"level": {"type": "string", + "enum": ["public", "internal", "confidential", "restricted"]}}}, + "expected": "public", + "difficulty": "easy", + "category": "security", + }, + # 36. Time urgency + { + "id": "t36_time_soon", + "context": "Please review the draft proposal and provide feedback within the next few days. The deadline is Friday.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["immediate", "soon", "later", "anytime"]}}}, + "expected": "soon", + "difficulty": "hard", + "category": "time", + }, + # 37. Time urgency - later + { + "id": "t37_time_later", + "context": "When you have time over the next few weeks, please update the project documentation with the new API changes.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["immediate", "soon", "later", "anytime"]}}}, + "expected": "later", + "difficulty": "hard", + "category": "time", + }, + # 38. Direction + { + "id": "t38_direction_right", + "context": "The carousel should advance to the next item. The user clicked the right arrow button to move forward.", + "schema": {"properties": {"direction": {"type": "string", + "enum": ["left", "right", "up", "down", "none"]}}}, + "expected": "right", + "difficulty": "medium", + "category": "direction", + }, + # 39. Direction - left + { + "id": "t39_direction_left", + "context": "The user wants to go back to the previous slide. Click the left arrow to navigate backwards.", + "schema": {"properties": {"direction": {"type": "string", + "enum": ["left", "right", "up", "down", "none"]}}}, + "expected": "left", + "difficulty": "hard", + "category": "direction", + }, + # 40. Component type + { + "id": "t40_component_card", + "context": "Each user profile is displayed in a bordered box with a shadow. It contains the profile picture, name, and status below.", + "schema": {"properties": {"component": {"type": "string", + "enum": ["modal", "sidebar", "header", "footer", "card"]}}}, + "expected": "card", + "difficulty": "medium", + "category": "ui", + }, + # 41. Component - sidebar + { + "id": "t41_component_sidebar", + "context": "The navigation panel on the left side of the screen contains links to Dashboard, Settings, and Logout.", + "schema": {"properties": {"component": {"type": "string", + "enum": ["modal", "sidebar", "header", "footer", "card"]}}}, + "expected": "sidebar", + "difficulty": "hard", + "category": "ui", + }, +] + + +def run_single_decision(test_case, timeout=120): + """Run a single decision test and return the result dict.""" + body = { + "contexts": [test_case["context"]], + "schema": test_case["schema"], + "instructions": "Select the correct value from the allowed choices based on the context provided.", + "seed": 42, + } + + t0 = time.time() + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=timeout) + elapsed = time.time() - t0 + if resp.status_code == 200: + data = resp.json() + field_name = list(data["results"][0]["fields"].keys())[0] + fld = data["results"][0]["fields"][field_name] + return { + "success": True, + "decision": fld["value"], + "probability": fld["probability"], + "expected": test_case["expected"], + "elapsed_ms": round(elapsed * 1000), + "tokens": data["results"][0]["usage"].get("context_tokens", 0), + "timings": data.get("timings", {}), + "id": test_case["id"], + "context": test_case["context"], + "difficulty": test_case["difficulty"], + "category": test_case["category"], + } + else: + return { + "success": False, + "error": f"HTTP {resp.status_code}: {resp.text[:200]}", + "id": test_case["id"], + } + except Exception as e: + return { + "success": False, + "error": str(e), + "id": test_case["id"], + } + + +def evaluate_results(results): + """Evaluate test results and print a summary.""" + passed = 0 + failed = 0 + total_prob_correct = 0.0 + total_prob_wrong = 0.0 + correct_count = 0 + wrong_count = 0 + + print("\n" + "=" * 90) + print("DECISION SCORING TEST RESULTS (TEXT-ONLY)") + print("=" * 90) + print(f"{'ID':<18} {'Diff':<8} {'Category':<14} {'Expected':<14} {'Got':<14} {'Prob':<8} {'Time':<8} {'Status':<6}") + print("-" * 90) + + by_difficulty = {} + by_category = {} + + for r in results: + if not r["success"]: + print(f"{r['id']:<18} {'ERROR':<8} {'':<14} {'N/A':<14} {'N/A':<14} {'N/A':<8} {'N/A':<8} FAIL") + failed += 1 + continue + + is_correct = r["decision"] == r["expected"] + status = "PASS" if is_correct else "FAIL" + if is_correct: + passed += 1 + total_prob_correct += r["probability"] + correct_count += 1 + else: + failed += 1 + total_prob_wrong += r["probability"] + wrong_count += 1 + + diff = r["difficulty"] + if diff not in by_difficulty: + by_difficulty[diff] = {"passed": 0, "total": 0, "probs": []} + by_difficulty[diff]["total"] += 1 + if is_correct: + by_difficulty[diff]["passed"] += 1 + by_difficulty[diff]["probs"].append(r["probability"]) + + cat = r["category"] + if cat not in by_category: + by_category[cat] = {"passed": 0, "total": 0, "probs": []} + by_category[cat]["total"] += 1 + if is_correct: + by_category[cat]["passed"] += 1 + by_category[cat]["probs"].append(r["probability"]) + + print(f"{r['id']:<18} {diff:<8} {cat:<14} {r['expected']:<14} {r['decision']:<14} {r['probability']:.2%} {str(r['elapsed_ms'])+'ms':<8} {status}") + + print("-" * 90) + print(f"\nOverall: {passed}/{len(results)} passed ({passed/len(results)*100:.1f}%)\n") + + print("By Difficulty:") + for diff in sorted(by_difficulty.keys()): + d = by_difficulty[diff] + avg_prob = sum(d["probs"]) / len(d["probs"]) if d["probs"] else 0 + print(f" {diff:<8}: {d['passed']}/{d['total']} ({d['passed']/d['total']*100:.1f}%) avg_prob={avg_prob:.2%}") + + print("\nBy Category:") + for cat in sorted(by_category.keys()): + c = by_category[cat] + avg_prob = sum(c["probs"]) / len(c["probs"]) if c["probs"] else 0 + print(f" {cat:<14}: {c['passed']}/{c['total']} ({c['passed']/c['total']*100:.1f}%) avg_prob={avg_prob:.2%}") + + # Calibration analysis + if correct_count > 0 and wrong_count > 0: + avg_correct = total_prob_correct / correct_count + avg_wrong = total_prob_wrong / wrong_count + print(f"\nCalibration:") + print(f" Avg prob (correct): {avg_correct:.2%}") + print(f" Avg prob (wrong): {avg_wrong:.2%}") + if avg_correct > avg_wrong: + print(f" -> WELL CALIBRATED (correct answers have higher confidence)") + else: + print(f" -> POORLY CALIBRATED (wrong answers have higher confidence)") + elif correct_count > 0: + print(f"\nCalibration: All answers correct ({total_prob_correct/correct_count:.2%} avg prob)") + + # Timing summary + valid_times = [r["elapsed_ms"] for r in results if r["success"]] + if valid_times: + print(f"\nTiming:") + print(f" Avg per decision: {sum(valid_times)/len(valid_times):.0f}ms") + print(f" Min: {min(valid_times)}ms | Max: {max(valid_times)}ms") + total_time = sum(valid_times) + print(f" Total: {total_time}ms ({total_time/1000:.1f}s)") + + return passed, failed + + +def save_answer_key(): + """Save the answer key separately from the test code.""" + key = { + "description": "Answer key for decision scoring test suite (text-only)", + "total_tests": len(TESTS), + "tests": [ + { + "id": t["id"], + "context": t["context"][:80] + "..." if len(t["context"]) > 80 else t["context"], + "expected": t["expected"], + "difficulty": t["difficulty"], + "category": t["category"], + "schema_enum": list(t["schema"]["properties"].values())[0]["enum"], + } + for t in TESTS + ], + } + + key_path = os.path.join(os.path.dirname(__file__), "answer_key.json") + with open(key_path, "w") as f: + json.dump(key, f, indent=2) + print(f"Answer key saved to {key_path}") + return key_path + + +if __name__ == "__main__": + import sys + + save_key = "--save-key" in sys.argv + + if save_key: + save_answer_key() + print("Answer key saved. Run without --save-key to test.") + sys.exit(0) + + print(f"Running {len(TESTS)} decision scoring tests (text-only)...") + print(f"Server: {SERVER_URL}") + + results = [] + for i, test in enumerate(TESTS): + r = run_single_decision(test) + results.append(r) + if r["success"]: + status = "PASS" if r["decision"] == r["expected"] else "FAIL" + print(f"[{i+1:2d}/{len(TESTS)}] {test['id']}: {status} -> {r['decision']} ({r['probability']:.1%}) [{r['elapsed_ms']}ms]") + else: + print(f"[{i+1:2d}/{len(TESTS)}] {test['id']}: ERROR: {r.get('error', 'unknown')}") + + passed, failed = evaluate_results(results) + + # Save results + results_path = os.path.join(os.path.dirname(__file__), "test_results.json") + with open(results_path, "w") as f: + json.dump(results, f, indent=2) + print(f"\nDetailed results saved to {results_path}") + + sys.exit(0 if failed == 0 else 1) diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json new file mode 100644 index 000000000000..83409546fd77 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json @@ -0,0 +1,159 @@ +{ + "description": "Answer key for text-based sentiment decision test suite. All 20 files use the same schema and question for consistent batch evaluation.", + "question": "What is the sentiment expressed in the following text?", + "schema": { + "properties": { + "sentiment": { + "type": "string", + "enum": [ + "positive", + "negative", + "neutral" + ] + } + } + }, + "total_files": 20, + "tests": [ + { + "file": "t01.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "This product is absolutely amazing! I love it and would recommend it to everyone" + }, + { + "file": "t02.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "This service was terrible. The staff was rude, the food was cold, and I had to w" + }, + { + "file": "t03.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The meeting is scheduled for 3 PM in conference room B. Please bring the quarter" + }, + { + "file": "t04.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "Excellent experience from start to finish. Highly recommend this to anyone looki" + }, + { + "file": "t05.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "Absolutely disappointed with this purchase. The item arrived damaged and custome" + }, + { + "file": "t06.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The software update was installed successfully. System is functioning normally w" + }, + { + "file": "t07.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "The new interface is intuitive and the new features are genuinely useful. Great " + }, + { + "file": "t08.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "I'm very frustrated with this app. It keeps crashing and the latest update remov" + }, + { + "file": "t09.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The package contains 250 units as ordered. Shipping was completed within the agr" + }, + { + "file": "t10.txt", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "preview": "Despite some initial confusion, the support team was patient and helped resolve " + }, + { + "file": "t11.txt", + "expected": "negative", + "difficulty": "medium", + "category": "sentiment", + "preview": "The documentation is outdated and incomplete. Half the examples don't work and k" + }, + { + "file": "t12.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Monthly recurring revenue increased 2.3% quarter-over-quarter. Customer churn ra" + }, + { + "file": "t13.txt", + "expected": "positive", + "difficulty": "hard", + "category": "sentiment", + "preview": "The subtle improvements to the notification system really make a difference in d" + }, + { + "file": "t14.txt", + "expected": "negative", + "difficulty": "hard", + "category": "sentiment", + "preview": "The constant notifications are disruptive and I find the new design choices ques" + }, + { + "file": "t15.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "I'm thrilled with the results. The quality exceeded expectations and delivery wa" + }, + { + "file": "t16.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "Worst experience ever. The product broke within a week and I couldn't get a refu" + }, + { + "file": "t17.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Temperature reading: 72 degrees Fahrenheit. Humidity: 45%. No anomalies detected" + }, + { + "file": "t18.txt", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "preview": "The training workshop was well-organized and the instructors were knowledgeable." + }, + { + "file": "t19.txt", + "expected": "negative", + "difficulty": "medium", + "category": "sentiment", + "preview": "The conference was overcrowded and poorly organized. Sessions started late repea" + }, + { + "file": "t20.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Server uptime this month: 99.87%. Average response time: 142ms. Number of incide" + } + ] +} \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json new file mode 100644 index 000000000000..53c2d1ef300c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json @@ -0,0 +1,12 @@ +{ + "properties": { + "sentiment": { + "type": "string", + "enum": [ + "positive", + "negative", + "neutral" + ] + } + } +} \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt new file mode 100644 index 000000000000..85df534e13c1 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt @@ -0,0 +1 @@ +This product is absolutely amazing! I love it and would recommend it to everyone. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt new file mode 100644 index 000000000000..665b2b23bc14 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt @@ -0,0 +1 @@ +This service was terrible. The staff was rude, the food was cold, and I had to wait an hour. Never coming back. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt new file mode 100644 index 000000000000..465ccc8fa2e1 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt @@ -0,0 +1 @@ +The meeting is scheduled for 3 PM in conference room B. Please bring the quarterly report. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt new file mode 100644 index 000000000000..4fe834ea1977 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt @@ -0,0 +1 @@ +Excellent experience from start to finish. Highly recommend this to anyone looking to improve their workflow. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt new file mode 100644 index 000000000000..d4ad6be49851 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt @@ -0,0 +1 @@ +Absolutely disappointed with this purchase. The item arrived damaged and customer service was unresponsive. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt new file mode 100644 index 000000000000..99593ec45f93 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt @@ -0,0 +1 @@ +The software update was installed successfully. System is functioning normally with no errors reported. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt new file mode 100644 index 000000000000..ee8c7d259e6c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt @@ -0,0 +1 @@ +The new interface is intuitive and the new features are genuinely useful. Great job on the redesign! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt new file mode 100644 index 000000000000..99c217ac0c9f --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt @@ -0,0 +1 @@ +I'm very frustrated with this app. It keeps crashing and the latest update removed features I relied on. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt new file mode 100644 index 000000000000..092edc95a4ab --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt @@ -0,0 +1 @@ +The package contains 250 units as ordered. Shipping was completed within the agreed timeframe of 3-5 business days. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt new file mode 100644 index 000000000000..2b933f269759 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt @@ -0,0 +1 @@ +Despite some initial confusion, the support team was patient and helped resolve the issue quickly. Thank you! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt new file mode 100644 index 000000000000..776e2296739e --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt @@ -0,0 +1 @@ +The documentation is outdated and incomplete. Half the examples don't work and key features are undocumented. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt new file mode 100644 index 000000000000..c06f4595734d --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt @@ -0,0 +1 @@ +Monthly recurring revenue increased 2.3% quarter-over-quarter. Customer churn rate is 1.2% below industry average. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt new file mode 100644 index 000000000000..9133644f757a --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt @@ -0,0 +1 @@ +The subtle improvements to the notification system really make a difference in daily productivity. Worth the upgrade. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt new file mode 100644 index 000000000000..c08983c91a1c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt @@ -0,0 +1 @@ +The constant notifications are disruptive and I find the new design choices questionable at best. Regretting this update. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt new file mode 100644 index 000000000000..52354cff20bc --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt @@ -0,0 +1 @@ +I'm thrilled with the results. The quality exceeded expectations and delivery was faster than promised. Five stars! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt new file mode 100644 index 000000000000..2ed09b28b27e --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt @@ -0,0 +1 @@ +Worst experience ever. The product broke within a week and I couldn't get a refund. Stay away from this seller. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt new file mode 100644 index 000000000000..f271e1e6dbc6 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt @@ -0,0 +1 @@ +Temperature reading: 72 degrees Fahrenheit. Humidity: 45%. No anomalies detected in the past 24-hour monitoring period. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt new file mode 100644 index 000000000000..09d654200723 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt @@ -0,0 +1 @@ +The training workshop was well-organized and the instructors were knowledgeable. I learned several valuable techniques. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt new file mode 100644 index 000000000000..c519d34837db --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt @@ -0,0 +1 @@ +The conference was overcrowded and poorly organized. Sessions started late repeatedly and the venue was subpar. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt new file mode 100644 index 000000000000..b787a856a306 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt @@ -0,0 +1 @@ +Server uptime this month: 99.87%. Average response time: 142ms. Number of incidents reported: 3. All resolved within SLA. \ No newline at end of file diff --git a/tools/server/server-context.cpp b/tools/server/server-context.cpp index 41c6da3e1bbc..28a6b3b75050 100644 --- a/tools/server/server-context.cpp +++ b/tools/server/server-context.cpp @@ -9,6 +9,7 @@ #include "build-info.h" #include "common.h" +#include "base64.hpp" #include "fit.h" #include "llama.h" #include "log.h" @@ -39,6 +40,10 @@ constexpr int HTTP_POLLING_SECONDS = 1; +// Toggle debug output with LLAMA_DECISION_DEBUG env var +namespace { bool decision_debug_enabled() { static bool v = std::getenv("LLAMA_DECISION_DEBUG") != nullptr; return v; } } +#define DECISION_DEBUG(fmt, ...) do { if (decision_debug_enabled()) { fprintf(stderr, "[decision-debug] " fmt "\n", ##__VA_ARGS__); } } while(0) + static common_speculative_output_limits server_output_limits(const common_params & params) { if (params.embedding || (params.pooling_type != LLAMA_POOLING_TYPE_UNSPECIFIED && params.pooling_type != LLAMA_POOLING_TYPE_NONE)) { @@ -2387,21 +2392,99 @@ struct server_context_impl { if (!body.contains("contexts") || !body.at("contexts").is_array() || body.at("contexts").empty() || body.at("contexts").size() > 256) { throw std::invalid_argument("\"contexts\" must be an array of 1-256 strings"); } + // Optional: images array (base64-encoded), one per context, in the same order as contexts. + // An empty/null entry in images means "no image for this context". + // Each entry can also be an array of base64 strings for multi-image contexts. std::vector contexts; - for (const auto & c : body.at("contexts")) { + std::vector context_bitmaps; + bool has_images = body.contains("images") && body.at("images").is_array(); + DECISION_DEBUG("handle_decision: has_images=%d", (int)has_images); + if (has_images && body.at("images").size() != body.at("contexts").size()) { + throw std::invalid_argument("\"images\" array length must match \"contexts\" length"); + } + contexts.reserve(body.at("contexts").size()); + context_bitmaps.reserve(body.at("contexts").size()); + for (size_t i = 0; i < body.at("contexts").size(); ++i) { + const auto & c = body.at("contexts")[i]; if (!c.is_string() || c.get().empty()) { throw std::invalid_argument("every entry of \"contexts\" must be a non-empty string"); } - contexts.push_back(c.get()); + std::string ctx_text = c.get(); + mtmd::bitmaps ctx_bitmaps; + if (has_images && i < body.at("images").size()) { + // images[i] can be a single base64 string OR an array of base64 strings + const auto & img_entry = body.at("images")[i]; + std::vector img_list; + if (img_entry.is_array()) { + for (const auto & img : img_entry) { + if (img.is_string() && !img.get().empty()) { + img_list.push_back(img.get()); + } + } + } else if (img_entry.is_string() && !img_entry.get().empty()) { + img_list.push_back(img_entry.get()); + } + for (const std::string & b64 : img_list) { + DECISION_DEBUG("handle_decision: decoding image for context %zu", i); + // decode base64 image + std::string raw_b64 = b64; + // strip optional data URL prefix: "data:image/...;base64,...." + if (raw_b64.find("data:") == 0) { + auto pos = raw_b64.find(",base64,"); + if (pos != std::string::npos) { + raw_b64 = raw_b64.substr(pos + 8); + } else { + auto pos2 = raw_b64.find(","); + if (pos2 != std::string::npos) { + raw_b64 = raw_b64.substr(pos2 + 1); + } + } + } + std::string raw = base64::decode(raw_b64); + DECISION_DEBUG("handle_decision: image raw size=%zu", raw.size()); + if (!raw.empty()) { + auto out = mtmd_helper_bitmap_init_from_buf(mctx, reinterpret_cast(raw.data()), raw.size(), false, init_opt); + DECISION_DEBUG("handle_decision: bitmap init result.bitmap=%p", (void*)out.bitmap); + if (out.bitmap) { + ctx_bitmaps.entries.emplace_back(out.bitmap); + } else { + throw std::runtime_error("failed to decode image at context " + std::to_string(i)); + } + } + } + } + if (!ctx_bitmaps.entries.empty()) { + // Media markers should already be in the context text at this point. + // If the context text doesn't contain any media markers, prepend them + // (backward compatibility with single-image contexts). + // The harness can now include media markers inline in context text + // for multi-image contexts where image order matters relative to text. + const char * marker = mctx ? mtmd_get_marker(mctx) : nullptr; + if (marker && ctx_text.find(marker) == std::string::npos) { + // No media markers in text, so prepend them (legacy behavior) + std::string markers; + for (size_t j = 0; j < ctx_bitmaps.entries.size(); ++j) { + markers += marker; + } + ctx_text = markers + ctx_text; + DECISION_DEBUG("handle_decision: prepended %zu media markers", ctx_bitmaps.entries.size()); + } + // If markers are already in the text, mtmd_tokenize will find and use them + } + contexts.push_back(ctx_text); + context_bitmaps.emplace_back(std::move(ctx_bitmaps)); } if (!body.contains("schema")) { throw std::invalid_argument("\"schema\" must be provided"); } if (!decision_engine) { + DECISION_DEBUG("handle_decision: creating decision engine"); decision_engine = std::make_unique(ctx_tgt, (llama_seq_id) params_base.n_parallel, - params_base.n_seq_decision); + params_base.n_seq_decision, mctx); } + DECISION_DEBUG("handle_decision: compiling schema"); const auto cs = llama_decision::compile_schema(body.at("schema"), body.value("instructions", std::string())); + DECISION_DEBUG("handle_decision: rendering prompts"); std::string shared; std::vector dynamic; for (const auto & c : contexts) { @@ -2418,6 +2501,8 @@ struct server_context_impl { opt.tree_max = (size_t) body.value("tree_max", 128); opt.allow_cache = body.value("cache_prompt", true); + // If any context has images, pass the bitmaps through options for multimodal tokenization + opt.context_bitmaps = std::move(context_bitmaps); const auto b = decision_engine->decide_batch(shared, dynamic, cs.inputs, opt); size_t context_tokens = 0; From acf2dde10305c533a1af458d6908c20a60da43ee Mon Sep 17 00:00:00 2001 From: graydini Date: Tue, 29 Sep 2026 02:22:17 -0700 Subject: [PATCH 2/3] docs : document the decision endpoint, images, and the harness Link parallel-decision from the root README so the feature is reachable from the front page instead of only from its own directory. Explain that contexts can carry images, and that chunks are encoded in the same batched pass that scores the branches, so a decision over a screenshot stays a single pass rather than a run of generate calls. Add an Images section with the request shape, the positional rule for the images array, the multi-image form, and how media markers decide where each image lands. Fix the endpoint path: the docs said /v1/decision, the server serves POST /decision. --- README.md | 7 ++++++ tools/parallel-decision/README.md | 39 ++++++++++++++++++++++++++++--- 2 files changed, 43 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index aae3bcd35ad9..76f002985ea4 100644 --- a/README.md +++ b/README.md @@ -89,6 +89,13 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or ## Documentation +#### Constrained decisions + +- [parallel-decision](tools/parallel-decision/README.md) - answer a JSON schema in one batched + pass, with images, over `POST /decision` +- [vision decision harness](tools/parallel-decision/examples/vision-decision-harness/) - example + web UI for driving the endpoint over images and text + #### Tools - [cli](tools/cli/README.md) diff --git a/tools/parallel-decision/README.md b/tools/parallel-decision/README.md index eb372d9e794e..27e37f0ffad0 100644 --- a/tools/parallel-decision/README.md +++ b/tools/parallel-decision/README.md @@ -7,8 +7,12 @@ After the context, each field's allowed values are scored as token paths that fo fields are answered in one `llama_decode` and cannot see each other. Each answer comes back with a probability, and the JSON object is assembled by code, so it always matches the schema. +Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and the chunks are encoded +through the same mtmd path the completion endpoint uses, so a decision runs over a screenshot, a document scan, or a +folder of images without leaving the single-pass scoring model. + This directory holds the engine (`decision-engine.*`), a CLI (`llama-parallel-decision`), and the engine is also -served by `llama-server` as `POST /v1/decision`. +served by `llama-server` as `POST /decision`. ## Build @@ -59,13 +63,13 @@ sequence (about 50 MB each for Qwen3.5 4B and 9B), and llama.cpp only batches th the same number of tokens. The engine right-pads each group of branches to its longest one, so they still score in a single pass; the padding comes after the token that is read, so it doesn't change the result. -## POST /v1/decision +## POST /decision `contexts` is a list of 1-256 strings. They share one schema, one set of instructions, and one cached prefix; results come back in the same order. ```bash -curl http://localhost:8096/v1/decision -H "Content-Type: application/json" -d '{ +curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{ "model": "gemma-4-12b", "instructions": "Answer each question about this support request from its state.", "schema": { @@ -122,6 +126,35 @@ Numeric fields take `aggregate`: `mode` (default), `median` or `mean`. | `tree_max` | 128 | per-field switch between tree and greedy | | `cache_prompt` | true | reuse the cached instructions + schema prefix | +## Images + +Add `images` alongside `contexts`. It is positional: entry *i* belongs to context *i*. An entry is either one +base64 string or an array of base64 strings when a single context should see several images. A data URL prefix +is accepted and stripped. + +```bash +curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{ + "instructions": "Answer each question about this screenshot.", + "schema": { + "properties": { + "page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]}, + "error": {"type": "boolean"} + } + }, + "contexts": ["What kind of page is this?"], + "images": ["iVBORw0KGgoAAAANSUhEUg..."] +}' +``` + +Images are placed by the media marker that `mtmd` reserves for the loaded projector. If the context text already +contains that marker, the marker is left where the caller put it, so you can interleave markers with your own labels +and control which part of the text each image belongs to. If the text has no marker, one is prepended per image. + +Because the chunks are encoded in the same batched pass that scores the branches, adding images does not turn a +decision into a sequence of generate calls. + +Set `LLAMA_DECISION_DEBUG` in the environment to trace tokenization, chunk encoding, and decode on this path. + ## CLI `llama-parallel-decision` runs the same engine from a worker process (stdin/stdout protocol, one JSON request per From 66fa6c95629fb89ef2969db0c57992193471c053 Mon Sep 17 00:00:00 2001 From: graydini Date: Tue, 29 Sep 2026 02:33:55 -0700 Subject: [PATCH 3/3] docs : lead the README with constrained decisions Promote the decision branch from a footnote to the main topic. The feature was previously two bullets under Documentation, which put the mechanism, the vision work, and the harness behind the generic tool list. Constrained decisions now comes first, with what the batched scoring does, a runnable server command, a request and a response, and the positional rule for images. The vision harness gets its own section rather than a list item, and the debug toggle is documented next to it. The upstream content keeps its own heading below a divider, with the section levels demoted one step so the hierarchy stays correct. The example response is a real one from the server. --- README.md | 88 ++++++++++++++++++++++++++++++++++++++++++++++--------- 1 file changed, 74 insertions(+), 14 deletions(-) diff --git a/README.md b/README.md index 76f002985ea4..6570d55cc56f 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,7 @@
-LLM inference in C/C++ +LLM inference in C/C++, with batched constrained decisions over text and images [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://opensource.org/licenses/MIT) [![Release](https://img.shields.io/github/v/release/ggml-org/llama.cpp?filter=v*&color=brightgreen)](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0) @@ -17,7 +17,74 @@
-## Quick start +## Constrained decisions + +Instead of generating a JSON object one token at a time, this branch scores a whole schema in a single batched +forward pass. Every field has a fixed set of allowed values, so the values are scored as token paths that fork from +the same KV cache. All fields are answered in one `llama_decode`, cannot see each other, and the object is +assembled by code, so the output always matches the schema. Each field comes back with a probability. + +Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and chunks encode through +the same mtmd path the completion endpoint uses. A decision over a screenshot, a document scan, or a whole folder of +images stays a single pass rather than becoming a run of generate calls. + +```bash +./build/bin/llama-server -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \ + --decision-seqs 8 --host 0.0.0.0 --port 8081 +``` + +```bash +curl http://localhost:8081/decision -H "Content-Type: application/json" -d '{ + "instructions": "Answer each question about this screenshot.", + "schema": { + "properties": { + "page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]}, + "error": {"type": "boolean"} + } + }, + "contexts": ["What kind of page is this?"], + "images": ["iVBORw0KGgoAAAANSUhEUg..."] +}' +``` + +```json +{ + "object": "decision", + "results": [ + { + "decision": {"page": "settings", "error": true}, + "fields": { + "page": {"value": "settings", "probability": 0.868, "scored_nodes": 1, "tree": true}, + "error": {"value": true, "probability": 0.966, "scored_nodes": 1, "tree": true} + }, + "usage": {"context_tokens": 516, "scored_rows": 9} + } + ] +} +``` + +`images` is positional: entry *i* belongs to context *i*. An entry is one base64 string, or an array when a single +context should see several images. Media markers already present in the context text are left where the caller put +them, so images can be interleaved with the caller's own labels. + +Full reference: [parallel-decision](tools/parallel-decision/README.md). + +### Vision decision harness + +`tools/parallel-decision/examples/vision-decision-harness/` is a runnable example UI for the endpoint. It does +folder upload, batch runs over images and text with SSE progress, image selection across a folder, a text +classification suite, and snippet export. It is an example, not a dependency, and nothing in the server links +against it. See its [README](tools/parallel-decision/examples/vision-decision-harness/README.md). + +### Set `LLAMA_DECISION_DEBUG` + +Traces tokenization, chunk encoding, and decode on the decision path. + +## Standard llama.cpp + +Everything below this point is the upstream project as usual. + +### Quick start A few options to get `llama.cpp` installed on your machine: @@ -49,7 +116,7 @@ llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF -## Description +### Description The main goal of `llama.cpp` is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. @@ -65,7 +132,7 @@ a wide range of hardware - locally and in the cloud. The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-org/ggml) library. -## Supported backends +### Supported backends | Backend | Target devices | | --- | --- | @@ -87,14 +154,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or | [WebGPU](docs/build.md#webgpu) | All | | [ZenDNN](docs/build.md#zendnn) | AMD CPU | -## Documentation - -#### Constrained decisions - -- [parallel-decision](tools/parallel-decision/README.md) - answer a JSON schema in one batched - pass, with images, over `POST /decision` -- [vision decision harness](tools/parallel-decision/examples/vision-decision-harness/) - example - web UI for driving the endpoint over images and text +### Documentation #### Tools @@ -116,7 +176,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or - [Models](docs/models.md) - [Release process](docs/release.md) -## Contributing +### Contributing - Contributors can open PRs - Collaborators will be invited based on contributions @@ -124,7 +184,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or - Any help with managing issues, PRs and projects is very appreciated! - Read the [CONTRIBUTING.md](CONTRIBUTING.md) for more information -## Acknowledgements +### Acknowledgements - [yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) - Single-header HTTP server, used by `llama-server` - MIT license - [nothings/stb](https://github.com/nothings/stb) - Single-header image format decoder, used by multimodal subsystem - Public domain