Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions tools/mtmd/mtmd.h
Original file line number Diff line number Diff line change
Expand Up @@ -507,6 +507,11 @@ struct bitmap {
struct bitmaps {
std::vector<bitmap> entries;
~bitmaps() = default;
bitmaps() = default;
bitmaps(bitmaps && other) noexcept = default;
bitmaps & operator=(bitmaps && other) noexcept = default;
bitmaps(const bitmaps &) = delete;
bitmaps & operator=(const bitmaps &) = delete;
// return list of pointers to mtmd_bitmap
// example:
// auto bitmaps_c_ptr = bitmaps.c_ptr();
Expand Down
2 changes: 1 addition & 1 deletion tools/parallel-decision/CMakeLists.txt
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# shared engine: used by llama-parallel-decision and by llama-server's /decision endpoint
add_library(llama-decision STATIC decision-engine.cpp decision-engine.h)
target_include_directories(llama-decision PUBLIC ${CMAKE_CURRENT_SOURCE_DIR})
target_link_libraries(llama-decision PUBLIC llama-common llama)
target_link_libraries(llama-decision PUBLIC llama-common llama mtmd)
target_compile_features(llama-decision PUBLIC cxx_std_17)
# llama-server links it into a shared library (libllama-server-impl)
set_target_properties(llama-decision PROPERTIES POSITION_INDEPENDENT_CODE ON)
Expand Down
33 changes: 33 additions & 0 deletions tools/parallel-decision/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,3 +132,36 @@ line). Environment: `DECIDE_TREE`, `DECIDE_TREE_MAX`, `DECIDE_NSEQ`, `DECIDE_SPL
[decision-playground](https://github.com/thecodacus/decision-playground) is a browser-only playground: it talks
straight to your llama-server, runs a decision and the same question as a chat completion side by side with live
timers, and has a small game whose agents decide through the endpoint.

### Vision Decision Harness (Example)

For multimodal decision testing with image support, this directory ships a runnable example under
`examples/vision-decision-harness/`. It is a small Flask web UI that:

- Accepts image folder uploads and runs them through `/decision` in a batch
- Adds an image selection mode: every image in a folder is scored in one decision pass and the
UI reports the single best match for a question
- Streams results back over SSE as each file is scored
- Ships a text test suite for measuring classification accuracy and calibration
- Ships `tests/scan_for_secrets.py`, which uses the decision endpoint itself to flag files that
look like they contain private data before you commit

Build and run the server first:

```bash
./build/bin/llama-server --host 0.0.0.0 --port 8081 \
-m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \
--decision-seqs 8
```

Then the harness:

```bash
cd tools/parallel-decision/examples/vision-decision-harness
pip install -r requirements.txt
python3 app.py
```

The harness listens on port 5786 and proxies to the server on port 8081. Set `LLAMA_SERVER_URL`
to point it somewhere else.

249 changes: 233 additions & 16 deletions tools/parallel-decision/decision-engine.cpp

Large diffs are not rendered by default.

46 changes: 44 additions & 2 deletions tools/parallel-decision/decision-engine.h
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@
// larger fields walk the trie greedily.

#include "llama.h"
#include "mtmd.h"
#include "mtmd-helper.h"
#include "json.h"

#include <string>
Expand All @@ -34,6 +36,10 @@ struct options {
size_t tree_max = 128;
bool split_boundary = false; // legacy: tokenise suffix and values separately
bool allow_cache = true; // reuse the cached static prefix when it matches
// Optional per-context bitmaps for multimodal decision. If non-empty, must have the same
// size as the contexts vector in decide_batch. Each entry holds bitmaps to prepend to
// that context (the context text should contain media markers at the corresponding positions).
std::vector<mtmd::bitmaps> context_bitmaps;
};

struct field_result {
Expand Down Expand Up @@ -72,7 +78,7 @@ struct batch_result {
// flight, then branches. The context needs a unified KV cache so branches share the trunk's cells.
class engine {
public:
engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs);
engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs, mtmd_context * mctx = nullptr);

result decide(const std::string & shared_text, const std::string & context_text,
const std::vector<field_input> & fields, const options & opt);
Expand All @@ -83,8 +89,11 @@ class engine {
const std::vector<field_input> & fields, const options & opt);

private:
// A prompt part can be either a list of text tokens or a media chunk (image/audio).
struct prompt_part {
const tokens_t * toks;
enum type { TOKS, CHUNK } kind;
const tokens_t * toks = nullptr; // when kind == TOKS
const mtmd_input_chunk * chunk = nullptr; // when kind == CHUNK
llama_pos pos0;
llama_seq_id seq;
};
Expand All @@ -102,7 +111,40 @@ class engine {
int n_pool;
bool pad_branches; // recurrent/hybrid model: branches in a decode need equal lengths
tokens_t cached;
mtmd_context * mctx; // optional multimodal context for vision input

// Tokenize text that may contain media markers, expanding them into chunks via mtmd.
// Returns text tokens with LLAMA_TOKEN_NULL at image positions, and the image chunks
// that need to be encoded separately in decode_parts.
struct multimodal_tokens {
tokens_t toks; // text tokens, LLAMA_TOKEN_NULL at image positions
std::vector<const mtmd_input_chunk *> chunks; // image/audio chunks to decode at those positions

~multimodal_tokens() {
for (const auto * chunk : chunks) {
mtmd_input_chunk_free(const_cast<mtmd_input_chunk *>(chunk));
}
}
multimodal_tokens() = default;
multimodal_tokens(multimodal_tokens && other) noexcept
: toks(std::move(other.toks)), chunks(std::move(other.chunks)) {}
multimodal_tokens & operator=(multimodal_tokens && other) noexcept {
if (this != &other) {
for (const auto * chunk : chunks) {
mtmd_input_chunk_free(const_cast<mtmd_input_chunk *>(chunk));
}
toks = std::move(other.toks);
chunks = std::move(other.chunks);
}
return *this;
}
// non-copyable (chunks are owned)
multimodal_tokens(const multimodal_tokens &) = delete;
multimodal_tokens & operator=(const multimodal_tokens &) = delete;
};

multimodal_tokens tokenize_mm(const std::string & text, bool add_special,
const mtmd::bitmaps * bitmaps = nullptr) const;
tokens_t tokenize(const std::string & text, bool add_special) const;
void decode_parts(const std::vector<prompt_part> & parts);
bool prepare_prefix(const tokens_t & shared, bool allow_cache);
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
__pycache__/
*.pyc
# written by the test runners on each run
test_results.json
evaluation_results.json
152 changes: 152 additions & 0 deletions tools/parallel-decision/examples/vision-decision-harness/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,152 @@
# Vision Decision Harness

A web-based UI for testing the multimodal `/decision` endpoint of llama-server, built on top of the [parallel-decision branch](https://github.com/thecodacus/llama.cpp) of llama.cpp.

## Overview

The harness provides a browser-based interface for evaluating multimodal decision-making capabilities of llama-server. It supports:

- **Image batch processing**: Upload folders of images and run classification questions
- **Image Selection mode**: Scan a folder of images and find which one matches a question (e.g. "which image contains a rubber duck?")
- **Text file batch processing**: Run text files through the decision endpoint for scoring/classification evaluation
- **Dynamic decision tags**: Add/remove decision options with a + button (tag-list style)
- **Prompt autocompletion**: Get suggestions from the llama-server `/completion` endpoint
- **Benchmarking**: Per-file timing, total batch timing, progress bar
- **Streaming results**: Results appear as each file completes (SSE)
- **Image preview toggle**: Show/hide thumbnails in results
- **Export snippet**: Generate a Python API request snippet with docstring
- **Secret scanning**: Automated scan for sensitive information before committing

## Prerequisites

1. **llama-server with decision + multimodal support** - built from the `multimodal-decision` branch of the [thecodacus/llama.cpp](https://github.com/thecodacus/llama.cpp) fork. The server must be started with `--decision-seqs N` (N >= 3) and `--mmproj` for vision support.

```bash
./build/bin/llama-server \
--host 0.0.0.0 \
--model model-Q4_K_M.gguf \
--mmproj mmproj-model.gguf \
--decision-seqs 8 \
--device Vulkan1 \
--port 8081
```

2. **Python 3** with Flask and requests:
```bash
pip install flask flask-cors requests
```

## Running

```bash
cd vision-decision-harness
python3 app.py
```

The UI is available at `http://0.0.0.0:5786`.

The server URL defaults to `http://0.0.0.0:8081`. Override with:
```bash
LLAMA_SERVER_URL=http://localhost:8081 python3 app.py
```

## API Endpoints

| Endpoint | Method | Description |
|---|---|---|
| `/` | GET | Main UI |
| `/api/health` | GET | Health check (proxies to server) |
| `/api/decision` | POST | Proxy to llama-server `/decision` |
| `/api/completion` | POST | Proxy to llama-server `/completion` |
| `/api/upload` | POST | Upload files/folders from browser |
| `/api/batch` | POST | Batch process files (non-streaming) |
| `/api/batch-stream` | POST | Batch process files with SSE streaming |
| `/api/image-selection/stream` | POST | Image selection mode with SSE streaming |
| `/api/export-snippet` | POST | Generate Python API request snippet |
| `/api/tag-suggestions` | POST | Get tag suggestions from `/completion` |

## Image Selection Mode

The image selection mode sends all images in a folder to the `/decision` endpoint in a single batched request with a yes/no schema. Each image is evaluated as a separate context, and results are aggregated to identify the single best match.

**Request:**
```json
{
"source": "/path/to/image/folder",
"question": "Which image contains a yellow rubber duck?",
"instructions": "Look at this image and determine if it contains the described object. Answer YES if it does, NO if it does not.",
"seed": 42
}
```

**SSE Events:**
- `start`: Found N images, question
- `progress`: Per-image file loaded
- `result`: Per-image yes/no evaluation
- `decision`: Single winning image with file name and probability
- `done`: Total time

## Text Test Suite

The `tests/` directory contains a comprehensive text-based decision scoring test suite:

- `text_samples/` - 20 text files with sentiment classification questions
- `answer_key.json` - Expected results for all test files
- `schema.json` - Shared JSON Schema for the tests
- `evaluate_text_tests.py` - Runner that sends each file through `/decision` and compares to the answer key
- `test_decision_scoring.py` - 41 built-in decision scoring tests (text-only)

Run with:
```bash
python3 tests/evaluate_text_tests.py
python3 tests/test_decision_scoring.py
```

## Secret Scanning

Before committing, scan all harness files and modified llama.cpp files for sensitive information:

```bash
python3 tests/scan_for_secrets.py
```

This script uses the llama-server `/decision` endpoint to classify each file as containing or not containing:
- Passwords
- API keys
- Email addresses
- Phone numbers
- Home directory paths
- IP addresses
- Credentials/tokens

Any files flagged as sensitive are reported for manual review before committing.

## Multimodal Decision Support

### Server-side Changes

The fork adds the following to `tools/server/server-context.cpp`:

1. **Multi-image contexts**: The `images` field can be an array of arrays - one sub-array per context, enabling multiple images per decision context.

2. **Media marker handling**: Context text can contain inline media markers (fetched from `/props`) to position images at specific points in the text. If no markers are present, they are prepended for backward compatibility.

3. **Batch vision encoding path**: Uses `mtmd_batch_init` -> `mtmd_batch_add_chunk` -> `mtmd_batch_encode` -> `mtmd_batch_get_output_embd` -> `mtmd_helper_decode_image_chunk` for vision encoding, matching the server's working completion path. Falls back to per-chene encoding for models that don't support batch encoding (e.g., SmolVLM2).

### Decision Engine Changes

- `decision-engine.h/cpp` - Added `multimodal_tokens` struct, `tokenize_mm()` method, `options::context_bitmaps` field, and batch vision encoding in `decode_parts()`
- `tools/mtmd/mtmd.h` - Added explicit move constructors for `bitmaps` and `bitmap` to support vector storage
- `tools/parallel-decision/CMakeLists.txt` - Links `mtmd` target

### Debug Toggle

All debug prints in the decision engine are controlled by the `LLAMA_DECISION_DEBUG` environment variable:

```bash
LLAMA_DECISION_DEBUG=1 ./build/bin/llama-server --model ...
```

## License

This harness is provided as an example for the parallel-decision multimodal extensions. See the main [llama.cpp README](https://github.com/ggml-org/llama.cpp) for the underlying project license.
Loading