Skip to content

Repository files navigation

PDF Markdown Studio

PDF Markdown Studio Demo

Install on Debian 13+ or Ubuntu 24.04/26.04+

For the first Open Research Tools installation on a system, copy and run this one command:

wget -qO /tmp/keyring.deb https://keyring.openresearchtools.com && sudo apt install -y /tmp/keyring.deb && sudo apt update && sudo apt install -y pdf-markdown-studio

If the Open Research Tools APT repository is already configured:

sudo apt install pdf-markdown-studio

PDF Markdown Studio is a desktop app for converting PDFs and images into clean Markdown.

It is built as an example integration client for OpenResearchTools Engine, showing how to run Engine PDF and VLM workflows from a GUI.

The PDF Markdown Studio application source code is licensed under the MIT License; third-party dependencies and bundled components remain licensed under their respective original licenses.

What You Can Do

  • Add multiple files (PDFs and images) into one workspace.
  • Preview the source document and generated Markdown side-by-side.
  • Search by text and jump through matching pages.
  • Convert selected files using either fast PDF extraction or VLM-based extraction.
  • Edit Markdown in place and save changes.

Supported Files

  • PDF (.pdf)
  • Images (.png, .jpg, .jpeg)

Conversion Modes (How To Choose)

1) FAST PDF

Use this for machine-readable digital PDFs.

  • Very fast.
  • Best when text is selectable in the PDF.
  • Output quality is strong for standard digital documents.
  • Limitation: table content is often flattened/inline in output; complex tables can become misstructured.

2) PDF VLM + Image VLM

Use this when layout is complex, scanned-like, visual-heavy, or FAST output is poor.

  • Uses your selected VLM model + MMProj.
  • Works for both PDFs and images.
  • Slower than FAST, but better for difficult pages.
  • Engine applies a per-page quality gate for PDF VLM output. Pages that end up obviously truncated, stuck in repetition/looping, or otherwise fail the gate are retried automatically before final output is written.
  • This is meant to reduce bad outputs on difficult pages, but manual inspection is still recommended for important documents.
  • Limitation: table quality depends on the selected model's capabilities.
  • On complex tables with heavy formatting/whitespace, models can misattribute values to wrong cells or rows.
  • If downstream automation depends on table values, compare Markdown against the original document before automated extraction.

3) FAST PDF + VLM fallback

Try FAST first, then automatically switch to VLM if FAST identifies non-machine-readable content.

  • Good default when PDF quality is mixed.
  • Balances speed and robustness.

Quick Start

  1. Open the app.
  2. In Settings, make sure runtime is healthy and model paths are set.
  3. Click File, Add PDFs / Images.
  4. Tick files in the Documents sidebar.
  5. Select a conversion mode.
  6. Click Convert Selected.

Output files are written next to the source document:

  • filenameFAST.md
  • filenameVLM.md

Viewing and Navigation

  • Left pane: original PDF/image.
  • Right pane: Markdown preview/edit.
  • Find, Prev/Next Hit, and page controls are above the workspace.
  • View zoom affects both panes together.

Runtime and Model Location

Default Engine runtime path:

  • Windows: C:\Users\<user>\AppData\Roaming\OpenResearchTools\PDF Markdown Studio\Engine
  • macOS: ~/Library/Application Support/OpenResearchTools/PDF Markdown Studio/Engine
  • Linux Vulkan: /opt/openresearchtools/engine/vulkan (package openresearchtools-engine)
  • Linux CUDA: /opt/openresearchtools/engine/cuda (package openresearchtools-engine-cuda)

On Linux the app does not download or maintain a private Engine copy. Its Vulkan/CUDA selector switches directly between those two system-installed package roots. The pdf-markdown-studio Debian package depends on both Engine packages so both choices are available after a normal APT installation.

Linux keeps the GUI and inference runtime in separate processes. Device enumeration, FAST conversion, PDF VLM, and image VLM run through example-cli from the selected package root. This prevents Vulkan and CUDA libraries with matching names from being loaded into the same GUI process and makes backend switching deterministic.

Default app settings/data path:

  • Windows config/data: C:\Users\<user>\AppData\Roaming\OpenResearchTools\PDF Markdown Studio
  • macOS config/data: ~/Library/Application Support/OpenResearchTools/PDF Markdown Studio
  • Linux config: ~/.config/OpenResearchTools/PDF Markdown Studio
  • Linux data: ~/.local/share/OpenResearchTools/PDF Markdown Studio

Shared VLM model folder:

  • Windows: C:\Users\<user>\AppData\Roaming\OpenResearchTools\models
  • macOS: ~/Library/Application Support/OpenResearchTools/models
  • Linux: ~/.local/share/OpenResearchTools/models

Each selected Qwen3.5 family downloads into its own shared repo folder under that global models root, including the required MMProj file for the chosen family.

GPU / CPU Execution

  • CPU mode: run without GPU acceleration.
  • GPU mode: select one GPU in settings.
  • The app sends one selected GPU for VLM execution paths through Engine runtime.

Troubleshooting

Unsigned Build Notice

This app is an open-source hobby development effort by the repository owner. We do not currently have funding for full paid code-signing and notarization pipelines across all platforms/releases.

Because of that, operating-system protections or hardened security environments (for example Windows SmartScreen, enterprise endpoint controls, or macOS Gatekeeper policies) may block unsigned binaries.

If your environment blocks unsigned binaries, the recommended path is:

  • build this desktop app from source on the target device,
  • build Openresearchtools-Engine from source on the same target device,
  • and use those locally-built artifacts in your deployment.

Windows (when blocked)

  • If SmartScreen shows "Windows protected your PC", use More info -> Run anyway only if your policy allows it.
  • In the app, go to Settings -> Runtime Setup and run:
    • Download/Repair runtime
    • Unblock unsigned runtime
    • Recheck
  • The Windows unblock script clears Mark-of-the-Web flags in the selected runtime directory by running Unblock-File recursively on runtime files.

macOS (when blocked)

  • Try Right click -> Open on first launch.
  • If blocked by Gatekeeper, use System Settings -> Privacy & Security -> Open Anyway when available and policy permits.
  • In the app, after runtime install/repair, click Unblock unsigned runtime then Recheck.
  • The macOS unblock script removes quarantine attributes recursively (xattr -dr com.apple.quarantine) and restores executable bits for runtime binaries/scripts where needed (chmod +x on relevant files).

If conversion fails or setup is incomplete

  1. Open Settings.
  2. On Windows/macOS, use the runtime health/check and download/repair actions. On Linux, reinstall the system packages if the selected runtime check fails: sudo apt install --reinstall openresearchtools-engine openresearchtools-engine-cuda.
  3. Confirm model and MMProj paths exist.
  4. Check Jobs and logs for the exact error.

If adding many files feels slow

  • Wait for background imports to finish before converting.
  • Large PDFs can take time to rasterize and preview.

Acknowledgements (What This App Uses)

This app uses OpenResearchTools Engine runtime components for PDF and VLM execution. For this app's active feature set, key upstream technologies include:

  • Openresearchtools-Engine: embeddable runtime used by this app (llama-server-bridge, runtime orchestration, and model/device execution path).
  • egui / eframe: native immediate-mode GUI framework used to build this desktop application UI.
  • llama.cpp and ggml: core inference runtime and device/offload mechanics used through Openresearchtools-Engine.
  • Docling: reference logic for VLM document-conversion behavior used by Engine pdfvlm, including page-wise rendering/scaling heuristics (scale, oversample) and Catmull-Rom style downscale before inference.
  • PDFium and pdfium-render: PDF rasterization/page access primitives used by the app's native PDF rendering and by Engine PDF conversion paths.

Current VLM Model Lineup

PDF Markdown Studio now uses the Qwen3.5 GGUF + MMProj model family Qwen: upstream Qwen3.5 model family reference used for the app's current vision model lineup for PDF VLM and Image VLM conversion:

  • Qwen3.5 9B (Q4_K_M and Q8_0)
  • Qwen3.5 4B (Q4_K_M and Q8_0)
  • Qwen3.5 2B (Q4_K_M and Q8_0)

The app downloads the text model and the matching MMProj for the selected family automatically into the shared OpenResearchTools model store.

Note the use of the models in our app does not imply affiliations or endorsements from original model authors. This is just a personal recommendation after testing many currently available models for speed/quality of the outputs. You are also free to use any other vision model that can run on GGML (llama.cpp backend). The app allows for manual model selection.

Recommended guidance:

  • 9B 4-bit is the recommended default for documents that need higher precision, denser layout understanding, or more reliable structure recovery.
  • 2B models often still produce surprisingly strong results at a fraction of the compute cost, and are a good option when you want speed or need to run on lighter hardware.
  • 4B is the middle ground when you want a better quality/speed balance.

openresearchtools/Qwen3.5-9B-GGUF, openresearchtools/Qwen3.5-4B-GGUF, and openresearchtools/Qwen3.5-2B-GGUF: converted GGUF + MMProj model repositories used by the app for PDF VLM and Image VLM conversion.

  • Qwen: upstream Qwen3.5 model family reference used for the app's current vision model lineup.

This project is independent and is not affiliated with, sponsored by, or endorsed by egui, llama.cpp, Docling, PDFium, Qwen, or other upstream projects/vendors.

How to cite

Suggested citation:

Rutkauskas, L. (2026). PDF Markdown Studio (Version 1.1.2) [Computer software]. OpenResearchTools. https://github.com/openresearchtools/pdfmarkdownstudio.

BibTeX:

@software{Rutkauskas_PDFMarkdownStudio_2026,
  author    = {Rutkauskas, L.},
  title     = {PDF Markdown Studio},
  version   = {1.1.2},
  date      = {2026-03-04},
  url       = {https://github.com/openresearchtools/pdfmarkdownstudio},
  publisher = {OpenResearchTools},
  license   = {MIT}
}

Linux ARM64 Vulkan release

Release 1.1.3 adds a native Linux ARM64 .deb. ARM64 always uses Vulkan, including when migrating saved CUDA settings, and has no CUDA runtime selector or CUDA package dependency. The package depends on openresearchtools-engine (>= 1.17); APT selects and downloads the ARM64 engine automatically.

The manual Linux ARM64 Vulkan release workflow builds and tests on ubuntu-24.04-arm. It copies the Windows x64, macOS ARM64 and Linux AMD64 binaries unchanged from release 1.1.2, verifies their GitHub SHA-256 digests, and writes fresh SHA256SUMS.txt. Their embedded versions are retained. The default draft release allows hardware testing before publication. The linux-arm64-hardware-tests Actions artifact contains the native test executable for testing with an installed engine.

Manual installation without the Open Research Tools APT repository:

wget https://github.com/openresearchtools/engine/releases/download/v1.17/engine-arm64.deb
wget https://github.com/openresearchtools/PDF-Markdown-Studio/releases/download/1.1.3/pdf-markdown-studio-ubuntu-arm64.deb
sudo apt install ./engine-arm64.deb ./pdf-markdown-studio-ubuntu-arm64.deb

About

Convert PDFs and Images to Markdown using either FAST text extractors or local vision models.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages