Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 5 additions & 4 deletions en/advanced/OCR.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# OCR

[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports two OCR engines: [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) and [Docling](https://github.com/docling-project/docling).
[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) (with a selectable OCR backend: Tesseract, EasyOCR, PaddleOCR, or AppleOCR) and [Docling](https://github.com/docling-project/docling).

Check warning on line 3 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.TooWordy] 'multiple' is too wordy. Raw Output: {"message": "[write-good.TooWordy] 'multiple' is too wordy.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 3, "column": 381}}}, "severity": "WARNING"}

Check warning on line 3 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.Passive] 'be used' may be passive voice. Use active voice if you can. Raw Output: {"message": "[write-good.Passive] 'be used' may be passive voice. Use active voice if you can.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 3, "column": 369}}}, "severity": "WARNING"}

Check warning on line 3 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.TooWordy] 'it is' is too wordy. Raw Output: {"message": "[write-good.TooWordy] 'it is' is too wordy.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 3, "column": 262}}}, "severity": "WARNING"}

Check warning on line 3 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.TooWordy] 'Consequently' is too wordy. Raw Output: {"message": "[write-good.TooWordy] 'Consequently' is too wordy.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 3, "column": 227}}}, "severity": "WARNING"}

Check warning on line 3 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.Passive] 'is defined' may be passive voice. Use active voice if you can. Raw Output: {"message": "[write-good.Passive] 'is defined' may be passive voice. Use active voice if you can.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 3, "column": 100}}}, "severity": "WARNING"}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

apple OCR? wtf?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a joke — it's real as of 5 days ago: JabRef/jabref#17219 ("Add OCRmyPDF plugins as new engines support (EasyOCR, PaddleOCR, AppleOCR)") added APPLEOCR as a selectable value in EngineSelection, alongside TESSERACT, EASYOCR, and PADDLEOCR, wired into OcrEngineFactory to run through OCRmyPDF's appleocr plugin. It shows up in the "OCR engine" dropdown in File → Preferences → OCR on current main.

Happy to drop it from this PR if it's considered out of scope / not ready for users yet — let me know and I'll push an update.


Generated by Claude Code


## How to install an OCR engine

Expand Down Expand Up @@ -48,8 +48,9 @@

* Available engines:

1. **OCRmyPDF**: the default engine. Well suited for general-purpose OCR on scanned PDFs.
2. **Docling**: an alternative engine with strong handling of complex layouts and documents containing tables or figures (slower than OCRmyPDF).
1. **Tesseract**: the default engine. Runs via OCRmyPDF and is well suited for general-purpose OCR on scanned PDFs.
2. **EasyOCR**, **PaddleOCR**, **AppleOCR**: alternative OCR backends, also run via OCRmyPDF through its respective plugin. Each plugin must be installed separately in addition to OCRmyPDF itself.

Check warning on line 52 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.TooWordy] 'in addition' is too wordy. Raw Output: {"message": "[write-good.TooWordy] 'in addition' is too wordy.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 52, "column": 166}}}, "severity": "WARNING"}

Check warning on line 52 in en/advanced/OCR.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.Passive] 'be installed' may be passive voice. Use active voice if you can. Raw Output: {"message": "[write-good.Passive] 'be installed' may be passive voice. Use active voice if you can.", "location": {"path": "en/advanced/OCR.md", "range": {"start": {"line": 52, "column": 142}}}, "severity": "WARNING"}
3. **Docling**: a separate engine with strong handling of complex layouts and documents containing tables or figures (slower than the OCRmyPDF-based engines).

* Changing the selected engine automatically re-runs **auto-detection** for that engine's path (see below), so if the newly selected engine is installed in a standard location, its path field will populate automatically.

Expand All @@ -74,7 +75,7 @@

* JabRef also **auto-detects the path automatically** whenever you change the selected engine in the dropdown, you don't need to trigger this manually. When you switch engines, JabRef tries the following commands, in order, and fills in the path field with the first one that works:

**For OCRmyPDF:**
**For Tesseract, EasyOCR, PaddleOCR, and AppleOCR** (all run through OCRmyPDF, so they share the same path):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Paddle OCR?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same source as the AppleOCR one above: JabRef/jabref#17219, merged 2026-09-26, adds PADDLEOCR to EngineSelection (identifier ocrmypdf_paddleocr, run via OCRmyPDF's PaddleOCR plugin). It's a real option in the current "OCR engine" dropdown alongside Tesseract, EasyOCR, and AppleOCR.

Let me know if you'd rather hold off documenting these three until they've had more real-world testing — happy to scope this PR down to just Tesseract/Docling if so.


Generated by Claude Code

1. `ocrmypdf`
2. `python -m ocrmypdf`
3. `py -m ocrmypdf`
Expand Down
6 changes: 6 additions & 0 deletions en/collect/add-entry-using-an-id.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,12 @@

![Screenshot of new entry dialog](../.gitbook/assets/newentrychoosetype-idgeneratorhighlighted-ads.png)

### Software Heritage

[Software Heritage](https://www.softwareheritage.org) is a universal archive that collects, preserves, and shares the source code of publicly available software.

ID search is carried out using a [SoftWare Heritage persistent IDentifier (SWHID)](https://www.softwareheritage.org/swhid/), e.g. `swh:1:dir:d198bc9d7a6bcf6db04f476d29314f157507d505`. Both a bare SWHID and a full `https://archive.softwareheritage.org/swh:1:...` URL are accepted.

Check warning on line 107 in en/collect/add-entry-using-an-id.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.Passive] 'are accepted' may be passive voice. Use active voice if you can. Raw Output: {"message": "[write-good.Passive] 'are accepted' may be passive voice. Use active voice if you can.", "location": {"path": "en/collect/add-entry-using-an-id.md", "range": {"start": {"line": 107, "column": 267}}}, "severity": "WARNING"}

Check warning on line 107 in en/collect/add-entry-using-an-id.md

View workflow job for this annotation

GitHub Actions / vale-lint

[vale] reported by reviewdog 🐶 [write-good.Passive] 'is carried' may be passive voice. Use active voice if you can. Raw Output: {"message": "[write-good.Passive] 'is carried' may be passive voice. Use active voice if you can.", "location": {"path": "en/collect/add-entry-using-an-id.md", "range": {"start": {"line": 107, "column": 11}}}, "severity": "WARNING"}

### Title

Based on the title of your publication, JabRef call Crossref, which return the corresponding DOI. Then JabRef fetches the reference based on this DOI.
Expand Down
8 changes: 8 additions & 0 deletions en/collect/import-using-online-bibliographic-database.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,10 @@ The [ACM Portal](https://dl.acm.org) includes two catalogs ([Wikipedia](https://

[ArXiv](https://arxiv.org) is a repository of scientific preprints in the fields of mathematics, physics, astronomy, computer science, quantitative biology, statistics, and quantitative finance ([Wikipedia](https://en.wikipedia.org/wiki/ArXiv)).

### BASE (Bielefeld Academic Search Engine)

[BASE](https://www.base-search.net) is one of the world's most voluminous search engines for academic open access web resources, operated by Bielefeld University Library ([Wikipedia](https://en.wikipedia.org/wiki/BASE_(search_engine))).

### Bibliotheksverbund Bayern (BVB)

The [Bibliotheksverbund Bayern (BVB)](https://www.bib-bvb.de) provides bibliographic information from all public libraries in Bavaria, Germany. The format used is [MarcXML](https://www.loc.gov/marc/bibliographic/), [which has been modified](https://www.bib-bvb.de/documents/10792/9f51a033-5ca1-42e2-b2d3-a75e7f1512d4), which in turn is [based on other modifications](https://www.dnb.de/marc21).
Expand Down Expand Up @@ -94,6 +98,10 @@ To fetch entries from Unpaywall indirectly through Crossref, choose **Search →

[DBLP](https://dblp.uni-trier.de/db/) is a computer science bibliography website listing more than 3.1 million journal articles, conference papers, and other publications on computer science ([Wikipedia](https://en.wikipedia.org/wiki/DBLP)).

### DNB

The [Deutsche Nationalbibliothek (DNB)](https://www.dnb.de), the German National Library, provides bibliographic data for German-language publications via its SRU interface. JabRef queries it using MARC XML records, the same format used by the Bibliotheksverbund Bayern (BVB) catalog above.

### DOAB

[DOAB (Directory of Open Access Books)](https://doabooks.org) is a community-driven discovery service that indexes and provides access to scholarly, peer-reviewed open access books and helps users to find trusted open access book publishers.
Expand Down
13 changes: 13 additions & 0 deletions en/jabkit.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,12 +81,25 @@ Commands:
generate-bib-from-aux Generate small bib from aux file.
preferences Manage JabKit preferences.
pdf Manage PDF metadata.
git Git integration for .bib files.
get-cited-works Get the cited works (bibliography).
get-citing-works Get the works citing the work at hand.
```

Hint: Using `jabkit <COMMAND> --help` will show the supported options for each command.

Any command that reads an input file also accepts a PostgreSQL connection URL of a JabRef shared library instead of a file path; JabKit exports the shared library read-only to a temporary `.bib` file before running the command.

### Git merge driver for `.bib` files

`jabkit git merge-driver` performs a semantic three-way merge of `.bib` files, so that concurrent changes to different entries (or different fields) merge automatically instead of producing textual Git conflicts. To use it for every `.bib` file in a repository:

```bash
git config --global merge.jabref.name "JabRef semantic .bib merge"
git config --global merge.jabref.driver "jabkit git merge-driver --porcelain %O %A %B"
echo "*.bib merge=jabref" >> .gitattributes
```

## Updating JabKit

Make use of `--fresh` to update JabKit
Expand Down
Loading