-
Notifications
You must be signed in to change notification settings - Fork 145
docs: catch up with jabref changes from the past month (automated monthly review) #696
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -1,6 +1,6 @@ | ||
| # OCR | ||
|
|
||
| [OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports two OCR engines: [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) and [Docling](https://github.com/docling-project/docling). | ||
| [OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) (with a selectable OCR backend: Tesseract, EasyOCR, PaddleOCR, or AppleOCR) and [Docling](https://github.com/docling-project/docling). | ||
|
Check warning on line 3 in en/advanced/OCR.md
|
||
|
|
||
| ## How to install an OCR engine | ||
|
|
||
|
|
@@ -48,8 +48,9 @@ | |
|
|
||
| * Available engines: | ||
|
|
||
| 1. **OCRmyPDF**: the default engine. Well suited for general-purpose OCR on scanned PDFs. | ||
| 2. **Docling**: an alternative engine with strong handling of complex layouts and documents containing tables or figures (slower than OCRmyPDF). | ||
| 1. **Tesseract**: the default engine. Runs via OCRmyPDF and is well suited for general-purpose OCR on scanned PDFs. | ||
| 2. **EasyOCR**, **PaddleOCR**, **AppleOCR**: alternative OCR backends, also run via OCRmyPDF through its respective plugin. Each plugin must be installed separately in addition to OCRmyPDF itself. | ||
|
Check warning on line 52 in en/advanced/OCR.md
|
||
| 3. **Docling**: a separate engine with strong handling of complex layouts and documents containing tables or figures (slower than the OCRmyPDF-based engines). | ||
|
|
||
| * Changing the selected engine automatically re-runs **auto-detection** for that engine's path (see below), so if the newly selected engine is installed in a standard location, its path field will populate automatically. | ||
|
|
||
|
|
@@ -74,7 +75,7 @@ | |
|
|
||
| * JabRef also **auto-detects the path automatically** whenever you change the selected engine in the dropdown, you don't need to trigger this manually. When you switch engines, JabRef tries the following commands, in order, and fills in the path field with the first one that works: | ||
|
|
||
| **For OCRmyPDF:** | ||
| **For Tesseract, EasyOCR, PaddleOCR, and AppleOCR** (all run through OCRmyPDF, so they share the same path): | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Paddle OCR?
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Same source as the AppleOCR one above: JabRef/jabref#17219, merged 2026-09-26, adds Let me know if you'd rather hold off documenting these three until they've had more real-world testing — happy to scope this PR down to just Tesseract/Docling if so. Generated by Claude Code |
||
| 1. `ocrmypdf` | ||
| 2. `python -m ocrmypdf` | ||
| 3. `py -m ocrmypdf` | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
apple OCR? wtf?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Not a joke — it's real as of 5 days ago: JabRef/jabref#17219 ("Add OCRmyPDF plugins as new engines support (EasyOCR, PaddleOCR, AppleOCR)") added
APPLEOCRas a selectable value inEngineSelection, alongsideTESSERACT,EASYOCR, andPADDLEOCR, wired intoOcrEngineFactoryto run through OCRmyPDF'sappleocrplugin. It shows up in the "OCR engine" dropdown in File → Preferences → OCR on currentmain.Happy to drop it from this PR if it's considered out of scope / not ready for users yet — let me know and I'll push an update.
Generated by Claude Code