Skip to content

docs: catch up with jabref changes from the past month (automated monthly review) - #696

Open
koppor wants to merge 1 commit into
mainfrom
docs/monthly-review-2026-10-01
Open

koppor wants to merge 1 commit into
mainfrom
docs/monthly-review-2026-10-01

Conversation

@koppor

@koppor koppor commented Oct 1, 2026

Copy link
Copy Markdown
Member

Pull Request Description

Summary

Automated monthly review of the JabRef user-documentation against user-facing changes merged into JabRef/jabref main over the past month (2026-09-01 to 2026-09-30). This is a low-risk, docs-only update — no code changes. Four pages were updated to cover features that shipped recently but were missing from the docs:

  • en/jabkit.md — added the new jabkit git merge-driver command and the new support for passing a PostgreSQL shared-database URL as input to any jabkit command.
  • en/advanced/OCR.md — updated the OCR engine list: the "OCR engine" dropdown now offers Tesseract, EasyOCR, PaddleOCR, and AppleOCR (all via OCRmyPDF) in addition to Docling, not just "OCRmyPDF" and "Docling".
  • en/collect/import-using-online-bibliographic-database.md — added the two new online-search catalogs: BASE (Bielefeld Academic Search Engine) and DNB (Deutsche Nationalbibliothek).
  • en/collect/add-entry-using-an-id.md — added the new Software Heritage (SWHID) identifier-based fetcher.

Each change was verified directly against the current jabref source (enum/class names, CLI command wiring) rather than only against the CHANGELOG, to avoid documenting anything speculative.

Steps to test

  1. Check the preview link in this PR for each changed page.
  2. en/jabkit.md: compare against jabkit --help output on a current main build of jabkit — it should list a git subcommand, and jabkit git merge-driver --help should match the documented setup snippet.
  3. en/advanced/OCR.md: open File → Preferences → OCR in a current build and confirm the OCR engine dropdown shows Tesseract, EasyOCR, PaddleOCR, AppleOCR, and Docling.
  4. en/collect/import-using-online-bibliographic-database.md: open View → Web search and confirm "BASE" and "DNB" appear in the catalog dropdown.
  5. en/collect/add-entry-using-an-id.md: open Library → New entry and confirm "Software Heritage" appears as an ID type.

Related issues and pull requests

Closes _____

Documents JabRef/jabref#16838 (git merge-driver), JabRef/jabref#16970 (shared-database URL as jabkit input), JabRef/jabref#17219 (EasyOCR/PaddleOCR/AppleOCR engines), JabRef/jabref#16530 (BASE fetcher), JabRef/jabref#17070 (DNB fetcher), JabRef/jabref#16849 (Software Heritage/SWHID fetcher)

AI usage

Claude Code (model claude-sonnet-5), running as a scheduled automated monthly documentation review. All factual claims in this PR were cross-checked against the jabref source code (fetcher classes, CLI command registration, OCR engine enum) rather than taken only from the CHANGELOG or commit messages.

Checklist

  • I reviewed and take ownership of all content in this PR, including any AI-assisted text — this is an automated PR; please give it a human review before merging
  • I opened JabRef and followed these instructions myself to confirm they still work (tested on version: _____) — not done; verified against source code only, a maintainer should spot-check the live UI
  • Any links I added or changed work (no 404s) — new external links point to well-known existing sites (base-search.net, dnb.de, softwareheritage.org); not run through an automated link checker in this session
  • Any images or tables I added display correctly — no new images or tables were added
  • I checked spelling/grammar

Given this is a documentation-only change describing already-shipped, verified behavior, it should be low risk to review and merge quickly.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XtzicSVFTs29DBkudE3PrX


Generated by Claude Code

- jabkit: document the new `git merge-driver` command and the new
  support for passing a shared-database (PostgreSQL) URL as input
  (JabRef/jabref#16838, JabRef/jabref#16970)
- OCR: document EasyOCR, PaddleOCR, and AppleOCR as new selectable
  OCRmyPDF backends, alongside Tesseract and Docling
  (JabRef/jabref#17219)
- Add the BASE (Bielefeld Academic Search Engine) and DNB (Deutsche
  Nationalbibliothek) catalogs to the online search fetcher list
  (JabRef/jabref#16530, JabRef/jabref#17070)
- Add the Software Heritage (SWHID) fetcher to the "Add entry using
  an ID" page (JabRef/jabref#16849)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XtzicSVFTs29DBkudE3PrX
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0)

Grey Divider

Great, no issues found!

Qodo reviewed your code and found no material issues that require review

Grey Divider

Tip of the day
💡 Did you know, you can keep summaries lean with Findings visible per group, which tucks the rest behind a View link

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Document recent JabKit, OCR, and bibliography lookup features

📝 Documentation 🕐 10-20 Minutes

Grey Divider

AI Description

• Document JabKit’s Git merge driver and PostgreSQL shared-library URL inputs.
• Clarify selectable OCR backends and their installation and path requirements.
• Add BASE, DNB, and Software Heritage to the relevant lookup guides.
Diagram

graph TD
  J["JabKit guide"] --> C(["CLI capabilities"])
  O["OCR guide"] --> E(["OCR backends"])
  S["Search guide"] --> W(["Web catalogs"])
  I["ID guide"] --> F(["SWHID lookup"])
Loading
High-Level Assessment

Adding these capabilities to the existing task-specific guides keeps them where readers will look. Separate pages would fragment short additions without improving discoverability; a maintainer should still spot-check the UI and CLI examples before merging.

Files changed (4) +32 / -4

Documentation (4) +32 / -4
OCR.mdClarify OCR engines and selectable backends +5/-4

Clarify OCR engines and selectable backends

• Replaces the two-engine description with Tesseract, EasyOCR, PaddleOCR, and AppleOCR as OCRmyPDF-backed choices alongside Docling. Clarifies plugin installation and the shared OCRmyPDF executable path used for detection.

en/advanced/OCR.md

add-entry-using-an-id.mdDocument Software Heritage identifier lookup +6/-0

Document Software Heritage identifier lookup

• Adds Software Heritage to the identifier-based entry guide, with an example SWHID and the accepted bare-identifier and archive-URL forms.

en/collect/add-entry-using-an-id.md

import-using-online-bibliographic-database.mdAdd BASE and DNB to the online catalog guide +8/-0

Add BASE and DNB to the online catalog guide

• Adds descriptions of BASE and the Deutsche Nationalbibliothek to the supported web-search catalogs, including DNB’s SRU and MARC XML context.

en/collect/import-using-online-bibliographic-database.md

jabkit.mdDocument JabKit merge driver and shared-library inputs +13/-0

Document JabKit merge driver and shared-library inputs

• Adds the Git command to the command list and provides a '.bib' merge-driver setup example. Explains that commands reading input files can also accept a PostgreSQL URL for a JabRef shared library.

en/jabkit.md

Comment thread en/advanced/OCR.md
# OCR

[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports two OCR engines: [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) and [Docling](https://github.com/docling-project/docling).
[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) (with a selectable OCR backend: Tesseract, EasyOCR, PaddleOCR, or AppleOCR) and [Docling](https://github.com/docling-project/docling).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

apple OCR? wtf?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a joke — it's real as of 5 days ago: JabRef/jabref#17219 ("Add OCRmyPDF plugins as new engines support (EasyOCR, PaddleOCR, AppleOCR)") added APPLEOCR as a selectable value in EngineSelection, alongside TESSERACT, EASYOCR, and PADDLEOCR, wired into OcrEngineFactory to run through OCRmyPDF's appleocr plugin. It shows up in the "OCR engine" dropdown in File → Preferences → OCR on current main.

Happy to drop it from this PR if it's considered out of scope / not ready for users yet — let me know and I'll push an update.


Generated by Claude Code

Comment thread en/advanced/OCR.md
* JabRef also **auto-detects the path automatically** whenever you change the selected engine in the dropdown, you don't need to trigger this manually. When you switch engines, JabRef tries the following commands, in order, and fills in the path field with the first one that works:

**For OCRmyPDF:**
**For Tesseract, EasyOCR, PaddleOCR, and AppleOCR** (all run through OCRmyPDF, so they share the same path):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Paddle OCR?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same source as the AppleOCR one above: JabRef/jabref#17219, merged 2026-09-26, adds PADDLEOCR to EngineSelection (identifier ocrmypdf_paddleocr, run via OCRmyPDF's PaddleOCR plugin). It's a real option in the current "OCR engine" dropdown alongside Tesseract, EasyOCR, and AppleOCR.

Let me know if you'd rather hold off documenting these three until they've had more real-world testing — happy to scope this PR down to just Tesseract/Docling if so.


Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants