From aca5b86650cc1dccc45e81449323599d3aed2e0f Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 08:22:30 +0000 Subject: [PATCH] docs: catch up with jabref changes from the past month - jabkit: document the new `git merge-driver` command and the new support for passing a shared-database (PostgreSQL) URL as input (JabRef/jabref#16838, JabRef/jabref#16970) - OCR: document EasyOCR, PaddleOCR, and AppleOCR as new selectable OCRmyPDF backends, alongside Tesseract and Docling (JabRef/jabref#17219) - Add the BASE (Bielefeld Academic Search Engine) and DNB (Deutsche Nationalbibliothek) catalogs to the online search fetcher list (JabRef/jabref#16530, JabRef/jabref#17070) - Add the Software Heritage (SWHID) fetcher to the "Add entry using an ID" page (JabRef/jabref#16849) Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01XtzicSVFTs29DBkudE3PrX --- en/advanced/OCR.md | 9 +++++---- en/collect/add-entry-using-an-id.md | 6 ++++++ .../import-using-online-bibliographic-database.md | 8 ++++++++ en/jabkit.md | 13 +++++++++++++ 4 files changed, 32 insertions(+), 4 deletions(-) diff --git a/en/advanced/OCR.md b/en/advanced/OCR.md index bdecef347..11d994483 100644 --- a/en/advanced/OCR.md +++ b/en/advanced/OCR.md @@ -1,6 +1,6 @@ # OCR -[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports two OCR engines: [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) and [Docling](https://github.com/docling-project/docling). +[OCR](https://en.wikipedia.org/wiki/Optical_character_recognition) (Optical Character Recognition) is defined as the electronic or mechanic conversion of images of typed, handwritten or printed text into machine-encoded text. Consequently, with this technology it is possible to add editable and searchable data to PDFs and other files in your Jabref library. OCR can be used via multiple tools and engines. Currently, JabRef supports [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/) (with a selectable OCR backend: Tesseract, EasyOCR, PaddleOCR, or AppleOCR) and [Docling](https://github.com/docling-project/docling). ## How to install an OCR engine @@ -48,8 +48,9 @@ The OCR engine selected in your preferences must be installed on your system to * Available engines: -1. **OCRmyPDF**: the default engine. Well suited for general-purpose OCR on scanned PDFs. -2. **Docling**: an alternative engine with strong handling of complex layouts and documents containing tables or figures (slower than OCRmyPDF). +1. **Tesseract**: the default engine. Runs via OCRmyPDF and is well suited for general-purpose OCR on scanned PDFs. +2. **EasyOCR**, **PaddleOCR**, **AppleOCR**: alternative OCR backends, also run via OCRmyPDF through its respective plugin. Each plugin must be installed separately in addition to OCRmyPDF itself. +3. **Docling**: a separate engine with strong handling of complex layouts and documents containing tables or figures (slower than the OCRmyPDF-based engines). * Changing the selected engine automatically re-runs **auto-detection** for that engine's path (see below), so if the newly selected engine is installed in a standard location, its path field will populate automatically. @@ -74,7 +75,7 @@ Performing OCR will fail if wrong engine path is provided, make sure that the co * JabRef also **auto-detects the path automatically** whenever you change the selected engine in the dropdown, you don't need to trigger this manually. When you switch engines, JabRef tries the following commands, in order, and fills in the path field with the first one that works: - **For OCRmyPDF:** + **For Tesseract, EasyOCR, PaddleOCR, and AppleOCR** (all run through OCRmyPDF, so they share the same path): 1. `ocrmypdf` 2. `python -m ocrmypdf` 3. `py -m ocrmypdf` diff --git a/en/collect/add-entry-using-an-id.md b/en/collect/add-entry-using-an-id.md index 475957af7..5e1ab57c1 100644 --- a/en/collect/add-entry-using-an-id.md +++ b/en/collect/add-entry-using-an-id.md @@ -100,6 +100,12 @@ ID search is carried out using the [ADS Bibcode](http://adsabs.github.io/help/ac ![Screenshot of new entry dialog](../.gitbook/assets/newentrychoosetype-idgeneratorhighlighted-ads.png) +### Software Heritage + +[Software Heritage](https://www.softwareheritage.org) is a universal archive that collects, preserves, and shares the source code of publicly available software. + +ID search is carried out using a [SoftWare Heritage persistent IDentifier (SWHID)](https://www.softwareheritage.org/swhid/), e.g. `swh:1:dir:d198bc9d7a6bcf6db04f476d29314f157507d505`. Both a bare SWHID and a full `https://archive.softwareheritage.org/swh:1:...` URL are accepted. + ### Title Based on the title of your publication, JabRef call Crossref, which return the corresponding DOI. Then JabRef fetches the reference based on this DOI. diff --git a/en/collect/import-using-online-bibliographic-database.md b/en/collect/import-using-online-bibliographic-database.md index 1c4a74cd2..ff02fdafe 100644 --- a/en/collect/import-using-online-bibliographic-database.md +++ b/en/collect/import-using-online-bibliographic-database.md @@ -58,6 +58,10 @@ The [ACM Portal](https://dl.acm.org) includes two catalogs ([Wikipedia](https:// [ArXiv](https://arxiv.org) is a repository of scientific preprints in the fields of mathematics, physics, astronomy, computer science, quantitative biology, statistics, and quantitative finance ([Wikipedia](https://en.wikipedia.org/wiki/ArXiv)). +### BASE (Bielefeld Academic Search Engine) + +[BASE](https://www.base-search.net) is one of the world's most voluminous search engines for academic open access web resources, operated by Bielefeld University Library ([Wikipedia](https://en.wikipedia.org/wiki/BASE_(search_engine))). + ### Bibliotheksverbund Bayern (BVB) The [Bibliotheksverbund Bayern (BVB)](https://www.bib-bvb.de) provides bibliographic information from all public libraries in Bavaria, Germany. The format used is [MarcXML](https://www.loc.gov/marc/bibliographic/), [which has been modified](https://www.bib-bvb.de/documents/10792/9f51a033-5ca1-42e2-b2d3-a75e7f1512d4), which in turn is [based on other modifications](https://www.dnb.de/marc21). @@ -94,6 +98,10 @@ To fetch entries from Unpaywall indirectly through Crossref, choose **Search → [DBLP](https://dblp.uni-trier.de/db/) is a computer science bibliography website listing more than 3.1 million journal articles, conference papers, and other publications on computer science ([Wikipedia](https://en.wikipedia.org/wiki/DBLP)). +### DNB + +The [Deutsche Nationalbibliothek (DNB)](https://www.dnb.de), the German National Library, provides bibliographic data for German-language publications via its SRU interface. JabRef queries it using MARC XML records, the same format used by the Bibliotheksverbund Bayern (BVB) catalog above. + ### DOAB [DOAB (Directory of Open Access Books)](https://doabooks.org) is a community-driven discovery service that indexes and provides access to scholarly, peer-reviewed open access books and helps users to find trusted open access book publishers. diff --git a/en/jabkit.md b/en/jabkit.md index 1607eea23..92f5d5f4d 100644 --- a/en/jabkit.md +++ b/en/jabkit.md @@ -81,12 +81,25 @@ Commands: generate-bib-from-aux Generate small bib from aux file. preferences Manage JabKit preferences. pdf Manage PDF metadata. + git Git integration for .bib files. get-cited-works Get the cited works (bibliography). get-citing-works Get the works citing the work at hand. ``` Hint: Using `jabkit --help` will show the supported options for each command. +Any command that reads an input file also accepts a PostgreSQL connection URL of a JabRef shared library instead of a file path; JabKit exports the shared library read-only to a temporary `.bib` file before running the command. + +### Git merge driver for `.bib` files + +`jabkit git merge-driver` performs a semantic three-way merge of `.bib` files, so that concurrent changes to different entries (or different fields) merge automatically instead of producing textual Git conflicts. To use it for every `.bib` file in a repository: + +```bash +git config --global merge.jabref.name "JabRef semantic .bib merge" +git config --global merge.jabref.driver "jabkit git merge-driver --porcelain %O %A %B" +echo "*.bib merge=jabref" >> .gitattributes +``` + ## Updating JabKit Make use of `--fresh` to update JabKit