A Python tool for retrieving academic abstracts from 9 different APIs including PubMed, OpenAlex, arXiv, Semantic Scholar, and more.
git clone <repository-url>
cd information-retrieval
pip install -r requirements.txtCopy env_template.txt to .env and add your keys:
cp env_template.txt .envRequired: NCBI API key (get one here)
Highly Recommended: Semantic Scholar key (get one here) - 50x faster!
python run.py# All sources (recommended)
python run.py
# Only fast APIs (skip Semantic Scholar)
python run.py --phases 1 2
# Start fresh (ignore existing abstracts)
python run.py --fresh
# Custom input/output
python src/fetch_abstracts.py --input my_pubs.xlsx --output results.json- Input:
data/input/publications.xlsx(must have 'title' column) - Output:
data/output/alternative_abstracts.json - Failed:
data/output/alternative_failed.json
- 9 Data Sources - Comprehensive coverage across disciplines
- Smart 3-Phase Search - Fast APIs first, specialized sources second, rate-limited last
- Auto Retry - Handles rate limits automatically (up to 5 retries with progressive backoff)
- Progress Tracking - Real-time updates and source statistics
- High Success Rate - 85-95% coverage across all sources
- Modular Architecture - Easy to extend with new sources
- Europe PMC - 40M+ biomedical papers, fast and reliable, no key needed
- OpenAlex - 200M+ open access works, all disciplines, unlimited rate
- CrossRef - 130M+ DOIs, publisher metadata, polite pool
- PubMed - 35M+ biomedical citations, requires NCBI API key
- arXiv - 2M+ preprints (physics, CS, math, biology), always has abstracts
- CORE - 200M+ open access papers, optional API key recommended
- bioRxiv - Biology preprints, free API
- medRxiv - Medical preprints, free API
- Semantic Scholar - 200M+ papers, AI-powered search, API key highly recommended (50x faster)
| Key | Required? | Rate Limit | Get It |
|---|---|---|---|
| NCBI | ✅ Required | 3→10 req/s | Link |
| Semantic Scholar | 🌟 Highly Recommended | 100→5,000 req/5min | Link |
| CORE | Optional | Limited→10 req/s | Link |
| Unpaywall Email | Optional | Any valid email | Link |
Without Semantic Scholar API Key:
- 100 requests per 5 minutes (~1 every 3 seconds)
- Processing 1,000 documents: ~58 minutes
With Semantic Scholar API Key:
- 5,000 requests per 5 minutes (~16/second)
- Processing 1,000 documents: ~16 minutes
- Time saved: ~72% faster!
-
NCBI API Key (Required):
- Create free account at https://www.ncbi.nlm.nih.gov/account/settings/
- Generate API key in Settings → API Key Management
- Add to
.env:NCBI_API_KEY=your_key_here
-
Semantic Scholar API Key (Recommended):
- Go to https://www.semanticscholar.org/product/api
- Request API access (usually instant for academic emails)
- Add to
.env:SEMANTIC_SCHOLAR_API_KEY=your_key_here
-
CORE API Key (Optional):
- Register at https://core.ac.uk/services/api
- Add to
.env:CORE_API_KEY=your_key_here
-
Unpaywall Email (Optional):
- Just needs any valid email
- Add to
.env:UNPAYWALL_EMAIL=your@email.com
{
"12345678": {
"title": "Publication Title",
"abstract": "Abstract text...",
"source": "Europe PMC"
}
}{
"total_failed": 5,
"total_processed": 100,
"success_rate": "95.00%",
"failed_documents": [
{
"index": 10,
"title": "Publication Title",
"error": "Abstract not found in any source"
}
]
}Edit src/config.py to customize:
- API URLs and timeouts
- Rate limiting delays (default: 0.5s general, 1.0s Semantic Scholar)
- Retry attempts (default: 5 retries)
- File paths
- Backoff multipliers
- Create a new fetcher class in
src/fetchers/:
from .base import AbstractFetcher
class MyNewSourceFetcher(AbstractFetcher):
def __init__(self):
super().__init__(name="MyNewSource", delay=0.5)
def fetch_by_title(self, title: str):
# Implementation
return (identifier, abstract)- Add to
abstract_retriever.pyin appropriate phase - Done!
Before training, use the inspection utility to verify your data:
python inspect_data.pyThis will show you:
- Number of documents in each file
- Document structure and fields
- Text length statistics
- Sample documents
- Dataset balance
- Most common words
python transformer_encoder.pyThis will:
- Load data from
relevant_documents.jsonandnon_relevant_documents.json - Split into train (70%), validation (15%), test (15%)
- Build vocabulary from training data
- Check if there is any model already trained or train the Transformer Encoder for 15 epochs
- Evaluate on the test set
- Save the trained model to
transformer_ir_model.pth
If you want to check the predictions with any custom text that you want to input to the model, you can do:
python inference.pyThis way you can check the confidence on any text input that you use and check if the predictions are okay!
If you want to check the evaluation metrics in a visual way, just run:
python visualizations.pyAll the images containing the different evaluations will be stored in a folder called figures.