An end-to-end, local-first Agentic Retrieval-Augmented Generation (RAG) system built for enterprise reliability, absolute data privacy, and deterministic grounding. Designed specifically for deep financial filings, corporate annual reports, 10-Ks, and sensitive internal documentation, this system executes 100% locally with Zero Data Leakage.
Enterprise documentsβsuch as earnings reports before public disclosure, board minutes, M&A filings, and healthcare recordsβcannot be routed through third-party cloud inference APIs (e.g., OpenAI, Anthropic, or cloud providers) without violating strict data privacy regulations (GDPR, HIPAA, SOC 2 Type II) and non-disclosure agreements.
This repository implements a deliberate, uncompromising architectural shift to a 100% local, air-gapped pipeline:
| Component | Technology | Enterprise Role |
|---|---|---|
| Local Orchestration | CrewAI | Multi-agent deliberation, goal decomposition, and dynamic tool selection on local threads. |
| Persistent Retrieval | ChromaDB | On-disk local vector index (./company_db) with automatic SQLite FTS5 fallback. |
| Local Inference | Ollama | Local execution of open-weight foundation models (ollama/llama3.2, llama3.1, mistral). |
| Local Embeddings | Sentence-Transformers | all-MiniLM-L6-v2 dense vector embeddings (384 dimensions) generated in-process on CPU. |
| Zero Cloud Keys | Sanitized Environment | All external API keys and cloud dependencies stripped to eliminate outbound data leakage. |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β 100% AIR-GAPPED LOCAL PROCESS BOUNDARY β
β β
ββββββββββββββββββββββββ β ββββββββββββββββββββββ ββββββββββββββββββββββββ β
β Enterprise Report β ββββ> β β DocumentLoader β βββ> β DocumentSplitter β β
β (PDF / TXT / 10-K) β β β (pypdf / metadata) β β (800 char / 150 ovlp)β β
ββββββββββββββββββββββββ β ββββββββββββββββββββββ ββββββββββββββββββββββββ β
β β β
β βΌ β
ββββββββββββββββββββββββ β ββββββββββββββββββββββ ββββββββββββββββββββββββ β
β Glass Box Live Audit β <ββββ β β CrewAI Agent β <βββ β ChromaDB (Local) β β
β (Streamlit / UI) β β β (Ollama Llama 3.2) β β (all-MiniLM-L6-v2) β β
ββββββββββββββββββββββββ β ββββββββββββββββββββββ ββββββββββββββββββββββββ β
β β² β β
β β (Circuit Breaker) β (Fallback) β
β βββββββ SQLite FTS5 ββββββββββ β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Production RAG systems fail when they encounter missing context, vector database latency spikes, or hallucination vulnerabilities. The system enforces a Three-Tier Reliability Guardrail architecture (src/agentic_rag/guardrails.py):
User Query
β
βΌ
[ Tier 1: Input Guardrail ] βββββββββ> [ BLOCKED ] (Adversarial Injection / Empty)
β Passed
βΌ
[ Tier 2: Retrieval Guardrail ]
βββ Vector DB Timeout? ββββββββββ> [ CIRCUIT BREAKER ] ββ> SQLite FTS5 Fallback
β
βββ Similarity < Threshold (0.65)?
β
βββ Yes βββββββββββββββββ> [ DISCIPLINED REFUSAL ] (Missing Context Detected)
βββ No βββββββββββββββββ> High-Confidence Evidence Chunks
β
βΌ
[ Tier 3: Output Guardrail ] ββββββββ> [ HALLUCINATION AUDITOR ] ββ> Verified Answer
- Adversarial Pattern Interception: Detects and neutralizes prompt injections, system override signatures (
ignore previous instructions,DAN mode,developer mode), and malicious inputs before token transmission. - Query Normalization: Validates length, strips control tokens, and deconstructs question intent.
- Dynamic Confidence Gating: Calculates cosine similarity for all retrieved chunks. If the top similarity score falls below the configurable confidence gate (default: 0.65), the system refuses to speculate.
- Missing-Context Refusal Protocol: When an out-of-domain query is detected (e.g., asking about Apple or Tesla in a TCS annual report), the pipeline immediately triggers an explicit refusal stating that document coverage is absent, preventing factual hallucination.
- Vector DB Timeout Circuit Breaker: ChromaDB vector operations are wrapped with a timeout threshold (default: 3.0s). If the vector store hangs or fails, the pipeline automatically falls back to local SQLite FTS5 full-text keyword search without crashing.
- Numerical Metric Extraction: Automatically extracts all numbers, currency figures (
$,βΉ), percentages (%), and financial metrics from the agent's output. - Grounding Cross-Check: Verifies each extracted figure against the verbatim text in the retrieved chunks.
- Hallucination Risk Score: Computes a real-time risk metric (
0.0%for fully grounded outputs) and enforces page-level citations ([Source X β’ Page Y]).
Rather than hiding agent decisions inside a black box, the system provides a Glass Box Live Audit Console (built with Streamlit) that exposes the agent's internal state in real time:
- Step-by-Step Reasoning Traces (Chain of Thought): Live stage-by-stage inspection showing intent deconstruction, tool selection rationale, search parameter formulation, and metric verification.
- ChromaDB Similarity Scores & Distance Inspector: Every retrieved chunk displays its exact cosine distance, normalized similarity percentage, chunk ID, page number, and comparison to the confidence threshold.
- Real-Time Fallback Trigger Monitor: Prominent visual status alerts show whether the query executed nominally, triggered a low-confidence refusal, engaged the SQLite FTS5 circuit breaker, or was quarantined by input guardrails.
- Interactive Reliability Test Presets:
- π’ Grounded In-Domain Audit: Runs a standard financial inquiry (Operating Margin & RoE) with high similarity (
> 0.85), passing all guardrails. - π‘ Missing Context Fallback: Runs an out-of-domain query (e.g., Apple iPhone revenue in a TCS filing), demonstrating automatic similarity drop (
~0.42 < 0.65) and disciplined refusal. - π Vector Timeout Circuit Breaker: Simulates a ChromaDB timeout to verify zero-downtime fallback to SQLite FTS5.
- π΄ Adversarial Injection Block: Demonstrates instant prompt injection quarantine.
- π’ Grounded In-Domain Audit: Runs a standard financial inquiry (Operating Margin & RoE) with high similarity (
- Operating System: macOS, Linux, or Windows (WSL2 recommended)
- Python: 3.10, 3.11, 3.12, or 3.13
- Ollama: Installed locally (ollama.com)
Download and install Ollama from ollama.com/download.
# Verify Ollama installation
ollama --version
# Start the Ollama daemon (runs on http://localhost:11434)
ollama serve
# In another terminal, pull the recommended lightweight model
ollama pull llama3.2(Optional: You can also use other local models such as ollama pull llama3.1 or ollama pull mistral).
git clone https://github.com/JohnnyWilson16/agentic-rag-system.git
cd agentic-rag-system
# Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Install dependencies
pip install -r requirements.txtCopy the template to .env. Note that no cloud API keys are required or supported:
cp .env.example .envYour .env file should contain the local-only configuration:
LLM_PROVIDER=ollama
LLM_MODEL=ollama/llama3.2
OLLAMA_BASE_URL=http://localhost:11434
CONFIDENCE_THRESHOLD=0.65
VECTOR_DB_TIMEOUT_SECONDS=3.0
ENABLE_GUARDRAILS=true
ZERO_DATA_LEAKAGE_ENFORCED=true
EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L6-v2
VECTOR_STORE_DIR=./company_db
RETRIEVER_K=3Index the included sample corporate report or your own PDF:
# Ingest sample report
python -m agentic_rag.cli ingest data/sample_annual_report.txt --reset
# Ingest a PDF document
python -m agentic_rag.cli ingest data/annual_report_2025_2026.pdf --resetOutput:
Document Ingestion Summary
βββββββββββββββββββββββ³ββββββββββββββββββββββββββββββββββββββ
β Property β Value β
β‘ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ©
β Document Name β annual_report_2025_2026.pdf β
β Raw Pages / Sectionsβ 360 β
β Chunks Created β 1,818 β
β Total Vectors in DB β 1,818 β
β Embedding Model β sentence-transformers/all-MiniLM-L6 β
β Persist Directory β ./company_db β
βββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββββββββ
β Ingestion completed successfully.
Launch the interactive audit interface with live telemetry, fallback triggers, and confidence controls:
./run_demo.sh
# Or directly:
streamlit run app.pyOpen http://localhost:8501 in your browser.
Launch the FastAPI backend serving the single-page application:
./run_app.sh
# Or directly:
python -m uvicorn agentic_rag.api:app --host 127.0.0.1 --port 8000Open http://127.0.0.1:8000 in your browser.
Run queries directly from your terminal:
# Targeted query
python -m agentic_rag.cli query "What was the operating margin and Return on Equity for FY 2026?"
# Full multi-section audit
python -m agentic_rag.cli analyze --company "Tata Consultancy Services"
# Check system status
python -m agentic_rag.cli statusagentic-rag-system/
βββ app.py # Streamlit Glass Box Live Audit Console
βββ main.py # ASGI server root entrypoint
βββ run_demo.sh # One-click launcher for Streamlit Audit Console
βββ run_app.sh # One-click launcher for Web Console & FastAPI
βββ data/
β βββ sample_annual_report.txt # Sample corporate financial text document
β βββ annual_report_2025_2026.pdf # Comprehensive 360-page financial report PDF
βββ scripts/
β βββ create_sample_pdf.py # Utility to generate structured benchmark PDFs
β βββ run_analysis.py # Quickstart offline analysis script
βββ src/
β βββ agentic_rag/
β βββ __init__.py # Package initialization and exports
β βββ agents.py # CrewAI agent definitions & grounding prompts
β βββ api.py # FastAPI REST API & session manager
β βββ cli.py # Command-line interface with Rich formatting
β βββ config.py # Pydantic settings with Zero Data Leakage enforcer
β βββ demo_engine.py # Glass-box execution engine with guardrails telemetry
β βββ document_loader.py # PDF/TXT loader with page metadata extraction
β βββ embeddings.py # Low-memory ONNX / CPU sentence embeddings
β βββ guardrails.py # Three-Tier Reliability Guardrails & Enforcer
β βββ llm_factory.py # CrewAI local Ollama LLM constructor
β βββ pipeline.py # End-to-End Agentic RAG pipeline coordinator
β βββ tasks.py # Structured analytical task definitions
β βββ text_splitter.py # Recursive semantic chunking (800 chars / 150 ovlp)
β βββ tools.py # CrewAI tool wrapper for vector retrieval
β βββ vector_store.py # ChromaDB persistence & SQLite FTS5 fallback
β βββ static/ # Web Console SPA (HTML, CSS, JS)
βββ tests/
β βββ conftest.py # Pytest fixtures and environment configuration
β βββ test_api.py # API endpoint and health tests
β βββ test_config.py # Configuration & Zero Data Leakage tests
β βββ test_document_processing.py # PDF loading and chunking tests
β βββ test_guardrails.py # Comprehensive tests for three-tier guardrails
β βββ test_pipeline.py # Pipeline orchestration & CrewAI mock tests
β βββ test_tools_and_agent.py # Search tool formatting and task tests
β βββ test_vector_store.py # ChromaDB indexing & similarity search tests
βββ .env.example # Environment template for local Ollama
βββ LICENSE # MIT License
βββ pyproject.toml # Build metadata and package dependencies
βββ requirements.txt # Frozen dependency specifications
Execute the comprehensive test suite with pytest:
pytest tests/ -vThe test suite covers:
- Zero Data Leakage: Verifies that cloud API keys are stripped and external URLs are rejected.
- Input Guardrail: Tests valid queries, length bounds, and prompt injection signatures.
- Retrieval Guardrail: Tests confidence gating, missing-context fallback, and vector timeout circuit breaker.
- Output Guardrail: Validates numerical metric cross-referencing and hallucination risk calculation.
- Vector Store & Ingestion: Tests chunking, embedding generation, ChromaDB persistence, and SQLite FTS5 fallback.
- Pipeline & API: Verifies CrewAI tool binding, task coordination, and REST endpoint contracts.
MIT License. See LICENSE for details.