A production-ready, full-stack framework for building intelligent RAG applications with an interactive web interface.
This framework provides everything you need to build sophisticated AI agents that can:
- Retrieve Information: Use semantic vector search (RAG) to find relevant documents
- Make Intelligent Decisions: Automatically determine when to search vs. answer directly
- Maintain Context: Persist conversation history across sessions
- Interactive UI: Modern, responsive web interface with smooth typing animations
- Flexible Deployment: Supports OpenAI, Azure OpenAI, or local models (LM Studio)
- π€ Intelligent Tool Use: LLM calls
search_docsonly when necessary - π― Semantic Filtering: Relevance threshold (β₯0.5) removes irrelevant results
- π Optimized Retrieval: Configurable TOP_K and chunk sizing for accuracy
- π Session Memory: Persists chat history to disk
- π Multi-Provider: OpenAI, Azure OpenAI, or LM Studio (local)
- β‘ FastAPI Server: Production-ready REST API
- π¨ Modern UI: Dark theme with glassmorphism effects
- β¨ Smooth Typing Animation: ChatGPT-like letter-by-letter responses
- π Copy-to-Clipboard: Easy code snippet copying
- π« Interactive Animations: Hover effects, loading states, transitions
- π Source Citations: Clear attribution for all retrieved information
- π― Real-time Feedback: Shake validation, pulse effects, status indicators
Agentic-RAG-Framework/
βββ backend/agentic_rag/ # Backend Python package
β βββ api.py # FastAPI server
β βββ agent.py # Core agent logic
β βββ tools.py # Search tools with semantic filtering
β βββ index.py # Document indexing
β βββ chat.py # CLI interface
β βββ ...
βββ frontend/ # React + Vite web app
β βββ src/
β β βββ components/ # UI components
β β βββ index.css # Styles & animations
β β βββ App.jsx
β βββ tailwind.config.js # Tailwind CSS config
β βββ package.json
βββ data/
β βββ docs/ # Your documents (.txt, .md, .pdf)
β βββ index/ # Generated embeddings
β βββ sessions/ # Chat history
βββ .env # Configuration
βββ requirements.txt # Python dependencies
- Python 3.9+
- Node.js 18+ and npm
- LM Studio (for local models) OR OpenAI/Azure API keys
Windows (PowerShell)
cd Agentic-RAG-Framework
py -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txtMac/Linux
cd Agentic-RAG-Framework
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtCreate .env file (or copy from .env.example):
# Provider: openai, azure, or use LM Studio
PROVIDER=openai
OPENAI_BASE_URL=http://127.0.0.1:1234/v1 # LM Studio endpoint
OPENAI_API_KEY=lm-studio-local # Dummy key for LM Studio
# Model names (use deployment names for Azure)
CHAT_MODEL=qwen2.5-7b-instruct-1m
EMBEDDING_MODEL=text-embedding-nomic-embed-text-v1.5
# Retrieval settings (optimized defaults)
TOP_K=8 # Number of chunks to retrieve
CHUNK_TOKENS=600 # Chunk size for faster processing
CHUNK_OVERLAP=120 # Overlap between chunksFor OpenAI:
PROVIDER=openai
OPENAI_API_KEY=sk-your-key-here
CHAT_MODEL=gpt-4o-mini
EMBEDDING_MODEL=text-embedding-3-smallFor Azure OpenAI:
PROVIDER=azure
AZURE_OPENAI_API_KEY=your-key
AZURE_OPENAI_BASE_URL=https://your-resource.openai.azure.com/openai/v1/
CHAT_MODEL=gpt-4o-mini # deployment name
EMBEDDING_MODEL=text-embedding-ada-002 # deployment namecd frontend
npm installPlace your documents (.txt, .md, .pdf) in data/docs/:
# Sample document included for testing
data/docs/sample_policy.mdGenerate embeddings for your documents:
python -m agentic_rag.indexTerminal 1 - Backend:
python -m uvicorn agentic_rag.api:app --reload --port 8000Terminal 2 - Frontend:
cd frontend
npm run devOpen your browser:
http://localhost:5173
Start the interactive CLI (without frontend):
python -m agentic_rag.chatResume a session:
python -m agentic_rag.chat --session <SESSION_ID>- Smooth letter-by-letter reveal (10ms per character)
- Blinking cursor during typing
- ChatGPT/Gemini-like experience
- Hover effects on all components
- Shake animation on validation errors
- Pulse glow on active send button
- Copy-to-clipboard for code blocks
- Smooth transitions throughout
Every response shows its sources:
SOURCES
π sample_policy.md
π document.pdf
TOP_K (default: 8)
- Number of chunks to retrieve from vector store
- Higher = better coverage, but more noise
- Recommended: 5-10
CHUNK_TOKENS (default: 600)
- Size of each document chunk
- Smaller = faster processing
- Recommended: 400-800
CHUNK_OVERLAP (default: 120)
- Overlap between consecutive chunks
- Prevents information loss at boundaries
- Recommended: 15-20% of CHUNK_TOKENS
The framework automatically filters results with relevance score < 0.5:
# In tools.py
RELEVANCE_THRESHOLD = 0.5 # Only keep relevant resultsThis prevents irrelevant answers and improves accuracy.
Edit backend/agentic_rag/tools.py:
TOOLS = [
{
"type": "function",
"function": {
"name": "your_custom_tool",
"description": "What this tool does",
"parameters": {...}
}
}
]- Colors: Edit CSS variables in
frontend/src/index.css - Animations: Adjust keyframes in
frontend/src/index.css - Components: Modify files in
frontend/src/components/
In frontend/src/components/ChatInterface.jsx:
const typingSpeed = 3; // milliseconds per character# Build and deploy the FastAPI app
pip install -r requirements.txt
uvicorn agentic_rag.api:app --host 0.0.0.0 --port 8000cd frontend
npm run build # Creates dist/ folder
# Deploy dist/ folder to your hosting serviceUpdate frontend proxy for production:
Edit frontend/vite.config.js:
server: {
proxy: {
'/ask': {
target: process.env.VITE_API_URL || 'https://your-backend.com',
changeOrigin: true,
}
}
}1. "What is the POSH policy?"
β Should return Prevention of Sexual Harassment policy
2. "Tell me about leave policy"
β Should return vacation/leave information
3. "Remote work guidelines"
β Should return remote work policies
4. "What is Python?" (general knowledge)
β Should answer directly without searching docs
- β Typing animation is smooth
- β Sources are displayed
- β Hover effects work
- β Copy-to-clipboard on code blocks
- β Empty submit triggers shake animation
Query the agent:
Request:
{
"query": "What is the vacation policy?",
"session_id": "sess_abc123" // Optional
}Response:
{
"answer": "Unused leave can be carried forward...",
"source": ["sample_policy.md", "hr_handbook.pdf"]
}- Backend: Python, FastAPI, OpenAI API, FAISS
- Frontend: React, Vite, Tailwind CSS
- Vector Store: FAISS (cosine similarity)
- Embeddings: OpenAI text-embedding-3-small or custom models
Contributions welcome! Please feel free to submit a Pull Request.