A production-grade LLM application architecture built from scratch in Python, implementing essential enterprise patterns: FastAPI API server, semantic caching, input/output guardrails, prompt versioning & A/B testing, cost tracking, retry mechanisms with fallback chains, and Google Gemini integration.
User Request
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Input Guardrails (Prompt injection defense & PII redaction) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Semantic Cache (Cosine similarity on query embeddings) │
│ ├─ Cache Hit ──► Return cached response (0ms, $0 cost) │
│ └─ Cache Miss ──► Continue to LLM pipeline │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Prompt Router (Template management & A/B test experiments) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ LLM Engine & Fallback Chain (Gemini / OpenAI / Anthropic) │
│ └─ Exponential backoff & retry with jitter │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Output Guardrails (Harmful content & code safety checks) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Observability & Tracking (Token usage, latency, unit costs) │
└─────────────────────────────────────────────────────────────┘
- FastAPI Server (
server.py): REST API endpoints for/v1/chat,/health,/v1/costs,/v1/cache/stats, supporting SSE streaming. - Production LLM Service (
production_app.py): Core pipeline logic incorporating guardrails, vector embeddings, semantic caching, exponential backoff retries, fallback model cascades, and per-user cost tracking. - Interactive Chat Application (
real_app.py): CLI chat client directly integrated with Google Gemini (gemini-3.6-flash).
.
├── .env.example # Environment variable templates
├── .gitignore # Keeps secrets and bytecode out of git
├── requirements.txt # Python dependencies
├── production_app.py # Full production pipeline & simulation suite
├── server.py # FastAPI web server and streaming endpoints
├── real_app.py # Live CLI application connected to Gemini
└── README.md # Documentation
# Clone the repository
git clone https://github.com/adityagupta27-cpu/Production-LLM-application.git
cd Production-LLM-application
# Install dependencies
pip install -r requirements.txtSet up your Gemini API key (get a free key at Google AI Studio):
cp .env.example .env
# Open .env and add your GEMINI_API_KEYpython real_app.pyuvicorn server:app --reload --port 8000Access the interactive Swagger API documentation at:
http://localhost:8000/docs