Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Production LLM Application

A production-grade LLM application architecture built from scratch in Python, implementing essential enterprise patterns: FastAPI API server, semantic caching, input/output guardrails, prompt versioning & A/B testing, cost tracking, retry mechanisms with fallback chains, and Google Gemini integration.

Architecture & Features

User Request
    │
    ▼
┌─────────────────────────────────────────────────────────────┐
│ Input Guardrails (Prompt injection defense & PII redaction) │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Semantic Cache (Cosine similarity on query embeddings)      │
│  ├─ Cache Hit  ──► Return cached response (0ms, $0 cost)    │
│  └─ Cache Miss ──► Continue to LLM pipeline                 │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Prompt Router (Template management & A/B test experiments)  │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ LLM Engine & Fallback Chain (Gemini / OpenAI / Anthropic)   │
│  └─ Exponential backoff & retry with jitter                 │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Output Guardrails (Harmful content & code safety checks)    │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Observability & Tracking (Token usage, latency, unit costs) │
└─────────────────────────────────────────────────────────────┘

Key Components

  • FastAPI Server (server.py): REST API endpoints for /v1/chat, /health, /v1/costs, /v1/cache/stats, supporting SSE streaming.
  • Production LLM Service (production_app.py): Core pipeline logic incorporating guardrails, vector embeddings, semantic caching, exponential backoff retries, fallback model cascades, and per-user cost tracking.
  • Interactive Chat Application (real_app.py): CLI chat client directly integrated with Google Gemini (gemini-3.6-flash).

Project Structure

.
├── .env.example        # Environment variable templates
├── .gitignore          # Keeps secrets and bytecode out of git
├── requirements.txt    # Python dependencies
├── production_app.py   # Full production pipeline & simulation suite
├── server.py           # FastAPI web server and streaming endpoints
├── real_app.py         # Live CLI application connected to Gemini
└── README.md           # Documentation

Quick Start

1. Installation

# Clone the repository
git clone https://github.com/adityagupta27-cpu/Production-LLM-application.git
cd Production-LLM-application

# Install dependencies
pip install -r requirements.txt

2. Configuration

Set up your Gemini API key (get a free key at Google AI Studio):

cp .env.example .env
# Open .env and add your GEMINI_API_KEY

3. Run the Live Chatbot

python real_app.py

4. Run the FastAPI Production Server

uvicorn server:app --reload --port 8000

Access the interactive Swagger API documentation at: http://localhost:8000/docs

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages