Skip to content

Repository files navigation

Mortgage Document Intelligence System

Live site (deployed on Vercel): https://triplemodelrag.vercel.app An interactive companion site (in site/) walks through the full RAG pipeline, chunking and embedding techniques, 3D vector retrieval, a guardrailed chatbot, security, evals, and impact. Source: site/, details: site/README.md.

Vercel Demo GIFs Next.js

Demos

Live RAG pipeline simulation with cost and review-time counters

Interactive 3D vector retrieval Grounded, guardrailed chatbot simulation Site tour

Run the notebooks

The three pipelines are Jupyter notebooks. You can run them offline or in Google Colab.

Offline (V3, no external OCR API): V3's 5-tier cascade (PyMuPDF, pdfplumber, PyPDF2, Tesseract, EasyOCR) runs entirely locally with no external OCR calls. For fully offline generation, point it at a local Llama 3 server (Ollama or vLLM) instead of a hosted endpoint.

pip install -r requirements.txt
jupyter notebook pipeline_v3_multitier.ipynb

Google Colab (V1, V2, V3): open a notebook in Colab and add your API keys (Replicate for V1/V2; an OpenAI-compatible key for V3 generation) as Colab Secrets.

Pipeline Open in Colab Needs
V1 DocTR + DeepSeek Open In Colab Replicate key
V2 Chandra OCR Open In Colab Replicate key
V3 Multi-tier Open In Colab Generation key (or local model)

Three Pipelines, Three OCR Strategies

Three production RAG (Retrieval-Augmented Generation) pipelines for processing 200+ page mortgage PDFs. Each pipeline uses a different OCR strategy, all delivering natural language Q&A with source citations, confidence scores, and a Gradio UI.

Pipeline File Primary OCR Embeddings LLM Key Strength
V1 pipeline_v1_doctr_deepseek.ipynb DocTR + DeepSeek OCR BGE-large-en-v1.5 + Reranker DeepSeek via Replicate Word-level bounding boxes for source highlighting; dual OCR comparison
V2 pipeline_v2_chandra.ipynb Chandra OCR BGE-large-en-v1.5 + Reranker Chandra via Replicate 40+ language support; 99.99% multilingual accuracy; 83.1 benchmark score
V3 pipeline_v3_multitier.ipynb 5-tier cascade (PyMuPDF, pdfplumber, PyPDF2, Tesseract, EasyOCR) sentence-transformers/all-mpnet-base-v2 Llama 3 70B Zero external OCR dependency; handles any PDF type via fallback

All three share the same core architecture: extract text, chunk semantically, embed, index in FAISS, retrieve, then generate an answer with source attribution.

Built during an AI Engineering externship at Outamation, where the goal was to find the most reliable extraction approach for complex mortgage documents: scanned pages, tables, degraded quality, and multi-hundred-page packages.

Sample documents

Tocr 1 Tocr 2
Extern 1 Extern 2
Extern 3 Extern 4
Extern 5 Extern 6
Extern 7 Extern 8

Why Three Versions?

Mortgage documents are unpredictable. A single loan package might contain native digital text, scanned pages from the 1980s, handwritten annotations, and complex fee tables, all in one PDF. No single OCR engine handles everything well.

V1 (DocTR + DeepSeek) was built first. DocTR provides word-level bounding boxes, enabling visual source highlighting on the original page. DeepSeek OCR runs as a secondary engine via Replicate, and the system compares outputs from both to select the highest-quality extraction per page. Uses BGE-large embeddings with a cross-encoder reranker for high-precision retrieval.

V2 (Chandra OCR) was built to test a newer engine. Chandra scored 83.1 on OCR benchmarks, supports 40+ languages with 99.99% multilingual accuracy, and processes pages at 0.025s on H100 GPUs. It runs via Replicate with retry logic and timeout handling. Same BGE + reranker retrieval stack as V1. Best for multilingual or math-heavy mortgage documents.

V3 (Multi-Tier Fallback) was built for zero-cost, zero-dependency reliability. Instead of relying on any external API for OCR, it cascades through five local extraction methods, keeping the highest-confidence result. Uses sentence-transformers for embeddings and a free-tier Llama 3 70B endpoint for generation. Best for environments where external API access is restricted, and the only path that runs fully offline for OCR.

System Architecture

                    +-----------------------------------+
                    |           PDF Upload              |
                    +----------------+------------------+
                                     |
                  +------------------+------------------+
                  v                  v                  v
          +------------+     +-------------+     +----------------+
          | V1: DocTR  |     | V2: Chandra |     | V3: 5-Tier     |
          | + DeepSeek |     |    OCR      |     | Cascade        |
          +-----+------+     +------+------+     +-------+--------+
                |                   |                    |
                +-------------------+--------------------                                    v
                         +-----------------------+
                         |  Semantic Chunking    |
                         |  (sentence boundary,  |
                         |   overlap, metadata)  |
                         +----------+------------+
                                    v
                         +-----------------------+
                         |  Embedding + Index    |
                         |  BGE-large or MPNet   |
                         |  FAISS L2 search      |
                         +----------+------------+
                                    v
                         +-----------------------+
                         |  Retrieval + Rerank   |
                         |  (V1/V2 cross-encoder |
                         |   reranker)           |
                         +----------+------------+
                                    v
                         +-----------------------+
                         |  Answer Generation    |
                         |  Llama 3 70B /        |
                         |  DeepSeek / Replicate |
                         +----------+------------+
                                    v
                         +-----------------------+
                         |  Gradio UI            |
                         |  Chat + Sources +     |
                         |  Confidence Scores    |
                         +-----------------------+

About

3 production RAG pipelines for mortgage PDFs — DocTR + DeepSeek, Chandra OCR, and 5-tier local fallback. FAISS vector search, BGE/MPNet embeddings, Gradio UI with source citations.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages