Skip to content

Latest commit

 

History

History
313 lines (214 loc) · 8.02 KB

File metadata and controls

313 lines (214 loc) · 8.02 KB

Setup Guide

Written by Raqibul Hasan Moon.

Getting this running locally, from nothing to a working chatbot. Roughly ten minutes, most of it waiting on downloads.


Before you start

Requirement Why Check
Node.js 18+ The code uses the built-in fetch node -v
Docker Runs Postgres with the pgvector extension docker --version
Gemini API key Embeddings How to get one
Groq API key Answer generation How to get one

Node 18 is the real floor — fetch became global there, and this project uses it directly instead of installing a HTTP client.

If you would rather not use Docker, any Postgres works provided pgvector is installed. Docker is simply the path of least resistance.


Step 1 — Install dependencies

npm install

Installs Express, pg, pdf-parse, Multer, uuid, and dotenv. Short list on purpose — no LangChain, no vector-database SDK.


Step 2 — Set up your environment file

cp .env.example .env

Open .env and paste in your two keys:

GEMINI_API_KEY=your-key-here
GROQ_API_KEY=gsk_your-key-here

The Postgres values can stay as they are for local development.

One rule: the credentials in DATABASE_URL must match POSTGRES_USER, POSTGRES_PASSWORD, and POSTGRES_DB above it. Docker creates the database from those three; the app connects using the URL. Changing one and forgetting the other is the single most common setup failure.

POSTGRES_USER=raguser
POSTGRES_PASSWORD=change-me-in-production
POSTGRES_DB=rag_db
                    ↓  these three must appear in the URL  ↓
DATABASE_URL=postgresql://raguser:change-me-in-production@localhost:5432/rag_db

Step 3 — Start Postgres

docker compose --env-file .env up -d

--env-file .env matters: the compose file reads your credentials from it. Without the flag, Docker cannot fill in the variables.

First run downloads the image (~400 MB) and takes a minute. Check it is up:

docker compose ps

Look for healthy in the status. The compose file includes a healthcheck that runs pg_isready, so healthy means the database is genuinely accepting queries — not merely that the port is open.

Why not the normal postgres image?

The compose file uses pgvector/pgvector:pg17, which ships the vector extension. The stock postgres image does not have it, and the app fails at boot with:

extension "vector" is not available

Step 4 — Start the app

npm run dev     # auto-restarts when you edit a file

or

npm start       # plain node

You want to see exactly this:

Database ready
RAG chatbot running on port 3000

Database ready means initDb() succeeded — the vector extension exists and the documents table is ready.

If the process exits instead, Postgres is unreachable. That is deliberate: the server refuses to accept requests it cannot serve. See Troubleshooting.


Step 5 — Ingest a PDF

There is a test.pdf in the project root:

curl -X POST http://localhost:3000/ingest -F "file=@test.pdf"
{ "message": "Ingested 8 chunks from \"test.pdf\"" }

Meanwhile the server logs Processing 8 chunks from "test.pdf".

This is the slow endpoint — one Gemini call per chunk, sent sequentially to respect rate limits. A large PDF takes a while, and that is expected.

Use your own PDF by pointing at it:

curl -X POST http://localhost:3000/ingest -F "file=@/path/to/your.pdf"

Note: there is no deduplication. Ingesting the same file twice stores its chunks twice.


Step 6 — Ask a question

curl -X POST http://localhost:3000/chat \
  -H "Content-Type: application/json" \
  -d '{ "question": "Webhook responded" }'
{
  "answer": "Webhook responded with status code 400.",
  "sources": ["test.pdf"],
  "topSimilarity": "0.694"
}

Working. topSimilarity tells you how well retrieval did — how to read it.

Now try something the PDF does not cover:

curl -X POST http://localhost:3000/chat \
  -H "Content-Type: application/json" \
  -d '{ "question": "What is the capital of France?" }'

A low score and a refusal to answer. That is correct behaviour — proof the model is answering from your documents rather than its own knowledge.

More example requests, including PowerShell versions, are in scripts/bash.sh.


Everyday commands

# Start the database (once per reboot)
docker compose --env-file .env up -d

# Start the app
npm run dev

# Stop the app
Ctrl-C

# Stop the database, keep the data
docker compose down

# Stop the database and DELETE all ingested documents
docker compose down -v

down -v removes the volume. Every embedding you paid for is gone and must be re-ingested — useful for a clean slate, painful by accident.


Inspecting the database

docker compose exec postgres psql -U raguser -d rag_db

Useful queries:

-- How many chunks are stored?
SELECT COUNT(*) FROM documents;

-- Which documents have been ingested?
SELECT source, COUNT(*) FROM documents GROUP BY source;

-- Look at some actual chunks
SELECT source, LEFT(content, 80) FROM documents LIMIT 5;

-- Remove one document's chunks (useful before re-ingesting it)
DELETE FROM documents WHERE source = 'test.pdf';

Exit with \q.

That third query is worth running once — seeing the real chunk boundaries makes chunking and overlap concrete in a way reading about it does not.


Troubleshooting

App exits immediately, no Database ready

Postgres is not reachable. Check docker compose ps shows healthy, then verify DATABASE_URL matches your POSTGRES_* values. Mismatched credentials are the usual cause.

ECONNREFUSED 127.0.0.1:5432

The database is not listening yet. First boot takes a few seconds to initialise the data directory — wait and retry. If something else already uses port 5432 (a native Postgres install, commonly), change the host side of the mapping in docker-compose.yml:

ports:
  - "5433:5432"     # host 5433 -> container 5432

Then update DATABASE_URL to use localhost:5433.

extension "vector" is not available

Postgres is running without pgvector — almost always because a plain postgres image was used before. Recreate the container:

docker compose down -v
docker compose --env-file .env up -d

pdfParse is not a function

Fixed in this codebase already. pdf-parse v2 exports a PDFParse class where v1 exported a callable function. If you see this, you are running older code — check src/ingest.js uses new PDFParse({ data: buffer }).

Ingest succeeds but says 0 chunks

Either every chunk fell under the 50-character minimum, or the PDF is image-only. A scanned PDF has no text layer — nothing to extract, nothing to embed. Run it through OCR first.

429 Too Many Requests

Free-tier rate limit, usually while ingesting something large. Wait a minute. For persistent trouble, add a delay between chunks in src/ingest.js.

Answers are vague or wrong

Check topSimilarity before touching anything else. Below 0.5 means retrieval failed and the right text never reached the model — the fix is in chunking or phrasing, not in the prompt. Confirm the document actually contains the answer, then try wording closer to the document's own language.

Environment variables not loading

dotenv reads .env once at startup. Restart the server after editing it. Confirm the file is named exactly .env, sits in the project root, and has no quotes around values.


Next