Ask questions about your own PDFs and get answers grounded in their contents, with the source document cited.
A deliberately small, dependency-light Retrieval-Augmented Generation service. Upload a PDF, it is split into overlapping chunks, each chunk is embedded and stored as a vector in Postgres. A question is embedded the same way, the closest chunks are retrieved by cosine distance, and an LLM answers using only those chunks as context.
No LangChain, no vector-database service, no ORM. Two HTTP endpoints and about 200 lines of application code — the retrieval mechanics stay visible rather than hidden behind a framework.
New here? Start with the docs. This README is the reference; the guides below are the explanation.
Guide What it covers 1. How RAG works The concepts in plain English — embeddings, vectors, chunking. Start here. 2. Getting API keys Gemini and Groq accounts, step by step. Includes swapping in OpenAI. 3. Setup guide Nothing to working chatbot in ~10 minutes. 4. Code walkthrough Every file explained, and why it is written that way.
- How it works
- Stack
- Project layout
- Prerequisites
- Setup
- Running
- API reference
- Interpreting results
- Configuration
- How the pieces work
- Deployment
- Troubleshooting
- Scaling notes and known limitations
┌──────────── INGEST ────────────┐
PDF ──> pdf-parse ──> chunk (500 chars, 50 overlap) ──> Gemini embeddings
│
v
Postgres + pgvector
documents(content,
source, embedding)
^
│
question ──> Gemini embedding ──> cosine search (top 5) ─────────┘
│
v
Groq / Llama 3.1 ──> answer + sources
└──────────── QUERY ─────────────┘
The core idea: text with similar meaning produces vectors pointing in similar directions. Finding relevant context becomes a distance comparison rather than a keyword match, so "Webhook responded" can retrieve a passage that never uses those exact words.
| Concern | Choice | Why |
|---|---|---|
| Runtime | Node.js 18+ (CommonJS) | Built-in fetch, no transpiler |
| HTTP | Express 5 | Two routes; nothing heavier is warranted |
| Uploads | Multer (memory storage) | File is parsed and discarded — never touches disk |
| PDF text | pdf-parse v2 |
Class-based API (see note below) |
| Database | Postgres 17 + pgvector | Vectors live beside relational data; one datastore |
| Embeddings | Gemini gemini-embedding-001 |
3072-dim, strong retrieval quality |
| Generation | Groq llama-3.1-8b-instant |
Fast and inexpensive for grounded answers |
Both AI providers are called with plain fetch rather than their SDKs. Google's
Node SDK defaults to the v1beta endpoint, where gemini-embedding-001 is not
available — only on v1. Calling directly sidesteps that and keeps the
dependency list short.
rag-chatbot-nodejs/
├── src/
│ ├── index.js Express app, routes, boot sequence
│ ├── db.js Postgres pool, schema bootstrap (initDb)
│ ├── ingest.js PDF -> text -> chunks -> embeddings -> rows
│ ├── query.js question -> embedding -> vector search -> answer
│ └── embeddings.js Gemini + Groq HTTP clients
├── docs/
│ ├── 01-how-it-works.md RAG concepts in plain English
│ ├── 02-api-keys.md Getting Gemini + Groq keys
│ ├── 03-setup.md Local setup, start to finish
│ ├── 04-code-walkthrough.md Every file explained
│ └── upload/ Local PDF drop (git-ignored)
├── scripts/
│ └── bash.sh Reference curl commands for both endpoints
├── docker-compose.yml Postgres with pgvector
├── .env.example Config template — copy to .env
└── package.json
| File | Responsibility |
|---|---|
src/index.js |
Route definitions, upload validation, error → HTTP mapping. Starts listening only after initDb() resolves. |
src/db.js |
Single pg connection pool. Creates the vector extension and documents table idempotently on boot. |
src/ingest.js |
Extracts text, chunks it with overlap, embeds each chunk sequentially, inserts one row per chunk. |
src/query.js |
Embeds the question, retrieves the 5 nearest chunks, joins them into context, returns answer + sources + score. |
src/embeddings.js |
The only file that talks to external APIs. Swap providers here. |
- Node.js 18 or newer —
fetchmust be global. Developed on Node 25. - Docker — for Postgres with pgvector. A native Postgres works too, but the
vectorextension must be installed. - A Gemini API key — free tier is sufficient. aistudio.google.com/apikey
- A Groq API key — free tier is sufficient. console.groq.com/keys
1. Install dependencies
npm install2. Create your environment file
cp .env.example .envOpen .env and fill in both API keys. .env is git-ignored and must stay that
way — it holds live credentials.
3. Start Postgres
docker compose --env-file .env up -dThe image is pgvector/pgvector:pg17. The stock postgres image will not
work: initDb() runs CREATE EXTENSION vector on every boot and fails without
it.
Confirm it is accepting connections:
docker compose ps
docker compose logs postgres | tail4. Start the application
npm run dev # nodemon, reloads on change
npm start # plain nodeExpected output:
Database ready
RAG chatbot running on port 3000
Database ready means the extension and table exist. If the process exits
instead, Postgres is unreachable — see Troubleshooting.
Ingest a document, then ask about it:
curl -X POST http://localhost:3000/ingest -F "file=@test.pdf"
# { "message": "Ingested 8 chunks from \"test.pdf\"" }
curl -X POST http://localhost:3000/chat \
-H "Content-Type: application/json" \
-d '{ "question": "Webhook responded" }'
# { "answer": "Webhook responded with status code 400.",
# "sources": ["test.pdf"], "topSimilarity": "0.694" }More examples, including PowerShell equivalents, are in
scripts/bash.sh.
Multipart upload of a single PDF under the field name file.
| Content-Type | multipart/form-data |
| Field | file — a PDF |
{ "message": "Ingested 8 chunks from \"test.pdf\"" }| Status | Meaning |
|---|---|
200 |
Chunks were embedded and stored |
400 |
No file attached, or the declared type is not PDF |
500 |
Parsing, embedding, or database failure — see detail |
Re-uploading the same file ingests it again; there is no deduplication.
Chunks accumulate under the same source name.
{ "question": "What status code did the webhook return?" }{
"answer": "Webhook responded with status code 400.",
"sources": ["test.pdf"],
"topSimilarity": "0.694"
}| Field | Meaning |
|---|---|
answer |
Generated using only the retrieved chunks as context |
sources |
Deduplicated filenames the context came from |
topSimilarity |
Cosine similarity of the single best chunk, 0–1 |
| Status | Meaning |
|---|---|
200 |
Answered (including "not enough information") |
400 |
question missing or not a string |
500 |
Embedding, search, or generation failure |
With no documents ingested, returns answer: "No relevant documents found."
topSimilarity reports how well retrieval performed — not how correct the
answer is. It is the most useful debugging signal in the system.
| Score | Reading |
|---|---|
| > 0.7 | Strong match; retrieved chunks genuinely address the question |
| 0.5 – 0.7 | Partial; related material, possibly not the specific answer |
| < 0.5 | Weak; matched loosely, likely without the answer present |
A low score with a hedged answer is the system working correctly. The system prompt instructs the model to state plainly when the context is insufficient, so it declines rather than inventing an answer.
Worth testing deliberately: ask about something your PDF never mentions and confirm the model declines. That behaviour is what makes the service safe to trust on the questions it can answer.
All configuration is environment variables, loaded via dotenv.
| Variable | Required | Purpose |
|---|---|---|
GEMINI_API_KEY |
yes | Embeddings for both ingestion and queries |
GROQ_API_KEY |
yes | Answer generation |
DATABASE_URL |
yes | Postgres connection string used by the app |
POSTGRES_USER |
compose | Database user created by Docker |
POSTGRES_PASSWORD |
compose | Database password created by Docker |
POSTGRES_DB |
compose | Database name created by Docker |
PORT |
no | HTTP port, defaults to 3000 |
The three POSTGRES_* values are consumed by docker-compose.yml when creating
the container. They must agree with the credentials embedded in DATABASE_URL —
changing one without the other is the most common setup mistake.
Tuning constants are in code rather than config: chunk size and overlap in
src/ingest.js, retrieval count (LIMIT 5) and the model's
system prompt in src/query.js and
src/embeddings.js.
Chunking — 500 characters, 50 of overlap. The overlap exists so a sentence straddling a boundary is not split into two halves that each make no sense when retrieved alone. 500 suits most technical documents; dense legal or financial text usually wants closer to 300. Chunks shorter than 50 characters are dropped as too small to carry meaning — which is why a very short PDF can ingest zero chunks and still report success.
Embedding. Every chunk becomes a 3072-dimension vector via
gemini-embedding-001. Requests are sent one at a time: firing them in
parallel with Promise.all() trips Gemini's per-minute rate limit on any
sizeable document.
Storage. One row per chunk in documents (id, content, source, embedding).
source is the original filename, which is what allows answers to cite where
they came from.
Retrieval. The question is embedded with the same model — vectors from
different models occupy unrelated coordinate spaces and cannot be compared.
pgvector's <=> operator gives cosine distance; 1 - distance converts it to
the similarity score returned to the client. Results are ordered by raw distance
ascending, i.e. closest first.
Generation. The top 5 chunks are joined into a context block and sent to Llama 3.1 with an instruction to answer only from that context. This grounding is what separates the system from a plain chatbot: answers are traceable to source documents, and gaps are admitted rather than filled in.
The application is stateless — all state is in Postgres — so it scales horizontally behind a load balancer without further coordination.
Before going to production:
- Rotate every key that has been in a local
.env, and inject secrets from your platform's secret manager rather than a file on disk. - Use managed Postgres with pgvector — AWS RDS, Google Cloud SQL, Supabase, and Neon all support the extension. Verify it is enabled before deploying.
- Require TLS on the database connection (
?sslmode=requireinDATABASE_URL). - Put authentication in front of
/ingest. As written, it is unauthenticated; anyone who can reach it can add documents to the knowledge base and influence every answer the system gives. - Bound upload size. Multer buffers the whole file in memory with no limit
configured, so a large upload is an easy memory-exhaustion vector. Set
multer({ limits: { fileSize: ... } }). - Add rate limiting on both routes — each request costs money at two external APIs.
- Run behind a process supervisor. The app exits deliberately if the database is unreachable at boot, and expects to be restarted.
Container build. No Dockerfile is included; the app runs directly on Node. A minimal one:
FROM node:22-alpine
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY src ./src
EXPOSE 3000
CMD ["node", "src/index.js"]Point DATABASE_URL at your managed instance and supply both API keys as
environment variables.
Health checking. There is no /health route. Add one before deploying
behind a load balancer that expects a probe.
pdfParse is not a function
pdf-parse v2 is a breaking rewrite: v1 exported a callable, v2 exports a
PDFParse class. The correct v2 usage — already in this codebase — is:
const { PDFParse } = require("pdf-parse");
const parser = new PDFParse({ data: buffer });
try {
const { text } = await parser.getText();
} finally {
await parser.destroy(); // releases the PDF.js worker
}extension "vector" is not available
Postgres is running without pgvector — almost always the stock postgres image
instead of pgvector/pgvector. Fix the image, then recreate the container.
Process exits immediately, no Database ready
Postgres is unreachable or the credentials are wrong. Confirm the container is
up (docker compose ps), then check that DATABASE_URL matches the
POSTGRES_* values. Mismatched credentials between the two are the usual cause.
ECONNREFUSED 127.0.0.1:5432
The database is not listening yet. Postgres takes a few seconds to initialise on
first run while it creates the data directory — wait, then retry. If something
else already occupies 5432, remap the host port in docker-compose.yml.
Ingest returns 200 but reports 0 chunks
Every chunk fell below the 50-character floor, or the PDF is image-only. A
scanned PDF contains no text layer; pdf-parse extracts nothing and there is
nothing to embed. OCR it first.
Gemini 400 / model not found
gemini-embedding-001 is served on the v1 endpoint, not v1beta. If you
switch to the official SDK, this is the failure you will hit.
429 from either provider
Free-tier rate limits. Ingestion is already sequential; for large documents, add
a short delay between chunks in the loop in src/ingest.js.
Answers are vague or wrong
Check topSimilarity first. A low score means retrieval failed, not the model —
so the fix is in chunking or the question's phrasing, not the prompt. Confirm
the document actually contains the answer, and try a phrasing closer to the
document's own wording.
The vector column cannot be indexed at 3072 dimensions. pgvector permits
vector columns up to 16,000 dimensions, but both HNSW and IVFFlat indexes cap
at 2,000. Every query therefore runs an exact sequential scan.
This is genuinely fine — preferable, even — at small scale: exact search returns true nearest neighbours where an approximate index only estimates them. It stops being fine in the tens of thousands of chunks, when the scan dominates latency. Two ways out, in increasing order of effort:
- Store as
halfvec(3072)and build an HNSW index —halfvecindexes support up to 4,000 dimensions. Costs a little precision, keeps the model. - Request a smaller embedding via Gemini's
outputDimensionality(768 or 1536) and index normally. Requires re-embedding every document.
Either path means regenerating all stored vectors: embeddings from different models or dimensions are not comparable.
Other current limitations:
- No deduplication. Re-ingesting a file duplicates its chunks. Delete by
sourcefirst, or add a uniqueness constraint. - Partial ingestion on failure. Each chunk commits in its own transaction, so an error midway leaves earlier chunks stored. Re-ingesting compounds it — see deduplication above.
- PDF only. The route rejects other types; the mimetype check is a convenience, not a security control, since the client supplies that header.
- No authentication or rate limiting on either route.
- No upload size limit. Files are buffered entirely in memory.
- No conversation memory. Each
/chatcall is independent; follow-up questions that depend on prior turns will not resolve. - No automated tests.
Built by Raqibul Hasan Moon as a learning project, following How to Build a RAG Chatbot with Node.js, Gemini and pgvector on freeCodeCamp, then extended with Docker setup, expanded documentation, and notes on the parts worth understanding more deeply.
ISC