A fully containerized Retrieval Augmented Generation (RAG) stack for any amd64 (or arm64) Linux host with Docker. Qdrant (vector database), SearXNG (private search engine), and AnythingLLM (web UI + RAG orchestrator) in a single Docker Compose deployment — with the LLM and embeddings served by any external OpenAI-compatible API (vLLM, llama.cpp server, OpenRouter, OpenAI, …).
- AnythingLLM: RAG orchestrator + web UI. Talks to your external OpenAI-compatible LLM/embedding API and to the local Qdrant/SearXNG services.
- Qdrant: High-performance vector database for efficient similarity search.
- SearXNG: Privacy-respecting metasearch engine to augment RAG answers with web results.
- Caddy: TLS reverse proxy in front of the UI/APIs (loopback-only).
There is no bundled LLM server: the stack expects an OpenAI-compatible
endpoint you already run (e.g. vLLM at http://<host>:8000/v1) and sends both
chat completions and embeddings there. The stack is hardware-agnostic: it
reserves no GPU devices and ships no accelerator-specific tuning. If your LLM
host happens to have a GPU, the provider process serves it — these containers
are CPU-only.
flowchart LR
UI[AnythingLLM UI :3001] --> CMD[caddy :8443 TLS]
CMD -- /api/vectors/* --> QDR[Qdrant :6333]
CMD -- /search/* --> SX[SearXNG :8080]
CMD -- /* --> UI
UI -- OpenAI-compatible /v1/chat/completions --> EXT[External LLM API]
UI -- OpenAI-compatible /v1/embeddings --> EXT
QDR --> STO[(qdrant_storage)]
- Docker Engine (version 20.10+)
- Docker Compose (v2.0+)
- An OpenAI-compatible API endpoint reachable from the AnythingLLM
container. The default
.envassumes a server on the Docker host athttp://host.docker.internal:8000/v1(vLLM's default port); the compose file adds thehost-gatewaymapping so that name resolves inside the container. On Linux hosts running Docker < 20.10 this mapping is implicit; on rootless/podman setups it may need--network=host-style alternatives.
-
Clone the repository:
git clone https://github.com/amasu/dgx-spark-rag.git cd dgx-spark-rag -
Copy the environment file:
cp .env.example .env
Point
GENERIC_OPEN_AI_BASE_PATH/EMBEDDING_BASE_PATHat your OpenAI-compatible endpoint and set the model IDs + API keys (see Configuration). -
Start the services:
docker compose pull docker compose up -d
Once the stack is running, access the services at:
- AnythingLLM UI: http://localhost:3001
- Qdrant API: http://localhost:6333
- SearXNG: http://localhost:8181
- TLS entry point (Caddy): https://localhost:8443
https://localhost:8443 is served with a self-signed certificate (tls internal). Browsers and API clients will show a certificate warning on first
connect — accept it (or point your client's CA bundle at Caddy's internal CA
in the caddy_data volume). The per-service ports above are plain HTTP on
loopback only; nothing is exposed beyond 127.0.0.1.
Everything is configured through .env — nothing is wired to a specific
vendor. Common layouts:
| Provider | *_BASE_PATH |
Model IDs |
|---|---|---|
| vLLM on this host | http://host.docker.internal:8000/v1 |
served model name (vllm list-served-models) |
| llama.cpp server | http://host.docker.internal:8080/v1 |
-m alias (--alias) |
| OpenRouter | https://openrouter.ai/api/v1 |
vendor/model:tag + API key |
| OpenAI | https://api.openai.com/v1 |
e.g. gpt-4o-mini, text-embedding-3-small |
The embedding endpoint must implement POST /v1/embeddings. If your
embedding API needs a different host than the LLM, set EMBEDDING_BASE_PATH
separately (both default to the same host in .env.example).
First-run behavior: AnythingLLM consumes the LLM_PROVIDER /
GENERIC_OPEN_AI_* / EMBEDDING_* / VECTOR_DB variables on boot to build
its system preferences — a fresh anythingllm_storage volume starts already
wired to your provider. If you previously ran an older stack variant, delete
or archive ./anythingllm_storage once so the new defaults take effect
(profiles/chats live there and keep stale provider settings).
SearXNG configuration is stored in the ./searxng directory. Modify settings
there and restart the container:
docker compose restart searxngAll services store data in named volumes or bind mounts:
- Qdrant:
./qdrant_storage - AnythingLLM:
./anythingllm_storage - SearXNG:
./searxng(configuration)
All runtime configuration lives in the .env file (copy it from
.env.example). Docker Compose reads these variables automatically. Every
variable is wired into docker-compose.yml with ${VAR:-default} fallbacks,
so the stack also renders without an .env (fresh-clone safe).
All services are pinned to multi-arch images and respect their standard
deploy.resources CPU/memory limits — no core-pinning or accelerator
configuration, so the same files deploy identically on any machine.
| Variable | Default | Description |
|---|---|---|
QDRANT_PORT |
6333 |
Host port for the Qdrant HTTP/REST API. |
SEARXNG_PORT |
8181 |
Host port for the SearXNG web UI. |
ANYTHINGLLM_PORT |
3001 |
Host port for the AnythingLLM web UI. |
CADDY_PORT |
8443 |
Host port for the Caddy TLS entry point. |
| Variable | Default | Description |
|---|---|---|
QDRANT_STORAGE |
./qdrant_storage |
Bind-mount path for Qdrant vector data. |
ANYTHINGLLM_STORAGE |
./anythingllm_storage |
Bind-mount path for AnythingLLM data on the host. |
SEARXNG_CONFIG |
./searxng |
Bind-mount path for SearXNG configuration files. |
| Variable | Default | Description |
|---|---|---|
LLM_PROVIDER |
generic-openai |
AnythingLLM provider selector — do not change unless you switch engines. |
GENERIC_OPEN_AI_BASE_PATH |
http://host.docker.internal:8000/v1 |
Base URL of the OpenAI-compatible chat-completions API. host.docker.internal resolves to the Docker host (compose maps it via host-gateway). |
GENERIC_OPEN_AI_API_KEY |
(empty) | API key sent as Authorization: Bearer …. Dummy value if your server doesn't check keys. |
GENERIC_OPEN_AI_MODEL_PREF |
(empty) | Model ID the provider serves (e.g. the vLLM --served-model-name). |
| Variable | Default | Description |
|---|---|---|
EMBEDDING_ENGINE |
generic-openai |
AnythingLLM embedding-engine selector. |
EMBEDDING_BASE_PATH |
http://host.docker.internal:8000/v1 |
Base URL of the /v1/embeddings endpoint. |
EMBEDDING_MODEL_PREF |
(empty) | Embedding model ID (e.g. text-embedding-3-small). |
GENERIC_OPEN_AI_EMBEDDING_API_KEY |
(empty) | API key for the embedding endpoint (can differ from the LLM key). |
| Variable | Default | Description |
|---|---|---|
VECTOR_DB |
qdrant |
AnythingLLM vector-store selector. |
QDRANT_ENDPOINT |
http://qdrant:6333 |
In-stack service DNS — change only if you vector-DB elsewhere. |
QDRANT_API_KEY |
(empty) | Qdrant API key. Leave empty when qdrant runs unauthenticated. |
| Variable | Default | Description |
|---|---|---|
SEARXNG_SECRET_KEY |
(tracked placeholder) | Secret key for SearXNG. Substituted at container start into settings.yml by searxng/entrypoint-wrapper.sh — set a unique random string in production. |
SEARXNG_INSTANCE_NAME |
DGX Spark RAG SearXNG |
Display name shown in the SearXNG UI. Same substitution mechanism. |
The tracked searxng/settings.yml contains placeholders
(<SEARXNG_SECRET_KEY>, <SEARXNG_INSTANCE_NAME>); the wrapper mounted into
the container substitutes them from .env at startup without mutating the
host file.
| Variable | Default | Description |
|---|---|---|
ANYTHINGLLM_STORAGE_DIR |
/app/server/storage |
In-container storage directory for AnythingLLM (the host path is ANYTHINGLLM_STORAGE). |
Per-service memory and CPU reservations (deploy.resources constraints).
Tune QDRANT_MEMORY upward if you plan to index millions of vectors.
| Variable | Default | Description |
|---|---|---|
QDRANT_MEMORY |
4G |
Memory limit for the Qdrant service. |
QDRANT_CPUS |
2 |
CPU limit for the Qdrant service. |
SEARXNG_MEMORY |
1G |
Memory limit for the SearXNG service. |
SEARXNG_CPUS |
1 |
CPU limit for the SearXNG service. |
ANYTHINGLLM_MEMORY |
2G |
Memory limit for the AnythingLLM service. |
ANYTHINGLLM_CPUS |
2 |
CPU limit for the AnythingLLM service. |
CADDY_MEMORY |
512M |
Memory limit for the Caddy service. |
CADDY_CPUS |
1 |
CPU limit for the Caddy service. |
This project is licensed under the GNU General Public License v3.0 - see the LICENSE file for details.
See CHANGELOG.md for a history of notable changes to this project.
Contributions are welcome! Please open an issue or submit a pull request.
Happy retrieval-augmented generation!