Skip to content

Repository files navigation

DGX Spark RAG

A fully containerized Retrieval Augmented Generation (RAG) stack for any amd64 (or arm64) Linux host with Docker. Qdrant (vector database), SearXNG (private search engine), and AnythingLLM (web UI + RAG orchestrator) in a single Docker Compose deployment — with the LLM and embeddings served by any external OpenAI-compatible API (vLLM, llama.cpp server, OpenRouter, OpenAI, …).

Services

  • AnythingLLM: RAG orchestrator + web UI. Talks to your external OpenAI-compatible LLM/embedding API and to the local Qdrant/SearXNG services.
  • Qdrant: High-performance vector database for efficient similarity search.
  • SearXNG: Privacy-respecting metasearch engine to augment RAG answers with web results.
  • Caddy: TLS reverse proxy in front of the UI/APIs (loopback-only).

There is no bundled LLM server: the stack expects an OpenAI-compatible endpoint you already run (e.g. vLLM at http://<host>:8000/v1) and sends both chat completions and embeddings there. The stack is hardware-agnostic: it reserves no GPU devices and ships no accelerator-specific tuning. If your LLM host happens to have a GPU, the provider process serves it — these containers are CPU-only.

Architecture

flowchart LR
    UI[AnythingLLM UI :3001] --> CMD[caddy :8443 TLS]
    CMD -- /api/vectors/* --> QDR[Qdrant :6333]
    CMD -- /search/* --> SX[SearXNG :8080]
    CMD -- /* --> UI
    UI -- OpenAI-compatible /v1/chat/completions --> EXT[External LLM API]
    UI -- OpenAI-compatible /v1/embeddings --> EXT
    QDR --> STO[(qdrant_storage)]
Loading

Prerequisites

  • Docker Engine (version 20.10+)
  • Docker Compose (v2.0+)
  • An OpenAI-compatible API endpoint reachable from the AnythingLLM container. The default .env assumes a server on the Docker host at http://host.docker.internal:8000/v1 (vLLM's default port); the compose file adds the host-gateway mapping so that name resolves inside the container. On Linux hosts running Docker < 20.10 this mapping is implicit; on rootless/podman setups it may need --network=host-style alternatives.

Installation

  1. Clone the repository:

    git clone https://github.com/amasu/dgx-spark-rag.git
    cd dgx-spark-rag
  2. Copy the environment file:

    cp .env.example .env

    Point GENERIC_OPEN_AI_BASE_PATH / EMBEDDING_BASE_PATH at your OpenAI-compatible endpoint and set the model IDs + API keys (see Configuration).

  3. Start the services:

    docker compose pull
    docker compose up -d

Usage

Once the stack is running, access the services at:

TLS note (Caddy)

https://localhost:8443 is served with a self-signed certificate (tls internal). Browsers and API clients will show a certificate warning on first connect — accept it (or point your client's CA bundle at Caddy's internal CA in the caddy_data volume). The per-service ports above are plain HTTP on loopback only; nothing is exposed beyond 127.0.0.1.

Choosing the external provider

Everything is configured through .env — nothing is wired to a specific vendor. Common layouts:

Provider *_BASE_PATH Model IDs
vLLM on this host http://host.docker.internal:8000/v1 served model name (vllm list-served-models)
llama.cpp server http://host.docker.internal:8080/v1 -m alias (--alias)
OpenRouter https://openrouter.ai/api/v1 vendor/model:tag + API key
OpenAI https://api.openai.com/v1 e.g. gpt-4o-mini, text-embedding-3-small

The embedding endpoint must implement POST /v1/embeddings. If your embedding API needs a different host than the LLM, set EMBEDDING_BASE_PATH separately (both default to the same host in .env.example).

First-run behavior: AnythingLLM consumes the LLM_PROVIDER / GENERIC_OPEN_AI_* / EMBEDDING_* / VECTOR_DB variables on boot to build its system preferences — a fresh anythingllm_storage volume starts already wired to your provider. If you previously ran an older stack variant, delete or archive ./anythingllm_storage once so the new defaults take effect (profiles/chats live there and keep stale provider settings).

Customizing SearXNG

SearXNG configuration is stored in the ./searxng directory. Modify settings there and restart the container:

docker compose restart searxng

Persistent Data

All services store data in named volumes or bind mounts:

  • Qdrant: ./qdrant_storage
  • AnythingLLM: ./anythingllm_storage
  • SearXNG: ./searxng (configuration)

Configuration

All runtime configuration lives in the .env file (copy it from .env.example). Docker Compose reads these variables automatically. Every variable is wired into docker-compose.yml with ${VAR:-default} fallbacks, so the stack also renders without an .env (fresh-clone safe).

All services are pinned to multi-arch images and respect their standard deploy.resources CPU/memory limits — no core-pinning or accelerator configuration, so the same files deploy identically on any machine.

Ports

Variable Default Description
QDRANT_PORT 6333 Host port for the Qdrant HTTP/REST API.
SEARXNG_PORT 8181 Host port for the SearXNG web UI.
ANYTHINGLLM_PORT 3001 Host port for the AnythingLLM web UI.
CADDY_PORT 8443 Host port for the Caddy TLS entry point.

Storage paths

Variable Default Description
QDRANT_STORAGE ./qdrant_storage Bind-mount path for Qdrant vector data.
ANYTHINGLLM_STORAGE ./anythingllm_storage Bind-mount path for AnythingLLM data on the host.
SEARXNG_CONFIG ./searxng Bind-mount path for SearXNG configuration files.

LLM provider (OpenAI-compatible)

Variable Default Description
LLM_PROVIDER generic-openai AnythingLLM provider selector — do not change unless you switch engines.
GENERIC_OPEN_AI_BASE_PATH http://host.docker.internal:8000/v1 Base URL of the OpenAI-compatible chat-completions API. host.docker.internal resolves to the Docker host (compose maps it via host-gateway).
GENERIC_OPEN_AI_API_KEY (empty) API key sent as Authorization: Bearer …. Dummy value if your server doesn't check keys.
GENERIC_OPEN_AI_MODEL_PREF (empty) Model ID the provider serves (e.g. the vLLM --served-model-name).

Embeddings (OpenAI-compatible)

Variable Default Description
EMBEDDING_ENGINE generic-openai AnythingLLM embedding-engine selector.
EMBEDDING_BASE_PATH http://host.docker.internal:8000/v1 Base URL of the /v1/embeddings endpoint.
EMBEDDING_MODEL_PREF (empty) Embedding model ID (e.g. text-embedding-3-small).
GENERIC_OPEN_AI_EMBEDDING_API_KEY (empty) API key for the embedding endpoint (can differ from the LLM key).

Vector DB (local Qdrant)

Variable Default Description
VECTOR_DB qdrant AnythingLLM vector-store selector.
QDRANT_ENDPOINT http://qdrant:6333 In-stack service DNS — change only if you vector-DB elsewhere.
QDRANT_API_KEY (empty) Qdrant API key. Leave empty when qdrant runs unauthenticated.

SearXNG

Variable Default Description
SEARXNG_SECRET_KEY (tracked placeholder) Secret key for SearXNG. Substituted at container start into settings.yml by searxng/entrypoint-wrapper.sh — set a unique random string in production.
SEARXNG_INSTANCE_NAME DGX Spark RAG SearXNG Display name shown in the SearXNG UI. Same substitution mechanism.

The tracked searxng/settings.yml contains placeholders (<SEARXNG_SECRET_KEY>, <SEARXNG_INSTANCE_NAME>); the wrapper mounted into the container substitutes them from .env at startup without mutating the host file.

AnythingLLM

Variable Default Description
ANYTHINGLLM_STORAGE_DIR /app/server/storage In-container storage directory for AnythingLLM (the host path is ANYTHINGLLM_STORAGE).

Resource Limits

Per-service memory and CPU reservations (deploy.resources constraints). Tune QDRANT_MEMORY upward if you plan to index millions of vectors.

Variable Default Description
QDRANT_MEMORY 4G Memory limit for the Qdrant service.
QDRANT_CPUS 2 CPU limit for the Qdrant service.
SEARXNG_MEMORY 1G Memory limit for the SearXNG service.
SEARXNG_CPUS 1 CPU limit for the SearXNG service.
ANYTHINGLLM_MEMORY 2G Memory limit for the AnythingLLM service.
ANYTHINGLLM_CPUS 2 CPU limit for the AnythingLLM service.
CADDY_MEMORY 512M Memory limit for the Caddy service.
CADDY_CPUS 1 CPU limit for the Caddy service.

License

This project is licensed under the GNU General Public License v3.0 - see the LICENSE file for details.

Acknowledgments

Changelog

See CHANGELOG.md for a history of notable changes to this project.

Contributing

Contributions are welcome! Please open an issue or submit a pull request.


Happy retrieval-augmented generation!

About

Retrieval Augmented Generation containerized for Nvidia DGX Spark

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages