RAG with Qdrant + LangChain
How to build a Retrieval-Augmented Generation (RAG) stack with Qdrant as the vector database and LangChain as the orchestrator on a Rollin Host VPS.
When to use RAG
Use Retrieval-Augmented Generation (RAG) when you need the AI to answer questions about specific content (manuals, an internal knowledge base, product documentation) without training a model from scratch. The LLM still generates the answer, but with context retrieved from a vector database.
Architecture
PDF / Markdown / Notion ──▶ embeddings ──▶ Qdrant
│
User question ──▶ embedding ──▶ search ──▶ context + LLM ──▶ answer
Stack
- Qdrant: vector database, runs in Docker
- LangChain (Python or Node): orchestrates chunks, embeddings and the prompt
- OpenAI ada-002 or self-hosted bge-m3: embeddings model
- Server: VPS Plus (8 GB RAM, 4 vCPU, 160 GB)
Next steps
We will soon publish the full tutorial with ready-to-use code. In the meantime:
- Install Ollama on a VPS (to run embeddings locally)
- n8n + EvolutionAPI + OpenAI (to connect RAG to WhatsApp)
Last updated: