Install Ollama on a VPS
Install Ollama on an Ubuntu 22.04 VPS with an NVIDIA GPU, expose the REST API over HTTPS via Caddy and serve Llama 3, Qwen and DeepSeek, OpenAI-style.
Ollama is the fastest way to run LLMs on your own server. It exposes an OpenAI-compatible REST API: you change the URL and your code keeps working.
When to use Ollama
- Dev / experimentation: switch models (Llama 3, Qwen, DeepSeek) with 1 command
- Light inference: up to ~5 concurrent calls
- Embeddings with
nomic-embed-textorbge-m3 - Small teams that do not want to set up Triton/vLLM yet
For high concurrency (10+ simultaneous requests), switch to vLLM.
Prerequisites
- VPS with an NVIDIA GPU (we recommend GPU Estação or GPU Estúdio)
- Ubuntu 22.04 LTS
- NVIDIA driver + CUDA 12.x already installed
- Your own domain (
ollama.sua-empresa.com.br)
Confirm the GPU is visible:
nvidia-smi
Expected output (example with an RTX 4090):
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 535.154.05 Driver Version: 535.154.05 CUDA Version: 12.2 |
+-----------------------------------------------------------------------------+
| GPU Name | Memory-Usage | GPU-Util |
| 0 NVIDIA RTX 4090 | 0MiB / 24576MiB | 0% |
+-----------------------------------------------------------------------------+
Step 1: install Ollama
curl -fsSL https://ollama.com/install.sh | sh
The script detects the GPU and installs CUDA support automatically. Confirm:
ollama --version
systemctl status ollama
The service is already running on localhost:11434.
Step 2: download and test a model
ollama pull llama3.1:8b
ollama run llama3.1:8b "Em uma frase, por que o céu é azul?"
The first run downloads ~5 GB. After that it runs right away.
Step 3: expose the API over HTTPS
By default, Ollama listens on 127.0.0.1:11434, local only. To access it from outside, we will expose it through Caddy with HTTPS.
Configure Ollama to listen on all interfaces:
sudo systemctl edit ollama.service
Add:
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=*"
Restart:
sudo systemctl restart ollama
Install Caddy:
sudo apt install -y caddy
Configure /etc/caddy/Caddyfile:
ollama.sua-empresa.com.br {
reverse_proxy localhost:11434
# Simple API key auth (Caddy header_up matcher)
@authorized header Authorization "Bearer rh_ollama_token_secreto_grande"
handle @authorized {
reverse_proxy localhost:11434
}
handle {
respond "Unauthorized" 401
}
}
Restart Caddy:
sudo systemctl restart caddy
Wait for DNS to propagate and for Caddy to issue the SSL certificate. Test:
curl https://ollama.sua-empresa.com.br/api/tags \
-H "Authorization: Bearer rh_ollama_token_secreto_grande"
Response: the list of downloaded models.
Step 4: use the OpenAI-compatible API
Ollama exposes an OpenAI-compatible /v1/chat/completions:
curl https://ollama.sua-empresa.com.br/v1/chat/completions \
-H "Authorization: Bearer rh_ollama_token_secreto_grande" \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [
{"role": "system", "content": "Responda em português, breve."},
{"role": "user", "content": "O que é Q4_K_M?"}
]
}'
Step 5: use it in n8n
Add an OpenAI Chat Model node with:
- Custom Credential: API Key =
rh_ollama_token_secreto_grande - Base URL (under advanced):
https://ollama.sua-empresa.com.br/v1 - Model:
llama3.1:8b
Done: your n8n now talks to your local GPU.
Embeddings
For RAG, download an embeddings model:
ollama pull nomic-embed-text
Use it in code:
curl https://ollama.sua-empresa.com.br/api/embed \
-H "Authorization: Bearer rh_ollama_token_secreto_grande" \
-d '{
"model": "nomic-embed-text",
"input": "Texto para gerar embedding"
}'
Returns a 768-dimension vector.
Monitoring
Watch GPU usage in real time:
watch -n 1 nvidia-smi
And the Ollama logs:
journalctl -u ollama -f
For Grafana, expose metrics via nvidia-dcgm-exporter:
docker run -d --gpus all --rm \
-p 9400:9400 \
nvcr.io/nvidia/k8s/dcgm-exporter:3.3.0-3.2.0-ubuntu22.04
Troubleshooting
| Problem | Likely cause | Solution |
|---|---|---|
cuda: out of memory | Model too large | Use a smaller quantization (:Q4, :Q5_K_M) |
| Slow response (~5 tok/s) | Running on CPU | Check nvidia-smi during inference: if it shows 0%, the GPU is not being used |
connection refused on port 11434 | OLLAMA_HOST was not applied | sudo systemctl daemon-reload && sudo systemctl restart ollama |
| Caddy does not issue SSL | DNS has not propagated | dig ollama.sua-empresa.com.br should return the VPS IP |
Next steps
- Serve Llama 3 with vLLM (for high concurrency)
- RAG with Qdrant + LangChain (Ollama + vector database)
- n8n + EvolutionAPI stack (replace the OpenAI node with your Ollama)
Last updated: