Migração 100% grátis na contratação semestral · nossa equipe migra tudo de qualquer provedor · novos clientes Migrar agora
Agente de IA no WhatsApp na API Oficial da Meta · somos Tech Provider aprovado · orçamento sob consulta Pedir orçamento
VPS com OpenClaw pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
VPS com Hermes Agent pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
Hospedagem com 30 dias de garantia · Não gostou? Devolvemos 100%, sem perguntas. Ver hospedagem

Install Ollama on a VPS

Install Ollama on an Ubuntu 22.04 VPS with an NVIDIA GPU, expose the REST API over HTTPS via Caddy and serve Llama 3, Qwen and DeepSeek, OpenAI-style.

Ollama is the fastest way to run LLMs on your own server. It exposes an OpenAI-compatible REST API: you change the URL and your code keeps working.

When to use Ollama

  • Dev / experimentation: switch models (Llama 3, Qwen, DeepSeek) with 1 command
  • Light inference: up to ~5 concurrent calls
  • Embeddings with nomic-embed-text or bge-m3
  • Small teams that do not want to set up Triton/vLLM yet

For high concurrency (10+ simultaneous requests), switch to vLLM.

Prerequisites

  • VPS with an NVIDIA GPU (we recommend GPU Estação or GPU Estúdio)
  • Ubuntu 22.04 LTS
  • NVIDIA driver + CUDA 12.x already installed
  • Your own domain (ollama.sua-empresa.com.br)

Confirm the GPU is visible:

nvidia-smi

Expected output (example with an RTX 4090):

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 535.154.05    Driver Version: 535.154.05    CUDA Version: 12.2  |
+-----------------------------------------------------------------------------+
| GPU  Name        | Memory-Usage      | GPU-Util  |
|  0  NVIDIA RTX 4090 | 0MiB / 24576MiB |     0%    |
+-----------------------------------------------------------------------------+

Step 1: install Ollama

curl -fsSL https://ollama.com/install.sh | sh

The script detects the GPU and installs CUDA support automatically. Confirm:

ollama --version
systemctl status ollama

The service is already running on localhost:11434.

Step 2: download and test a model

ollama pull llama3.1:8b
ollama run llama3.1:8b "Em uma frase, por que o céu é azul?"

The first run downloads ~5 GB. After that it runs right away.

Step 3: expose the API over HTTPS

By default, Ollama listens on 127.0.0.1:11434, local only. To access it from outside, we will expose it through Caddy with HTTPS.

Configure Ollama to listen on all interfaces:

sudo systemctl edit ollama.service

Add:

[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=*"

Restart:

sudo systemctl restart ollama

Install Caddy:

sudo apt install -y caddy

Configure /etc/caddy/Caddyfile:

ollama.sua-empresa.com.br {
  reverse_proxy localhost:11434

  # Simple API key auth (Caddy header_up matcher)
  @authorized header Authorization "Bearer rh_ollama_token_secreto_grande"
  handle @authorized {
    reverse_proxy localhost:11434
  }
  handle {
    respond "Unauthorized" 401
  }
}

Restart Caddy:

sudo systemctl restart caddy

Wait for DNS to propagate and for Caddy to issue the SSL certificate. Test:

curl https://ollama.sua-empresa.com.br/api/tags \
  -H "Authorization: Bearer rh_ollama_token_secreto_grande"

Response: the list of downloaded models.

Step 4: use the OpenAI-compatible API

Ollama exposes an OpenAI-compatible /v1/chat/completions:

curl https://ollama.sua-empresa.com.br/v1/chat/completions \
  -H "Authorization: Bearer rh_ollama_token_secreto_grande" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [
      {"role": "system", "content": "Responda em português, breve."},
      {"role": "user", "content": "O que é Q4_K_M?"}
    ]
  }'

Step 5: use it in n8n

Add an OpenAI Chat Model node with:

  • Custom Credential: API Key = rh_ollama_token_secreto_grande
  • Base URL (under advanced): https://ollama.sua-empresa.com.br/v1
  • Model: llama3.1:8b

Done: your n8n now talks to your local GPU.

Embeddings

For RAG, download an embeddings model:

ollama pull nomic-embed-text

Use it in code:

curl https://ollama.sua-empresa.com.br/api/embed \
  -H "Authorization: Bearer rh_ollama_token_secreto_grande" \
  -d '{
    "model": "nomic-embed-text",
    "input": "Texto para gerar embedding"
  }'

Returns a 768-dimension vector.

Monitoring

Watch GPU usage in real time:

watch -n 1 nvidia-smi

And the Ollama logs:

journalctl -u ollama -f

For Grafana, expose metrics via nvidia-dcgm-exporter:

docker run -d --gpus all --rm \
  -p 9400:9400 \
  nvcr.io/nvidia/k8s/dcgm-exporter:3.3.0-3.2.0-ubuntu22.04

Troubleshooting

ProblemLikely causeSolution
cuda: out of memoryModel too largeUse a smaller quantization (:Q4, :Q5_K_M)
Slow response (~5 tok/s)Running on CPUCheck nvidia-smi during inference: if it shows 0%, the GPU is not being used
connection refused on port 11434OLLAMA_HOST was not appliedsudo systemctl daemon-reload && sudo systemctl restart ollama
Caddy does not issue SSLDNS has not propagateddig ollama.sua-empresa.com.br should return the VPS IP

Next steps

Last updated: