Migração 100% grátis na contratação semestral · nossa equipe migra tudo de qualquer provedor · novos clientes Migrar agora
Agente de IA no WhatsApp na API Oficial da Meta · somos Tech Provider aprovado · orçamento sob consulta Pedir orçamento
VPS com OpenClaw pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
VPS com Hermes Agent pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
Hospedagem com 30 dias de garantia · Não gostou? Devolvemos 100%, sem perguntas. Ver hospedagem

Serve Llama 3 with vLLM

How to install and configure vLLM to serve Llama 3 (8B or 70B) with 5x higher throughput than Ollama thanks to continuous batching on the GPU.

Ollama is great for getting started and for casual use. But in production with concurrent requests, it serializes inference: one slow call holds up all the others. vLLM solves this with continuous batching: it processes multiple requests in parallel in the same GPU forward pass.

When to switch from Ollama to vLLM

ScenarioUse
1-5 simultaneous callsOllama
Dev / experimentationOllama
10+ concurrent callsvLLM
You want an OpenAI-compatible API with logprobsvLLM
Optimize cost per token in productionvLLM

Prerequisites

  • NVIDIA GPU with 16+ GB of VRAM (A100, A6000, RTX 4090). We recommend GPU Estação
  • CUDA 12.1+ installed
  • Python 3.9+
  • Hugging Face token with access to the Llama 3 model

Quick install via Docker

docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HUGGING_FACE_HUB_TOKEN=hf_seu_token" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model meta-llama/Meta-Llama-3-8B-Instruct \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.92

This starts an OpenAI-compatible server at http://localhost:8000/v1.

Test

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "messages": [{"role": "user", "content": "Em uma frase: por que o céu é azul?"}],
    "max_tokens": 100
  }'

Key parameters

  • --max-model-len 8192: context limit. The higher it is, the more VRAM it uses
  • --gpu-memory-utilization 0.92: fraction of VRAM used (leave ~8% free for overhead)
  • --tensor-parallel-size 2: use 2 GPUs in parallel (requires NVLink or fast PCIe interconnect)
  • --quantization awq: use an AWQ-quantized model (1/4 of the VRAM, minimal loss)

Serve Llama 3 70B on a single GPU

On an A100 80GB, use AWQ:

docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HUGGING_FACE_HUB_TOKEN=hf_seu_token" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model casperhansen/llama-3-70b-instruct-awq \
  --max-model-len 8192 \
  --quantization awq \
  --gpu-memory-utilization 0.95

Integrate with n8n

In the HTTP Request node:

  • URL: http://seu-servidor:8000/v1/chat/completions
  • Header: Authorization: Bearer dummy (vLLM does not require a token, but n8n needs the header)
  • Body: OpenAI-compatible payload

Or use the native OpenAI node with the base URL pointing to http://seu-servidor:8000/v1.

Next steps

Last updated: