Serve Llama 3 with vLLM
How to install and configure vLLM to serve Llama 3 (8B or 70B) with 5x higher throughput than Ollama thanks to continuous batching on the GPU.
Ollama is great for getting started and for casual use. But in production with concurrent requests, it serializes inference: one slow call holds up all the others. vLLM solves this with continuous batching: it processes multiple requests in parallel in the same GPU forward pass.
When to switch from Ollama to vLLM
| Scenario | Use |
|---|---|
| 1-5 simultaneous calls | Ollama |
| Dev / experimentation | Ollama |
| 10+ concurrent calls | vLLM |
| You want an OpenAI-compatible API with logprobs | vLLM |
| Optimize cost per token in production | vLLM |
Prerequisites
- NVIDIA GPU with 16+ GB of VRAM (A100, A6000, RTX 4090). We recommend GPU Estação
- CUDA 12.1+ installed
- Python 3.9+
- Hugging Face token with access to the Llama 3 model
Quick install via Docker
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=hf_seu_token" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--max-model-len 8192 \
--gpu-memory-utilization 0.92
This starts an OpenAI-compatible server at http://localhost:8000/v1.
Test
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3-8B-Instruct",
"messages": [{"role": "user", "content": "Em uma frase: por que o céu é azul?"}],
"max_tokens": 100
}'
Key parameters
--max-model-len 8192: context limit. The higher it is, the more VRAM it uses--gpu-memory-utilization 0.92: fraction of VRAM used (leave ~8% free for overhead)--tensor-parallel-size 2: use 2 GPUs in parallel (requires NVLink or fast PCIe interconnect)--quantization awq: use an AWQ-quantized model (1/4 of the VRAM, minimal loss)
Serve Llama 3 70B on a single GPU
On an A100 80GB, use AWQ:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=hf_seu_token" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model casperhansen/llama-3-70b-instruct-awq \
--max-model-len 8192 \
--quantization awq \
--gpu-memory-utilization 0.95
Integrate with n8n
In the HTTP Request node:
- URL:
http://seu-servidor:8000/v1/chat/completions - Header:
Authorization: Bearer dummy(vLLM does not require a token, but n8n needs the header) - Body: OpenAI-compatible payload
Or use the native OpenAI node with the base URL pointing to http://seu-servidor:8000/v1.
Next steps
- Install Ollama on a VPS
- n8n + EvolutionAPI + OpenAI: replace the OpenAI node with your vLLM
Last updated: