Self-hosted LLMs · Overview
When it makes sense to run LLMs on your own server with Ollama, vLLM or llama.cpp, and how to choose the right model, hardware and runtime.
Running your own LLM solves 3 problems: predictable cost at high volume, data that cannot leave your infrastructure (LGPD, compliance) and minimal latency when the LLM sits in the same data center as the application.
When it does NOT make sense
If you answer yes to any of these three, stick with GPT/Claude
- Volume < 10 million tokens/month? OpenAI/Anthropic is cheaper
- Need GPT-4o-level reasoning? Open models are still a step behind
- No one to monitor 24/7? A GPU going down at 3 a.m. has no Anthropic support
When it makes total sense
- High volume (50M+ tokens/month): the server pays for itself in 2-3 months
- Sensitive data that cannot go to an external API (healthcare, legal, financial)
- Embeddings at scale (RAG with millions of documents)
- Batch inference (call summaries, ticket classification)
Guides
Install Ollama on a VPS
The fastest way to run Llama 3, Qwen and DeepSeek with an OpenAI-compatible API.
Serve Llama 3 with vLLM
For production: 5x higher throughput than Ollama through continuous batching.
Recommended hardware
| Model | Quantization | Minimum VRAM | Suggested plan |
|---|---|---|---|
| Llama 3 8B | Q4_K_M | 8 GB | GPU Estação |
| Llama 3 70B | Q4_K_M | 48 GB | GPU Estúdio |
| Llama 3.1 405B | Q4_K_M | 240 GB | GPU Cluster |
| Qwen 2.5 72B | Q4_K_M | 48 GB | GPU Estúdio |
| DeepSeek V3 (671B) | FP8 | 8× H100 | on request |
Cost comparison
| Scenario | OpenAI GPT-4o | Llama 3 70B (self-hosted) |
|---|---|---|
| 10M tokens/month | R$ 750 | R$ 4.500 (idle server) |
| 100M tokens/month | R$ 7.500 | R$ 4.500 |
| 1B tokens/month | R$ 75.000 | R$ 4.500 (with queue) |
Above 20-30M tokens/month, self-hosting starts to make sense.
Last updated: