Migração 100% grátis na contratação semestral · nossa equipe migra tudo de qualquer provedor · novos clientes Migrar agora
Agente de IA no WhatsApp na API Oficial da Meta · somos Tech Provider aprovado · orçamento sob consulta Pedir orçamento
VPS com OpenClaw pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
VPS com Hermes Agent pré-instalado · a partir de R$ 56,90/mês · gerenciada pela Rollin Quero a VPS
Hospedagem com 30 dias de garantia · Não gostou? Devolvemos 100%, sem perguntas. Ver hospedagem

Self-hosted LLMs · Overview

When it makes sense to run LLMs on your own server with Ollama, vLLM or llama.cpp, and how to choose the right model, hardware and runtime.

Running your own LLM solves 3 problems: predictable cost at high volume, data that cannot leave your infrastructure (LGPD, compliance) and minimal latency when the LLM sits in the same data center as the application.

When it does NOT make sense

If you answer yes to any of these three, stick with GPT/Claude

  • Volume < 10 million tokens/month? OpenAI/Anthropic is cheaper
  • Need GPT-4o-level reasoning? Open models are still a step behind
  • No one to monitor 24/7? A GPU going down at 3 a.m. has no Anthropic support

When it makes total sense

  • High volume (50M+ tokens/month): the server pays for itself in 2-3 months
  • Sensitive data that cannot go to an external API (healthcare, legal, financial)
  • Embeddings at scale (RAG with millions of documents)
  • Batch inference (call summaries, ticket classification)

Guides

ModelQuantizationMinimum VRAMSuggested plan
Llama 3 8BQ4_K_M8 GBGPU Estação
Llama 3 70BQ4_K_M48 GBGPU Estúdio
Llama 3.1 405BQ4_K_M240 GBGPU Cluster
Qwen 2.5 72BQ4_K_M48 GBGPU Estúdio
DeepSeek V3 (671B)FP88× H100on request

Cost comparison

ScenarioOpenAI GPT-4oLlama 3 70B (self-hosted)
10M tokens/monthR$ 750R$ 4.500 (idle server)
100M tokens/monthR$ 7.500R$ 4.500
1B tokens/monthR$ 75.000R$ 4.500 (with queue)

Above 20-30M tokens/month, self-hosting starts to make sense.

Last updated: