Why We Run Our Own LLM Stack on Bare Metal Instead of Renting OpenAI Tokens

The cost and sovereignty math after 90 days in production

by
Why We Run Our Own LLM Stack on Bare Metal Instead of Renting OpenAI Tokens

After 90 days running a self-hosted LLM stack on bare metal, we have hard numbers on cost, latency, and sovereignty. Here is the unvarnished truth.

The Setup

We run a cluster of 4 nodes, each with dual AMD EPYC 7763 (64 cores), 512 GB RAM, and 4× NVIDIA A100 80GB SXM. Interconnect is InfiniBand HDR100. The stack: vLLM for inference, Hugging Face Text Generation Inference for embeddings, and a custom Rust-based router for load balancing and failover. We serve a mix of chat (Llama 3 70B), code completion (CodeLlama 34B), and embeddings (E5-mistral-7b).

The Cost Math

OpenAI Pricing (as of 2024)

  • GPT-4 Turbo: $10 per 1M input tokens, $30 per 1M output tokens
  • Codex (CodeLlama equivalent): $0.15 per 1M input, $0.60 per 1M output
  • Ada embeddings: $0.0001 per 1K tokens

Our Monthly Volume

  • 50M input tokens / 10M output tokens (chat)
  • 20M input / 5M output (code)
  • 100M embedding tokens

OpenAI monthly cost:

  • Chat: (50M * $10 + 10M * $30) / 1M = $500 + $300 = $800
  • Code: (20M * $0.15 + 5M * $0.60) / 1M = $3 + $3 = $6
  • Embeddings: 100M * $0.0001 = $10
  • Total: ~$816/month

Our Bare Metal Cost

  • Hardware amortized over 3 years: $200k / 36 = ~$5,556/month
  • Power and cooling: $1,200/month
  • Network and colo: $800/month
  • Total: ~$7,556/month

At first glance, OpenAI is 9x cheaper. But that ignores one key factor: we are not using all 4 GPUs at full capacity. Our average GPU utilization is 35%. The real cost per token is dominated by idle capacity.

Scaling Break-Even

If our volume grows 10x, OpenAI cost becomes $8,160/month. Our hardware cost stays flat. At 20x volume, OpenAI is $16,320/month vs our fixed $7,556. The break-even is around 12x current volume.

But volume is not the only factor. Latency, data control, and customization matter.

Latency Benchmarks

We measured end-to-end latency for a 500-token prompt with 100-token output:

Provider p50 (ms) p99 (ms)
OpenAI GPT-4 Turbo 850 3200
Self-hosted Llama 3 70B (vLLM) 420 1100
Self-hosted CodeLlama 34B 280 750

Self-hosted is 2-3x faster at p50 and 3x more consistent at p99. No shared infrastructure noise.

Data Sovereignty

Our customers handle sensitive legal and financial documents. Sending data to OpenAI exposes us to:

  • Data retention policies (up to 30 days by default)
  • Potential model training on customer data (opt-out required)
  • API availability dependency
  • No offline capability

With self-hosted, all data stays on-premises. We can audit every packet. We can fine-tune models on customer data without leaking it to third parties.

The Hidden Costs of OpenAI

  • Data egress: If you process large documents, tokenization is free, but any external API call incurs network overhead. Our average API call returns ~2 KB; at 50M requests/month, that's 100 GB egress. Not free.
  • Rate limits: OpenAI imposes tiered rate limits. We hit them regularly during batch jobs, requiring queue backpressure and retries. That adds complexity.
  • Model deprecation: OpenAI deprecates models without notice (e.g., GPT-3.5 Turbo 0301). We had to migrate code twice in 2023. Self-hosted models are frozen until we choose to upgrade.

The Fine-Tuning Advantage

We fine-tuned Llama 3 70B on 50K examples of our customer's internal Q&A. Using LoRA with rank 128 on 4 A100s, training took 12 hours. Inference latency increased by only 5% compared to base model. Accuracy on domain-specific questions jumped from 72% to 94%.

OpenAI fine-tuning is available for some models but:

  • You cannot export the fine-tuned model
  • You pay per token trained ($8/1M tokens for GPT-3.5)
  • Data must be sent to OpenAI

With self-hosted, we own the model. We can deploy it anywhere, audit it, and iterate without per-token fees.

Operational Overhead

Running bare metal is not free in labor. Our team of 3 DevOps/SRE engineers spends about 20% of their time on LLM infrastructure:

  • Model updates and rollbacks
  • GPU driver and kernel updates
  • Monitoring and alerting (Prometheus + Grafana)
  • Capacity planning

That's 0.6 FTE, roughly $120k/year. Add that to hardware cost.

Adjusted monthly cost: $7,556 (hardware) + $10,000 (labor) = $17,556/month.

Now break-even against OpenAI is at ~25x current volume. But we are already at 10x growth trajectory, so break-even is 2-3 months away.

When Self-Hosted Makes Sense

  • High volume: >10M tokens/month per model
  • Low latency requirements: <500ms p50
  • Data sensitivity: Cannot send data to third parties
  • Custom models: Fine-tuning or model customization needed
  • Predictable cost: No surprise bills from API usage spikes

When to Stay with OpenAI

  • Low volume: <1M tokens/month
  • Rapid prototyping: No ops team
  • Model diversity: Need access to many different models (GPT-4, Claude, Gemini) without managing infrastructure
  • No data sensitivity: Public data or fully anonymized

Our Architecture

┌─────────────┐     ┌──────────────┐     ┌─────────────────┐
│  Client App  │────▶│  Rust Router  │────▶│  vLLM Instance  │
└─────────────┘     └──────────────┘     └─────────────────┘
                           │
                           ▼
                    ┌──────────────┐
                    │  Model Store  │
                    │  (NFS + S3)   │
                    └──────────────┘

The Rust router handles:

  • Token-aware load balancing
  • Request queuing with priority
  • Model warm-up and scaling
  • Health checks and failover
// Simplified routing logic
async fn route_request(req: InferenceRequest) -> Result<Response> {
    let model = req.model.clone();
    let backend = select_backend(&model).await?;
    let result = backend.infer(req).await?;
    Ok(result)
}

Key Lessons Learned

  1. Don't over-provision. Start with 1 GPU, scale as needed. Idle GPUs are money wasted.
  2. Monitor everything. GPU memory fragmentation can silently degrade throughput. We track kv_cache_utilization per model.
  3. Use optimized inference engines. vLLM with PagedAttention gave us 2x throughput over vanilla Hugging Face.
  4. Batch dynamically. For small requests, batch size 1 is fine. For large, batch up to 64. Adaptive batching reduced p99 latency by 40%.
  5. Plan for model updates. Keep base model checkpoints and LoRA adapters separately. Use a versioned model registry.

The Sovereignty Principle

Beyond cost, the ability to control our own AI infrastructure is strategic. We can:

  • Audit model behavior
  • Block specific outputs
  • Run air-gapped for classified work
  • Fork and modify models

That is not possible with any API provider.

Final Numbers

After 90 days:

  • Total OpenAI avoided cost: ~$2,448
  • Total self-hosted cost: ~$22,668
  • Net loss: ~$20,220

But we now have the infrastructure to scale to 100x volume with zero marginal cost increase. And we have full data sovereignty.

If you are building AI products that handle sensitive data or expect high volume, the math favors self-hosting. If you are prototyping or have low volume, rent tokens.

We chose sovereignty.

#bare-metal#cost-engineering#llm-infrastructure#open-source#self-hosted#sovereign-ai
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related