Why We Self-Host Inference: The Sovereignty and Cost Case Against API-Only AI

Lessons from running 70B models on commodity hardware instead of paying per token to closed providers.

by
Why We Self-Host Inference: The Sovereignty and Cost Case Against API-Only AI

Every time you call an API like OpenAI, Anthropic, or Cohere, you're trading sovereignty for convenience. Your prompts, your data, your business logic — all pass through someone else's infrastructure. The cost isn't just monetary. It's a leak of strategic intelligence, a dependency on a pricing model that can change overnight, and a hard ceiling on customization.

At The Sovereign Stack, we've spent the last 18 months moving inference workloads from cloud APIs to self-hosted GPUs. This article explains why, with real numbers, hardware choices, and the operational trade-offs you need to know.

The Real Cost of API Tokens

Let's start with the obvious: money. A typical API call to GPT-4o costs about $0.01 per 1K input tokens and $0.03 per 1K output tokens. For a 2K-token conversation (common in customer support or coding assistants), that's $0.05 per call. At 10,000 calls/day, that's $500/day or $15,000/month.

Now consider self-hosting. We run Llama 3.1 70B on a single node with 4× NVIDIA RTX 6000 Ada (48GB each). Capital cost: ~$40,000. Power: ~1.5kW at full load, roughly $150/month. With vLLM and FP8 quantization, we serve ~30 tokens/second per user, handling 50 concurrent users easily. Our cost per 1K output tokens? About $0.003 — 10x cheaper than GPT-4o.

Amortized over 3 years, that's $40,000 + $5,400 power = $45,400, or $1,261/month. For $15,000/month of API calls, you'd spend $540,000 over the same period. The savings alone justify self-hosting for any workload above a few thousand calls daily.

But cost is only half the story.

Data Sovereignty: The Hidden Leak

Every API call sends your data to a third party. Even with zero-retention policies, the data transits their network, touches their logs, and may be used for model improvements unless explicitly opted out. For regulated industries (healthcare, finance, defense), that's a non-starter.

We've seen contracts where an API provider's terms changed mid-project, forcing a migration. Self-hosting eliminates that risk. Your data stays on your hardware, behind your firewall. No metadata leaks, no API key theft, no vendor lock-in.

Customization: LoRA and Beyond

APIs give you a black box. You can't fine-tune on proprietary data, adjust tokenizer behavior, or swap attention mechanisms. With self-hosted inference, we run multiple LoRA adapters on the same base model, switching between them per request via a simple routing layer.

# Example: routing to different LoRA adapters based on user context
from vllm import LLM, SamplingParams

llm = LLM(
    model="/models/llama-3.1-70b",
    enable_lora=True,
    max_lora_rank=64
)

# User from finance team gets compliance-tuned adapter
if user.team == "finance":
    lora_path = "/adapters/compliance-v2"
else:
    lora_path = "/adapters/general-v3"

output = llm.generate(
    prompts=[user_prompt],
    sampling_params=SamplingParams(temperature=0.1, max_tokens=512),
    lora_request={"lora_name": "adapter", "lora_path": lora_path}
)

No API provider offers this level of granularity without massive cost. Self-hosting means you control the entire stack.

The Hardware Reality Check

Self-hosting isn't free. You need GPUs, networking, cooling, and ops time. Here's our current stack:

  • GPUs: 4× NVIDIA RTX 6000 Ada (48GB each) — PCIe 4.0, NVLink bridges for inter-GPU communication.
  • CPU: AMD EPYC 7763 (64 cores) — enough to feed the GPUs without bottleneck.
  • RAM: 512GB DDR4-3200 — for KV cache and model weights.
  • Storage: 2× 3.84TB NVMe in RAID1 for model checkpoints, 2× 7.68TB for datasets.
  • Networking: 2× 100GbE ConnectX-6 for inference traffic and model sync.
  • Software: Ubuntu 22.04, Docker, vLLM 0.6.0, custom routing layer in Python.

Total build: ~$42,000. We run 24/7 with 99.97% uptime over 12 months.

Operational Lessons Learned

  1. Quantization is non-negotiable. FP8 halves memory bandwidth and VRAM usage with negligible quality loss. We use AWQ for 4-bit quantization on smaller models (7B-13B) and FP8 for 70B.

  2. KV cache management matters. With long contexts (32K tokens), KV cache can exceed model weights. We use PagedAttention in vLLM, which reduces memory fragmentation and allows higher concurrency.

  3. Batch scheduling is critical. Dynamic batching with continuous batching (vLLM's default) gives 2-3x throughput over static batching. We tune max_num_batched_tokens and max_num_seqs per workload.

  4. Monitoring is not optional. We use Prometheus + Grafana for GPU utilization, memory, latency p50/p99, and error rates. Alerts fire when VRAM usage exceeds 85% or latency p99 goes above 2 seconds.

  5. Firmware and driver hell. We maintain a golden image with CUDA 12.4, driver 550.54, and firmware updates for the EPYC. A misaligned firmware version once caused NVLink to fall back to PCIe, halving throughput. Test upgrades in staging.

When APIs Still Make Sense

I'm not dogmatic. For prototyping, bursty workloads, or when you need access to models you can't self-host (e.g., GPT-4o vision or Claude 3 Opus), APIs are fine. But for production inference with predictable load, self-hosting wins on cost, sovereignty, and customization.

We still use APIs for a few specialized tasks: multimodal embeddings (CLIP-based) and web search augmentation. But the core conversational and reasoning engines are self-hosted.

The Bottom Line

Self-hosting inference is a strategic investment. It cuts costs by 10x, eliminates data leakage, and gives you full control over model behavior. The upfront hardware cost is real, but the payback period is under 6 months for moderate workloads.

If you're running more than 5,000 API calls per day, do the math. Build a proof of concept with vLLM and a single 24GB GPU. You'll be surprised how far you can get.

Our next post will cover distributed inference across multiple nodes with Tensor Parallelism and how we handle failover. Stay tuned.

#data-sovereignty#inference-costs#llama-70b#on-prem-ai#open-source-ai#self-hosted-llm#vllm
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related