Open Model Delusion Calculator — the real cost to self-host an open-weights LLM at scale

The self-host LLM cost reality calculator. The weights are free; the cluster to serve them is not. Enter an open-weights LLM's parameters — total and active (MoE) params, precision (FP4/FP8/BF16/FP32), context length, and concurrency — and estimate the real NVIDIA GPU cluster needed to self-host it at scale: H100/H200 PCIe, B200, or GB200 NVL72 counts, node/rack topology, host CPU, RAM, storage, and the order-of-magnitude capex. The honest number behind the hype — not cloud-spot pricing or Mac-mini math.

Presets for current open models including Llama 4 Maverick, DeepSeek V4 Pro, GLM-5.2, Kimi K2.7, MiniMax M3, and Gemma 4.