Deployment estimate
Order-of-magnitude planning, not a hardware benchmark.
Cost scale
Estimated infrastructure cost from the current assumptions.
Planning summary
Useful numbers at a glance.
Assumptions & formulas
Transparent by design.
Weight and runtime memory
Model weights are estimated as parameter count ร bytes per weight. A configurable multiplier is then applied as a rough runtime memory allowance.
weight_memory_GB = parameters_billions ร bytes_per_weight runtime_memory_GB = weight_memory_GB ร overhead_multiplier
Sustainable request throughput
The entered token throughput is treated as aggregate useful token throughput for one replica.
sustainable_RPS = tokens_per_second ร utilization
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
input_tokens + output_tokens
Replicas
required_replicas = ceil(target_RPS / sustainable_RPS_per_replica)
Cost
hourly_cost = replicas ร GPUs_per_replica ร GPU_hourly_price daily_cost = hourly_cost ร 24 monthly_cost โ daily_cost ร 30
Important: Real inference performance depends on model architecture, runtime, GPU, quantization format, KV cache, context length, batch size, concurrency, speculative decoding and many other factors. Use this tool for first-pass planning, then benchmark your exact stack.