โšก Inferencing ยท Capacity & Economics

Inference Capacity Planner

Estimate model memory, sustainable request throughput, required replicas and serving cost before you benchmark or deploy.

๐Ÿ”’ Local-first: all calculations run in your browser. No model, workload, or cost data is sent anywhere.

Deployment estimate

Order-of-magnitude planning, not a hardware benchmark.

LIVE

Cost scale

Estimated infrastructure cost from the current assumptions.

Planning summary

Useful numbers at a glance.

Assumptions & formulas

Transparent by design.
Weight and runtime memory

Model weights are estimated as parameter count ร— bytes per weight. A configurable multiplier is then applied as a rough runtime memory allowance.

weight_memory_GB = parameters_billions ร— bytes_per_weight
runtime_memory_GB = weight_memory_GB ร— overhead_multiplier
Sustainable request throughput

The entered token throughput is treated as aggregate useful token throughput for one replica.

sustainable_RPS = tokens_per_second ร— utilization
                  โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
                  input_tokens + output_tokens
Replicas
required_replicas = ceil(target_RPS / sustainable_RPS_per_replica)
Cost
hourly_cost = replicas ร— GPUs_per_replica ร— GPU_hourly_price
daily_cost  = hourly_cost ร— 24
monthly_cost โ‰ˆ daily_cost ร— 30
Important: Real inference performance depends on model architecture, runtime, GPU, quantization format, KV cache, context length, batch size, concurrency, speculative decoding and many other factors. Use this tool for first-pass planning, then benchmark your exact stack.