Send Feedback

TCO Analysis Dashboard

Compare on-premises infrastructure, cloud, and AI API costs.

Calculating costs…

GPU Configuration

Use Case GPU Profiles

This requires 1 GPU servers

Affects both hardware pricing and cloud instance selection

GPU Configuration Summary

Use Case 1
8 GPUs
H2001 servers
p5en.48xlarge • 1 cloud unit
Total (All Use Cases)
8 GPUs
1 compute nodes
Includes the selected server and rack layouts

On-Premises Configuration

Hardware Breakdown

1 compute nodes across the selected server and rack layouts.

GPU server power

System power per server. Enter your server’s rating, or use the platform default.

Auto-selected Ethernet network

$45,000

Sizes switches and connections for your deployment.

Single-homed access

Network details

One server uses its internal GPU interconnect; no external compute fabric is required. In-band and management connectivity are priced separately. Single-homed access; 48 endpoint ports per switch. Uplink and WAN/router ports are quote-dependent. Equipment costs assume new purchases. Shared-fabric allocation depends on the site.

Conventional HGX 400GbE compute fabric: 1 plane, 0 leaf and 0 spine switches. Includes converged in-band and out-of-band management; dedicated storage fabric is opt-in.

Converged in-band switches × 1$30,000
Converged in-band pluggables × 4$3,000
Out-of-band management switches × 1$12,000
Edit unit prices

Hardware Maintenance

Personnel

Facilities Configuration

Platform operating model

A few choices set the inferred platform, security, and observability footprint.

Platform operating model

Secure AI: Enterprise controls for governed AI applications and platforms.

Platform servicesAI security & governanceAudit-ready observability

Recommended topology: 3 control nodes. Production Kubernetes uses a three-node control plane; large clusters add a dedicated management node.

Included platform services

Add prices for the highlighted services to include them in totals.

Shared controls

Cisco AI DefenseAdd price· 1 application
ServiceOn-premisesCloud IaaS
Platform & OS
Commercial OS supportAdd price· 160 coresAdd price· 160 cores
Kubernetes platform supportAdd price· 3 nodesAdd price· 3 nodes
Security & AI governance
Isovalent EnterpriseAdd price· 3 nodesAdd price· 3 nodes
Observability
Splunk Observability Cloud$540/yr· 3 hosts$540/yr· 3 hosts

Cloud Configuration

Cloud hardware to compare

Use Case 1

1 × p5en.48xlarge · 1,128 GB GPU memory · 7,916 BF16 TFLOPS

Instance Configuration

Pricing Configuration

Choose your preferred pricing model

Enter your negotiated enterprise discount

AWS Enterprise Support fee percentage

Additional Storage

Personnel

AWS platform services

Optional add-ons · US East list-rate defaults, editable for your region and usage.

Managed Kubernetes control plane

$0.00/mo

Application and infrastructure logs

$0.00/mo

Infrastructure monitoring

$0.00/mo

Monitoring dashboards

$0.00/mo

Runtime threat detection

$0.00/mo

AI content safety

$0.00/mo

Selected AWS services: $0.00/month · $0.00/year before discounts.

Amazon Linux and AWS Batch have no additional software fee. Compute and storage remain in infrastructure costs.

Calculating Costs

Updating summary...

Cloud Infrastructure

Total CostAdd price
Monthly Average
NPV

On-Premises Infrastructure

Total Cost$0.00
Monthly Average$0.00
NPV$0.00

Comparison

Key Assumptions

Cloud quantities follow the selected hardware or your VRAM/BF16 capacity match, rounded up to whole instances.

Management software for compute and networking is included in the per unit cost.

EC2 On-Demand prices come from AWS’s live pricing sources and are cached for up to 24 hours.

SaaS Token API pricing (OpenAI, Anthropic, Gemini, Bedrock) is updated periodically based on published rates.

Personnel costs include benefits and overhead at 30% of base salary.

GPU server power follows installed GPU TDP plus platform host overhead. Rack capacity uses peak power; training energy uses scheduled activity with a 30% idle-GPU allowance. Shared servers and switches use a 1.2 kW planning allowance each. PUE applies to electricity.

Inference runs continuously; training follows the configured run schedule.

Token API pricing uses the input/output token ratio from GenAI Sizing inputs (avgInputTokens/avgOutputTokens) and requires GenAI Sizing calculator results.

Effective cost per 1M output tokens accounts for both compute-bound prefill and memory-bound decode phases using: (requests/sec × output_tokens/request), which accurately reflects workloads with varying input/output ratios.

Networking includes converged in-band and out-of-band management switches. Compute fabric follows the selected topology policy; dedicated storage fabric is opt-in. WAN uplinks and cables require site-specific quotes.

Cloud compute uses the selected instance’s AWS price or your region-specific quote.