Self-Host or Subscribe? The AI Infrastructure Decision.
Model the real cost of running your own AI infrastructure versus managed subscriptions - across vendors, team sizes, and commitment tiers.
Your Setup
Your Team
Number of Developers50
Organisation Scale
⚠ Enterprise Scale - Use as Directional Estimate Only
At 2,000+ developers, self-hosting requires multiple GPU clusters, a dedicated MLOps team, enterprise SLAs, data residency compliance, and custom vendor contracts. This calculator gives a directional cost signal - not a deployment plan. Talk to Gyde for a proper architecture assessment.
Currency
Subscription Cost / Developer / Month
e.g. GitHub Copilot Business $19 · Cursor Teams $40 · Windsurf Teams $40 · Claude Code Team Premium $100
Infrastructure
Vendor
Plan
Model
Number of GPUs4
Selected model sets the minimum. Each additional set of min-GPUs adds a replica.
Effective cost / unit / hr-
Usage Assumptions
Daily Active Users (%)60%
% of developers using AI on an average day
Peak Concurrent Generation (%)30%
% of active users generating tokens simultaneously at peak
Typical AI Response Length (tokens)500
Code generation: 300-1,000. Long explanations: 1,000-4,000.
Hours Operated / Day16h
Infrastructure Detail
Model Replicas1
Each replica improves concurrency and adds redundancy - and doubles GPU count and cost.
Infrastructure Overhead (%)
Storage, networking, DevOps, monitoring, support. Industry norm: 20-35%.
Detailed Vendor Pricing (H100 SXM)
Click a row to apply
Plan
USD/unit/hr
In currency
GPU / Chip Cost / hr (manual override)
Commitment Discount
Applied on top of the rate above. Discount tiers vary by vendor.
Operational Cost (People)
⚠ GPU cost is visible. Engineering cost is invisible - and often larger. These inputs surface the true Total Cost of Ownership.
One-Time Setup Cost
One-time cost to deploy inference server, auth, monitoring, CI/CD, and agent harness. Typically 3-8 weeks of 1-2 senior engineers. Default reflects local market rates for selected currency.
Ongoing Maintenance (hrs/month)
Model upgrades, incident response, monitoring, capacity planning. Typically 1-2 days/week of one engineer.