Azure GPU and CPU Cost Reference for Open-Source LLM Fine-Tuning and Serving

Single-VM and AKS cluster costs for fine-tuning and serving open-source LLMs on Azure — GPU/CPU pricing (PAYG, Spot, Reserved), itemized cluster builds, alternatives to AKS, and leadership pivot tables. A cost addendum to the fine-tuning platforms reference.

#azure#gpu#aks#cost#llm#fine-tuning#serving#infrastructure

Addendum to Post-Training Techniques and Training Platforms. That reference covers the techniques and tools; this page covers what they cost to run on Azure.

Leadership-ready reference — Region: East US (primary) · Currency: USD · Prices checked: 11 August 2026

Read this first. All list prices below are Azure public pay-as-you-go (PAYG) rates for Linux VMs, verified against Vantage (instances.vantage.sh), CloudPrice, SpareCores and Microsoft pricing pages as of 11 Aug 2026. Azure prices change frequently and must be re-validated in the Azure Pricing Calculator before any budget commitment. As a large agricultural enterprise, your organization almost certainly has an Enterprise Agreement (EA) or CSP discount of 10–30% off these list prices — treat every number here as a ceiling, not a final cost. Azure Hybrid Benefit further reduces cost where you own Windows/SQL licensing (not usually relevant to Linux GPU nodes).

Table of Contents

  1. Part 1 — Single VM Costs (GPU & CPU)
  2. Part 2 — AKS Cluster Costs
  3. Part 3 — Alternatives to AKS
  4. Part 4 — Building a Platform that Scales GPU Capacity
  5. Part 5 — Leadership Pivot Tables
  6. Key Takeaways for Leadership

TL;DR

  • The single biggest cost lever is GPU utilization, not SKU choice — an idle A100 (~$2,681/month) costs exactly the same as a fully-loaded one. For your 8B LoRA + vLLM multi-adapter plan, a 1× A100 80GB (NC24ads_A100_v4, ~$3.67/hr PAYG) or 1× H100 NVL (NC40ads_H100_v5, ~$6.98/hr) node covers both fine-tuning and serving; running one node always-on on a 3-year reservation plus Spot for training bursts is the cost-optimal pattern.
  • Build-vs-buy crosses over at roughly 30–55 million tokens/day of sustained inference: below that, Azure AI Foundry serverless per-token APIs (Llama-class models at sub-$1 per 1M tokens) are cheaper and simpler than a self-hosted always-on GPU; above it, a reserved AKS GPU node wins decisively and the gap widens with volume.
  • A departmental production platform lands around $5,000–$12,000/month (one always-on A100 inference node + spot training + storage + monitoring on AKS Standard); an enterprise-scale multi-region platform runs $40,000–$120,000+/month. Reserved instances and Spot cut 35–65% off compute, and engineering FTE cost (~$450k–$900k/yr for 2–4 engineers) usually exceeds the infrastructure bill — model it explicitly.

Key Findings

  • A100 vs H100/H200 deployment posture: The A100 generation (NC A100 v4, ND A100 v4) is still fully offered and is the cost-efficient workhorse for 8B–32B work, but Azure is directing new capacity toward H100/H200 and Blackwell. NDv4 (A100 SXM) is described as “still available but being phased toward end-of-life for new workloads.” H200 (ND96isr_H200_v5, $110.24/hr) and GB200 (ND GB200 v6, now GA) currently have no published Spot or Reserved discount — you pay full on-demand.
  • GPU reserved-instance dollar rates are not published by aggregators (Vantage and CloudPrice show “N/A” for the 1-Year and 3-Year Reserved fields on every GPU SKU). The defensible general rule is ~35% off for 1-year reservations (confirmed by Thunder Compute); a ~55% 3-year figure is an industry assumption, not Microsoft-confirmed for GPU SKUs. Get exact reservation quotes from the Azure Portal purchase flow.
  • The surprise line items are monitoring and egress, not compute add-ons. Container Insights/Log Analytics ingest at $2.30/GB Analytics Logs (after 5 GB free/month per billing account; Microsoft Azure Monitor pricing page and MonitoringCost.com, 2026), and a chatty GPU cluster can push tens of GB/day; internet egress is $0.087/GB after 100 GB free.
  • Serverless GPU (Azure Container Apps) is GA with scale-to-zero and per-second billing for T4 and A100. Microsoft’s “Announcing GA for Azure Container Apps Serverless GPUs” (Microsoft Community Hub, modified 27 Mar 2025) states GA is “supported in West US 3, Australia East, and Sweden Central” with default A100/T4 quota for EA customers; a 2025 recap notes GA later expanded to 11 additional regions. This is the best fit for spiky, low-duty-cycle inference.
  • Region choice swings GPU cost 20–33%. US regions (East US, East US 2, South Central US) are cheapest; Western Europe ~+19%, Southeast Asia ~+30%, Australia East ~+33%.

Details

PART 1 — SINGLE VM COSTS (GPU and CPU)

1.1 GPU VMs — specs and PAYG/Spot pricing (East US, Linux, per hour; monthly at 730 h)

SKUvCPURAM GiBGPUVRAM totalTemp/NVMeInfiniBandPAYG $/hrSpot $/hr (disc.)PAYG $/mo
NC4as_T4_v34281× T4 16GB16 GB180 GBNo$0.526$0.153 (~71%)$383.98
NC8as_T4_v38561× T4 16GB16 GB360 GBNo$0.752$0.257 (~66%)$548.96
NC16as_T4_v3161101× T4 16GB16 GB360 GBNo~$1.20~$0.42~$876
NC64as_T4_v3644404× T4 16GB64 GB2880 GBNo$4.352$1.859 (~57%)$3,176.96
NC24ads_A100_v4242201× A100 80GB80 GB1026 GBNo$3.673$0.679 (~82%)$2,681.29
NC48ads_A100_v4484402× A100 80GB160 GB2052 GBNo$7.346~$1.36$5,362.58
NC96ads_A100_v4968804× A100 80GB320 GB4104 GBNo$14.692$2.715 (~82%)$10,725.16
ND96asr_v4969008× A100 40GB SXM320 GB10.7 TB200 Gb/s$27.197$5.983 (~78%)$19,853.81
ND96amsr_A100_v49619008× A100 80GB SXM640 GB10.7 TB200 Gb/s$32.77$8.52 (~74%)$23,922.10
NC40ads_H100_v5403201× H100 NVL 94GB94 GB3576 GBNo$6.98$1.403 (~80%)$5,095.40
NC80adis_H100_v5806402× H100 NVL 94GB188 GB7152 GBNo$13.96$2.806 (~80%)$10,190.80
ND96isr_H100_v59619008× H100 80GB SXM5640 GB28 TB3200 Gb/s$98.32$18.17 (~82%)$71,773.60
ND96isr_H200_v59618508× H200 141GB1128 GB31.8 TB3200 Gb/s$110.24none published$80,475.20
ND96isr_MI300X_v5 (AMD)9618508× MI300X 192GB1536 GB31.8 TBRDMA~$48 (region-dep., $8.87 low end)$8.87~$35,040
ND GB200 v6 (Blackwell)1289004× GB200 Grace-BlackwellNDR IBNot published (GA)none

A10 fractional GPUs (NVads_A10_v5 — VDI/visualization-oriented, usable for small inference):

SKUvCPURAM GiBA10 fraction / VRAMPAYG $/hrSpot $/hrPAYG $/mo
NV6ads_A10_v56551/6 A10 / 4 GB~$0.45$0.084~$330
NV12ads_A10_v5121101/3 A10 / 8 GB$0.908$0.168$662.84
NV18ads_A10_v5182201/2 A10 / 12 GB~$1.36~$0.25~$993
NV36ads_A10_v5364401× A10 / 24 GB$3.20$0.591$2,336.00
NV36adms_A10_v5368801× A10 / 24 GB$4.52$0.835$3,299.60
NV72ads_A10_v5728802× A10 / 48 GB~$6.40~$1.18~$4,672

AMD / other: NGads V620 v1 (AMD Radeon PRO V620; sizes ng8/ng16/ng32ads) exists for VDI/graphics but list pricing was not publicly retrievable — verify in the portal; not recommended for LLM training. MI300X (ND96isr_MI300X_v5) is available in 17 regions with 192 GB/GPU (1.5 TB/node) — attractive for 70B+ single-node inference — but ROCm software maturity vs CUDA is a real consideration for your Axolotl/Unsloth/vLLM stack; no reserved rate is published, so every hour is billed on-demand.

1.2 Reserved Instance & Savings Plan rates

Aggregators do not publish per-SKU GPU reserved dollar rates. Use these defensible rules and validate in-portal:

CommitmentDiscount vs PAYGNotes
Spot57–82% off (see table)<30 s eviction notice; training/burst only
1-Year Reserved~35% off (GPU, confirmed); ~31–41% (CPU)Thunder Compute: “one-year reserved… reduce on-demand rates by about 35%”
3-Year Reserved~55% off (assumption for GPU — unverified); ~53% (CPU)Label as estimate; no Microsoft source for 3-yr GPU
Savings Plan for Compute (1-yr)~31% realistic (up to 65% headline)Flexible across VM families
Savings Plan for Compute (3-yr)~53% realistic (Microsoft headline “up to 65%”)Best for mixed/always-on fleets

Worked RI estimates (East US, applying ~35%/~55%): NC24ads_A100_v4 → ~$2.39/hr (1yr, ~$1,743/mo) / ~$1.65/hr (3yr, ~$1,206/mo). NC40ads_H100_v5 → ~$4.54/hr (1yr) / ~$3.14/hr (3yr). NC8as_T4_v3 → ~$0.49/hr (1yr) / ~$0.34/hr (3yr). ND96isr_H100_v5 → ~$63.40/hr (1yr; secondary-source estimate from Spheron, based on Azure list rates 21 May 2026). H200, MI300X, and GB200 have no published reserved rate — full on-demand only.

1.3 CPU VMs (control plane, data prep, orchestration)

SKUvCPURAM GiBPAYG $/hrSpot $/hrPAYG $/moRole
D2s_v528~$0.096~$0.02~$70small system node
D4s_v5416~$0.192~$0.04~$140system node
D8s_v5832$0.384$0.081$280.32system/user node
D16s_v51664~$0.768~$0.16~$561system/orchestration
D32s_v532128~$1.536~$0.32~$1,121heavy orchestration
D8as_v5 (AMD)832$0.344$0.073$251cheaper GP node
E4s_v5432~$0.252~$0.05~$184data prep (mem)
E8s_v5864~$0.504~$0.10~$368tokenization/dataprep
E16s_v516128~$1.008~$0.21~$736large dataprep
E32s_v532256~$2.016~$0.42~$1,472large dataprep
E64s_v564512~$4.032~$0.84~$2,943very large dataprep
F8s_v2816~$0.338~$0.07~$247compute-opt
F16s_v21632~$0.676~$0.14~$493compute-opt
F32s_v23264~$1.352~$0.28~$987compute-opt
B2s (burstable)24~$0.042n/a~$30cheap always-on utility
Dpsv6 / Epsv6 (ARM, Cobalt 100)varies~10–15% cheaper than Dsv5/Esv5yescheap ARM utility/CI nodes

(CPU figures marked “~” are derived from the confirmed D8s_v5 anchor of $0.384/hr and standard family scaling; validate exact E/F/ARM rates in the Pricing Calculator. ARM Cobalt 100 is materially cheaper for stateless utility/CI workloads and worth adopting for non-GPU nodes.)

1.4 Quota, scarcity & deprecation flags

  • All N-series (GPU) SKUs are quota-gated. New subscriptions start at zero N-series vCPU quota; you must file per-region, per-family vCPU quota requests via the portal. Approval ranges from a few hours to several days depending on demand and account history.
  • Scarcity: H100/H200/GB200 capacity is constrained; ND96isr_H100_v5 is most consistently available in East US, South Central US, and West Europe, and scarce in Southeast Asia and UK South. ND A100 v4 remains widely available (27 regions). ND96asr_v4 (8× A100 40GB) is available in only ~5 regions (East US, West US 2, South Central US, Italy North, West Europe).
  • Deprecation posture: Direct new training clusters to H100/H200; keep A100 for inference and cost-sensitive fine-tuning. Do not build new long-lived commitments on the A100 SXM (NDv4) generation.

1.5 Regional price comparison — 8× H100 (ND96isr_H100_v5) as anchor

Region8× H100 $/hrVariance vs East US
East US~$88.49baseline
East US 2~$88.490%
South Central US~$88.490%
West US 3~$94.68+7%
North Europe~$105.30+19%
West Europe~$105.30+19%
Sweden Central~$100–105+13–19%
Southeast Asia~$115.04+30%

(Some trackers list ND96isr_H100_v5 East US at $98.32/hr; the $88.49 figure is from a March 2026 tracker. The relative regional spread of 20–33% is the reliable planning takeaway — put non-latency-sensitive training in the cheapest US region.)


PART 2 — AKS CLUSTER COSTS

2.1 Control plane tiers

Tier$/cluster/hr$/moSLAUse
Free$0$0No financially-backed SLA (best-effort ~99.5%)dev/test, <10 nodes
Standard$0.10~$7399.95% (with Availability Zones) / 99.9%production default; scales to 5,000 nodes
Premium$0.60~$43899.95% + Long-Term Support (2 yr per version)regulated/long-lived

2.2 Itemized billable components (everything, not just VMs)

  • System node pool: min 1 node; production recommendation 3 nodes across AZs of D8s_v5-class for HA of system pods.
  • Managed OS/data disks (Premium SSD v1): P10 (128 GiB) $19.71/mo, P15 (256 GiB) ~$38/mo, P20 (512 GiB) $73.22/mo, P30 (1 TiB) ~$135/mo. Snapshots ~$0.05/GB/mo.
  • Shared storage for datasets/checkpoints:
    • Azure Files Premium (provisioned SSD): $0.16/GiB/mo (v1); v2 ~$0.10/GiB/mo + provisioned IOPS/throughput charges.
    • Azure NetApp Files: Standard $0.1475/GiB/mo, Premium $0.2942/GiB/mo, Ultra $0.3927/GiB/mo (per Microsoft Learn ANF cost model; region-dependent). Reserved capacity ~34% off.
    • Azure Managed Lustre: priced per TiB provisioned (Durable-Premium-40/125/250/500 MB/s-per-TiB tiers), billed on provisioned capacity regardless of use, minimum cluster size applies; per-GB list not publicly retrievable — get from the Pricing Calculator. Only justified when multi-node training I/O is the bottleneck.
    • Guidance: Blob + BlobFuse2 for datasets/model weights (cheapest, use for HF cache); Azure Files Premium for shared checkpoints on small/medium clusters; NetApp Files / Managed Lustre only for multi-node distributed-training I/O at scale.
  • Blob Storage: Hot $0.018/GB/mo, Cool $0.01, Cold $0.0036, Archive $0.00099; transactions per 10k: Hot write ~$0.065 / read ~$0.005; Cool write ~$0.13; Archive read ~$5.50/10k + $0.022/GB retrieval (don’t put hot model weights in Archive).
  • Networking: Standard Load Balancer ~$0.025/hr + data-processed; NAT Gateway ~$0.045/hr + $0.045/GB; Application Gateway/AGIC (WAF_v2) ~$0.36/hr + capacity units; Private Endpoints ~$0.01/hr each; VNet peering ~$0.01/GB each direction; internet egress $0.087/GB after 100 GB free (Zone 1: US/EU/UK); Zone 2 (Asia-Pacific) $0.12/GB, Zone 3 (S. America/Africa/ME) $0.181/GB. Cross-Availability-Zone traffic within a region is free.
  • ACR: Basic $0.167/day ($5/mo, 10 GiB), Standard $0.667/day ($20/mo, 100 GiB), Premium $1.667/day ($50/mo, 500 GiB, geo-replication + Private Link); overage $0.10/GB/mo; geo-replication ~$50/mo per extra region. vLLM/Axolotl images are 10–25 GB, so Standard is the practical floor and Premium is realistic once you have several images plus retention.
  • Log Analytics / Container Insights: $2.30/GB Analytics Logs (after 5 GB free/month per billing account), Basic Logs $0.50/GB, Auxiliary $0.05/GB; retention free 31 days then ~$0.10/GB/mo (interactive) or ~$0.02/GB/mo (long-term). Commitment tier from 100 GB/day saves ~30%. Realistic GPU cluster: 5–30 GB/day → $350–$2,000/mo if unmanaged. This is the most common budget surprise.
  • Key Vault: ~$0.03/10k operations; Managed Identity: free; Defender for Containers: ~$7/vCPU/mo (billed per core — meaningful on large GPU nodes with 40–96 vCPU).
  • Azure Backup for AKS: backup instance ~$5/instance/mo + snapshot storage.

2.3 Three reference AKS clusters (itemized monthly, PAYG)

SMALL / DEV — Free control plane, scale-to-zero GPU:

Line itemConfig$/mo
Control planeFree$0
System nodes2× D4s_v5~$280
GPU node1× NC8as_T4_v3 (scale-to-zero, ~8h/day ≈ 240h)~$180
OS disks3× P10~$59
Blob (datasets)200 GB Hot~$4
ACRStandard~$20
Monitoring~3 GB/day~$130
Networking/LBminimal~$30
Total~$703/mo

MEDIUM / PRODUCTION — Standard control plane, autoscaling, private networking:

Line itemConfig$/mo (PAYG)
Control planeStandard$73
System nodes3× D8s_v5~$841
GPU inference1× NC24ads_A100_v4 always-on$2,681
GPU training (burst)1× NC24ads_A100_v4 Spot, ~150h/mo~$102
OS + data disks4× P15 + 1 TB P30~$287
Shared storage1 TB Azure Files Premium~$164
Blob2 TB Hot~$37
ACRPremium~$50
Load Balancer + NAT + Private Endpoints~$150
Egress~2 TB~$174
Monitoring~10 GB/day~$690
Key Vault + Defender~$120
Total~$5,369/mo

LARGE / SCALE — multi-region, full observability:

Line itemConfig$/mo (PAYG)
Control planeStandard ×2 regions~$146
System nodes3× D16s_v5 ×2 regions~$3,366
GPU inference pool4× NC40ads_H100_v5 always-on~$20,382
GPU training2× ND96isr_H100_v5 Spot, ~200h/mo~$7,268
Storage (disks + 10 TB NetApp Premium)~$3,500
Blob20 TB Hot~$369
ACRPremium + 1 geo-replica~$100
Networking + egress~10 TB egress~$1,200
Monitoring~40 GB/day~$2,760
Security/Backup~$500
Total~$40,000–$45,000/mo

2.4 Same three clusters — PAYG vs Spot vs 1-Yr Reserved (compute portion)

ClusterPAYG/moSpot (training + burst)1-Yr Reserved (baseline)
Small/Dev~$700~$550~$620
Medium/Prod~$5,370~$3,600~$3,900 (RI on inference node)
Large/Scale~$42,000~$26,000~$28,000–$30,000

PART 3 — ALTERNATIVES TO AKS

OptionBillingBest forWatch-outs
AKS + GPU node pools (baseline)VM-hours + platformFull control, multi-adapter vLLM, always-on + burstYou operate everything
Azure Container Apps serverless GPU (GA)Per-second, scale-to-zero, T4 & A100Spiky/low-duty inferenceInitially West US 3 / Australia East / Sweden Central (later +11 regions); cold starts; per-second GPU rate not surfaced in the calculator
Azure ML compute clusters + managed online endpointsVM-hours (no AzureML surcharge on the VM — you pay list VM price) + low-priority/Spot supportManaged training jobs, MLOpsLess control than AKS; online endpoint is always-on billed by VM
Azure BatchVM-hours + low-priority VMs (deep discount)Fire-and-forget training jobsNot for serving
IaaS VMs + CycleCloud / SkyPilotVM-hoursHPC-style multi-node training, multi-cloud burstYou build orchestration
Azure AI Foundry serverless (MaaS)Per-tokenBUY option; pilots, low/variable volumeNo infra control; per-token adds up at scale
Azure OpenAI PTU$1.00/PTU/hr; ~$260/PTU/mo reserved; ~$2,652/PTU/yr; min 15 PTU (Global)Predictable high-volume OpenAI models (GPT-4o/5)Not applicable to your open-source domain adapters

Foundry serverless per-1M-token prices: GPT-4o $2.50 in / $10.00 out (wrvishnu.com, Azure AI Foundry Pricing 2026); Phi-4-mini $0.07 in / $0.23 out (~35–40× cheaper than GPT-4o); DeepSeek R1 $1.35 / $5.40 and DeepSeek V3 $1.14 / $4.56 (Future AGI, verified 2 Jun 2026). Llama 3.3 70B: ~$0.59 in / $0.79 out per wrvishnu.com (May 2026 Global Standard) — note a pricing discrepancy: LiteLLM/Future AGI lists Llama 3.3 70B at ~$0.71/$0.71 (verified 2 Jun 2026). Use ~$0.6–0.8/1M as the planning band and confirm live.

Break-even (self-host vs serverless): An always-on NC24ads_A100_v4 (1× A100 80GB) costs $3.67/hr PAYG ($2,681/mo) or ~$1,206/mo on a 3-yr reservation. Serving an 8B model, one A100 can sustain roughly 1,000–5,000 output tok/s at good batch efficiency. Even at a conservative 1,000 tok/s sustained, that is 2.6B output tokens/month. At Llama-70B-class serverless output pricing ($0.79/1M) the equivalent serverless output spend would be ~$2,000+/month — so self-hosting on a reserved A100 becomes cheaper once you sustain roughly 30–55M tokens/day (≈1–1.6B tokens/month) at >40–50% GPU utilization. Below that, serverless wins on both cost and operational simplicity; above it, reserved AKS wins and the advantage grows with volume. Validate with a load test on your actual adapter — throughput depends heavily on sequence length, batch size, and quantization.


PART 4 — HOW TO BUILD A PLATFORM THAT SCALES GPU CAPACITY

  • Cluster Autoscaler (CAS) vs Node Autoprovisioning (NAP / Karpenter for Azure): CAS is fully supported and reliable — it scales predefined node pools on CPU/memory pressure. NAP (built on Karpenter) went GA in AKS mid-July 2025 (docs refreshed Sep/Oct 2025; public preview since Dec 2023). NAP picks the optimal VM SKU per pending pod via selectors (karpenter.azure.com/sku-gpu-name, sku-gpu-count, sku-gpu-manufacturer), consolidates underused nodes, and falls back across SKUs/zones during capacity shortages. Use NAP for GPU pools because it survives regional H100/H200 scarcity far better than fixed pools. Limits: no Windows/IPv6, you can’t stop a NAP-enabled cluster, requires Azure CNI Overlay/Cilium, and you still manage the GPU device plugin.
  • KEDA for event-driven scaling of inference — scale vLLM replicas on queue depth or custom vLLM metrics (pending-request count, batch fullness).
  • Scale-to-zero GPU pools (min-count 0): realistic cold start = node provision (3–5 min on Azure, slower than EC2) + image pull (10–25 GB vLLM image) + model load (tens of seconds to minutes). Mitigate with pre-pulled node images, ACR artifact caching, and a persistent HF-cache volume.
  • Spot node pools: set max price, eviction policy, and handle the <30 s eviction notice; checkpoint training every N steps to Blob/Azure Files so evictions are cheap. Spot discounts run 57–82% on GPU SKUs. Do not put always-on inference on Spot.
  • GPU sharing / partitioning: NVIDIA MIG on A100/H100 (partition one GPU into up to 7 instances), and time-slicing, via the NVIDIA GPU Operator (more control, DCGM metrics) or the simpler AKS GPU device plugin.
  • Multi-LoRA serving (directly relevant to your multi-species/feed-formulation plan): vLLM serves many LoRA adapters on one base model on a single GPU, loaded dynamically and routed per request. This is the single highest-leverage cost move — one A100/H100 can serve dozens of domain adapters instead of dedicating a node per adapter, driving marginal cost-per-adapter to near zero.
  • Quota management: per-region, per-family vCPU quotas. Request increases early, design multi-region fallback for scarce H100/H200, and use Capacity Reservations to guarantee scarce GPU availability for production.
  • Commitment blend (standard pattern): RI or Savings Plan for the always-on inference baseline; Spot for training and burst; PAYG only for unpredictable overflow.
  • Caching to cut cold starts: ACR artifact cache + Azure Container Storage + pre-pulled images + persistent PV for the HuggingFace cache.
  • Observability & cost attribution: namespace/label-based allocation via OpenCost/Kubecost on AKS + Azure Cost Management; DCGM exporter for GPU utilization. GPU utilization is the #1 cost lever — track it obsessively.
  • FinOps: idle-GPU detection, right-sizing, and scheduled shutdown of dev clusters (nights/weekends can roughly halve dev cost).

PART 5 — LEADERSHIP PIVOT TABLES

Pivot 1 — Scenario × Cost Component (monthly / annual, PAYG)

ScenarioTrainingInferenceStorageNetworkingPlatform/OpsMonthlyAnnual
Pilot~$100 (Spot T4/A100)~$180 (scale-to-zero)~$10~$30~$150~$700~$8,400
Departmental Prod~$102 (Spot A100)~$2,681 (1× A100)~$365~$324~$900~$5,370~$64,000
Enterprise Scale~$7,268 (Spot ND H100)~$20,382 (4× H100)~$3,900~$1,200~$3,400~$42,000~$500,000

Pivot 2 — Scenario × Commitment Model (compute, monthly)

ScenarioPAYGSpot1-Yr RI3-Yr RIMax saving
Pilot$700$550 (−21%)$620 (−11%)$560 (−20%)~21%
Departmental$5,370$3,600 (−33%)$3,900 (−27%)$3,100 (−42%)~42%
Enterprise$42,000$26,000 (−38%)$30,000 (−29%)$23,000 (−45%)~45%

Pivot 3 — Workload Type × Cost Driver × Lever

WorkloadCost driverOptimization lever
Training (periodic)GPU-hours × run frequencySpot + checkpointing; NAP; run in cheapest US region
Inference (always-on)GPU idle timeRI/Savings Plan + vLLM multi-LoRA + KEDA autoscale
Data prepE-series CPU-hoursSpot E-series; ARM (Cobalt 100) nodes
Dev/TestIdle clustersScheduled shutdown; scale-to-zero; Free control plane

Pivot 4 — Model Size × GPU need × monthly serve cost

ModelFine-tune (LoRA/QLoRA)Full fine-tuneServe (vLLM)Serve $/mo (PAYG / 3-yr RI)
3B1× T4 16GB / A101× A100 40–80GB1× T4$549 / ~$247
8B1× T4/A10 (QLoRA) or 1× A100 80GB2× A100 80GB1× A100 80GB or A10$2,681 / ~$1,206
32B1× A100 80GB (QLoRA)4–8× A100/H1001–2× A100 80GB or 1× H100$2,681–$5,362 / ~$1,206–$2,400
70B2× A100 80GB (QLoRA)8× A100/H100 (ND)2× A100 80GB or 2× H100$5,362–$10,191 / ~$2,400–$4,600

(Your stated target — an 8B domain adapter with LoRA/QLoRA/DPO/GRPO — fits comfortably on a single A100 80GB for both training and serving, which is why the departmental plan centers on one such node.)

Pivot 5 — Build vs Buy (8B/70B-class, monthly)

Sustained volumeFoundry serverless (per-token)Self-host AKS (1× A100, 3-yr RI)Azure OpenAI PTU (15 PTU min)Verdict
Low (~5M tok/day)~$300–600~$1,206 (+ ops)~$3,900Buy (serverless)
Medium (~40M tok/day)~$2,000–4,000~$1,206 (+ ops)~$3,900≈ Break-even → Build
High (~200M tok/day)~$10,000–20,000~$1,206–2,681 (+ scale)~$3,900–7,800Build (self-host)

Pivot 6 — Cost per business unit (the metrics leadership cares about)

MetricApprox value
Cost per 1M tokens (self-host A100, 3-yr RI, ~50% util)~$0.20–0.50
Cost per 1M tokens (Foundry serverless, ~70B)~$0.59–0.79
Cost per query (~500 tok avg)~$0.0001–0.0004
Cost per fine-tuning run (8B LoRA, ~6h on 1× A100 Spot)~$4–22
Cost per adapter per month (vLLM multi-LoRA, shared GPU)near-zero marginal (amortized on shared node)

Pivot 7 — 3-Year TCO (Departmental Production example)

ComponentYear 1Year 2Year 33-yr total
Infrastructure (3-yr RI blend)~$47,000~$47,000~$47,000~$141,000
Engineering FTE (2 eng @ ~$225k loaded)$450,000$450,000$450,000~$1,350,000
Total~$497,000~$497,000~$497,000~$1.49M

FTE cost dwarfs infrastructure — the platform decision should optimize engineer time, not just VM price. Prefer managed options (NAP, Container Apps serverless GPU, Azure ML) where they save engineer-hours.

Pivot 8 — Sensitivity / What-if (Departmental A100 inference node)

ChangeEffect on monthly cost
Utilization 20% vs 50% vs 80%VM cost is fixed; cost-per-token is 4× worse at 20% than 80% util
Queries double (within node headroom)$0 extra; beyond headroom, +1 node ($1,206–2,681)
Move region East US → West/North Europe+~19% on GPU
Move region East US → Southeast Asia+~30% on GPU
Spot vs on-demand (training)−57% to −82%
PAYG → 3-yr RI (inference)−~55%
EA/CSP negotiated discount−10% to −30% on top of everything

Key Takeaways for Leadership

  1. GPU utilization is the single biggest cost lever. An idle A100 costs the same ~$2,681/month as a busy one. Every optimization (multi-LoRA, autoscaling, right-sizing) exists to push utilization up.
  2. Buy before you build. Below ~30–55M tokens/day, Azure AI Foundry serverless per-token APIs are cheaper and far simpler than running your own always-on GPU. Only self-host once sustained volume clears that bar.
  3. One A100 80GB node serves your entire multi-adapter plan. vLLM multi-LoRA lets a single reserved A100 (~$1,206/mo on a 3-yr reservation) host dozens of species/feed-formulation adapters — cost-per-adapter approaches zero.
  4. Blend commitments: reserve the baseline, Spot the bursts. 3-year reservations cut ~55% off always-on inference; Spot cuts 57–82% off training. Combined, they roughly halve a naïve PAYG bill.
  5. Watch monitoring and egress — the silent budget killers. Log Analytics at $2.30/GB and egress at $0.087/GB can quietly add thousands per month; cap log ingestion and keep traffic in-region.
  6. Engineering headcount, not hardware, is the biggest 3-year number (~$1.35M FTE vs ~$141k infra in the departmental case). Optimize for engineer productivity with managed AKS features (NAP is GA), not just cheap VMs.
  7. You have negotiating leverage. As a large enterprise, apply your EA/CSP discount (10–30% off list) before committing, and get exact GPU reserved-instance quotes from the Azure Portal — public trackers do not publish them.

Caveats

  • GPU reserved-instance dollar rates are not publicly published (Vantage/CloudPrice show “N/A”). The ~35% 1-year rule is confirmed; the ~55% 3-year figure is an industry estimate, not Microsoft-confirmed for GPU SKUs. H200, MI300X, and GB200 have no Spot or RI discount today — full on-demand only.
  • Several CPU (E/F/ARM) and A10 sub-SKU prices are derived/estimated from anchor SKUs and marked “~”; validate before budgeting.
  • Azure Managed Lustre and Container Apps serverless-GPU per-unit prices could not be extracted from public pages — obtain them from the Azure Pricing Calculator/portal.
  • ND96isr_H100_v5 East US is quoted as both ~$88.49/hr and $98.32/hr by different trackers; the ~20–33% regional spread is the reliable planning figure.
  • Llama 3.3 70B serverless pricing shows a source discrepancy ($0.59/$0.79 vs $0.71/$0.71) — treat ~$0.6–0.8/1M as a band and confirm live.
  • Token-throughput and break-even figures are engineering estimates dependent on model, sequence length, batch size, and quantization — validate with a load test on your actual adapter.
  • All prices exclude your negotiated EA/CSP discount and Azure Hybrid Benefit, and must be re-validated in the Azure Pricing Calculator before budget commitment.

Read the companion reference: Post-Training Techniques and Training Platforms — LoRA, QLoRA, FFT, DPO, and GRPO across Axolotl, Oumi, LLaMA-Factory, Unsloth, TRL, torchtune, NeMo, LLM Foundry, Ludwig, and PEFT.


Part of the Univrs research ecosystem: