Azure GPU and CPU Cost Reference for Open-Source LLM Fine-Tuning and Serving
Single-VM and AKS cluster costs for fine-tuning and serving open-source LLMs on Azure — GPU/CPU pricing (PAYG, Spot, Reserved), itemized cluster builds, alternatives to AKS, and leadership pivot tables. A cost addendum to the fine-tuning platforms reference.
Leadership-ready reference — Region: East US (primary) · Currency: USD · Prices checked: 11 August 2026
Read this first. All list prices below are Azure public pay-as-you-go (PAYG) rates for Linux VMs, verified against Vantage (instances.vantage.sh), CloudPrice, SpareCores and Microsoft pricing pages as of 11 Aug 2026. Azure prices change frequently and must be re-validated in the Azure Pricing Calculator before any budget commitment. As a large agricultural enterprise, your organization almost certainly has an Enterprise Agreement (EA) or CSP discount of 10–30% off these list prices — treat every number here as a ceiling, not a final cost. Azure Hybrid Benefit further reduces cost where you own Windows/SQL licensing (not usually relevant to Linux GPU nodes).
The single biggest cost lever is GPU utilization, not SKU choice — an idle A100 (~$2,681/month) costs exactly the same as a fully-loaded one. For your 8B LoRA + vLLM multi-adapter plan, a 1× A100 80GB (NC24ads_A100_v4, ~$3.67/hr PAYG) or 1× H100 NVL (NC40ads_H100_v5, ~$6.98/hr) node covers both fine-tuning and serving; running one node always-on on a 3-year reservation plus Spot for training bursts is the cost-optimal pattern.
Build-vs-buy crosses over at roughly 30–55 million tokens/day of sustained inference: below that, Azure AI Foundry serverless per-token APIs (Llama-class models at sub-$1 per 1M tokens) are cheaper and simpler than a self-hosted always-on GPU; above it, a reserved AKS GPU node wins decisively and the gap widens with volume.
A departmental production platform lands around $5,000–$12,000/month (one always-on A100 inference node + spot training + storage + monitoring on AKS Standard); an enterprise-scale multi-region platform runs $40,000–$120,000+/month. Reserved instances and Spot cut 35–65% off compute, and engineering FTE cost (~$450k–$900k/yr for 2–4 engineers) usually exceeds the infrastructure bill — model it explicitly.
Key Findings
A100 vs H100/H200 deployment posture: The A100 generation (NC A100 v4, ND A100 v4) is still fully offered and is the cost-efficient workhorse for 8B–32B work, but Azure is directing new capacity toward H100/H200 and Blackwell. NDv4 (A100 SXM) is described as “still available but being phased toward end-of-life for new workloads.” H200 (ND96isr_H200_v5, $110.24/hr) and GB200 (ND GB200 v6, now GA) currently have no published Spot or Reserved discount — you pay full on-demand.
GPU reserved-instance dollar rates are not published by aggregators (Vantage and CloudPrice show “N/A” for the 1-Year and 3-Year Reserved fields on every GPU SKU). The defensible general rule is ~35% off for 1-year reservations (confirmed by Thunder Compute); a ~55% 3-year figure is an industry assumption, not Microsoft-confirmed for GPU SKUs. Get exact reservation quotes from the Azure Portal purchase flow.
The surprise line items are monitoring and egress, not compute add-ons. Container Insights/Log Analytics ingest at $2.30/GB Analytics Logs (after 5 GB free/month per billing account; Microsoft Azure Monitor pricing page and MonitoringCost.com, 2026), and a chatty GPU cluster can push tens of GB/day; internet egress is $0.087/GB after 100 GB free.
Serverless GPU (Azure Container Apps) is GA with scale-to-zero and per-second billing for T4 and A100. Microsoft’s “Announcing GA for Azure Container Apps Serverless GPUs” (Microsoft Community Hub, modified 27 Mar 2025) states GA is “supported in West US 3, Australia East, and Sweden Central” with default A100/T4 quota for EA customers; a 2025 recap notes GA later expanded to 11 additional regions. This is the best fit for spiky, low-duty-cycle inference.
Region choice swings GPU cost 20–33%. US regions (East US, East US 2, South Central US) are cheapest; Western Europe ~+19%, Southeast Asia ~+30%, Australia East ~+33%.
Details
PART 1 — SINGLE VM COSTS (GPU and CPU)
1.1 GPU VMs — specs and PAYG/Spot pricing (East US, Linux, per hour; monthly at 730 h)
SKU
vCPU
RAM GiB
GPU
VRAM total
Temp/NVMe
InfiniBand
PAYG $/hr
Spot $/hr (disc.)
PAYG $/mo
NC4as_T4_v3
4
28
1× T4 16GB
16 GB
180 GB
No
$0.526
$0.153 (~71%)
$383.98
NC8as_T4_v3
8
56
1× T4 16GB
16 GB
360 GB
No
$0.752
$0.257 (~66%)
$548.96
NC16as_T4_v3
16
110
1× T4 16GB
16 GB
360 GB
No
~$1.20
~$0.42
~$876
NC64as_T4_v3
64
440
4× T4 16GB
64 GB
2880 GB
No
$4.352
$1.859 (~57%)
$3,176.96
NC24ads_A100_v4
24
220
1× A100 80GB
80 GB
1026 GB
No
$3.673
$0.679 (~82%)
$2,681.29
NC48ads_A100_v4
48
440
2× A100 80GB
160 GB
2052 GB
No
$7.346
~$1.36
$5,362.58
NC96ads_A100_v4
96
880
4× A100 80GB
320 GB
4104 GB
No
$14.692
$2.715 (~82%)
$10,725.16
ND96asr_v4
96
900
8× A100 40GB SXM
320 GB
10.7 TB
200 Gb/s
$27.197
$5.983 (~78%)
$19,853.81
ND96amsr_A100_v4
96
1900
8× A100 80GB SXM
640 GB
10.7 TB
200 Gb/s
$32.77
$8.52 (~74%)
$23,922.10
NC40ads_H100_v5
40
320
1× H100 NVL 94GB
94 GB
3576 GB
No
$6.98
$1.403 (~80%)
$5,095.40
NC80adis_H100_v5
80
640
2× H100 NVL 94GB
188 GB
7152 GB
No
$13.96
$2.806 (~80%)
$10,190.80
ND96isr_H100_v5
96
1900
8× H100 80GB SXM5
640 GB
28 TB
3200 Gb/s
$98.32
$18.17 (~82%)
$71,773.60
ND96isr_H200_v5
96
1850
8× H200 141GB
1128 GB
31.8 TB
3200 Gb/s
$110.24
none published
$80,475.20
ND96isr_MI300X_v5 (AMD)
96
1850
8× MI300X 192GB
1536 GB
31.8 TB
RDMA
~$48 (region-dep., $8.87 low end)
$8.87
~$35,040
ND GB200 v6 (Blackwell)
128
900
4× GB200 Grace-Blackwell
—
—
NDR IB
Not published (GA)
none
—
A10 fractional GPUs (NVads_A10_v5 — VDI/visualization-oriented, usable for small inference):
SKU
vCPU
RAM GiB
A10 fraction / VRAM
PAYG $/hr
Spot $/hr
PAYG $/mo
NV6ads_A10_v5
6
55
1/6 A10 / 4 GB
~$0.45
$0.084
~$330
NV12ads_A10_v5
12
110
1/3 A10 / 8 GB
$0.908
$0.168
$662.84
NV18ads_A10_v5
18
220
1/2 A10 / 12 GB
~$1.36
~$0.25
~$993
NV36ads_A10_v5
36
440
1× A10 / 24 GB
$3.20
$0.591
$2,336.00
NV36adms_A10_v5
36
880
1× A10 / 24 GB
$4.52
$0.835
$3,299.60
NV72ads_A10_v5
72
880
2× A10 / 48 GB
~$6.40
~$1.18
~$4,672
AMD / other: NGads V620 v1 (AMD Radeon PRO V620; sizes ng8/ng16/ng32ads) exists for VDI/graphics but list pricing was not publicly retrievable — verify in the portal; not recommended for LLM training. MI300X (ND96isr_MI300X_v5) is available in 17 regions with 192 GB/GPU (1.5 TB/node) — attractive for 70B+ single-node inference — but ROCm software maturity vs CUDA is a real consideration for your Axolotl/Unsloth/vLLM stack; no reserved rate is published, so every hour is billed on-demand.
1.2 Reserved Instance & Savings Plan rates
Aggregators do not publish per-SKU GPU reserved dollar rates. Use these defensible rules and validate in-portal:
Commitment
Discount vs PAYG
Notes
Spot
57–82% off (see table)
<30 s eviction notice; training/burst only
1-Year Reserved
~35% off (GPU, confirmed); ~31–41% (CPU)
Thunder Compute: “one-year reserved… reduce on-demand rates by about 35%”
3-Year Reserved
~55% off (assumption for GPU — unverified); ~53% (CPU)
Label as estimate; no Microsoft source for 3-yr GPU
Savings Plan for Compute (1-yr)
~31% realistic (up to 65% headline)
Flexible across VM families
Savings Plan for Compute (3-yr)
~53% realistic (Microsoft headline “up to 65%”)
Best for mixed/always-on fleets
Worked RI estimates (East US, applying ~35%/~55%): NC24ads_A100_v4 → ~$2.39/hr (1yr, ~$1,743/mo) / ~$1.65/hr (3yr, ~$1,206/mo). NC40ads_H100_v5 → ~$4.54/hr (1yr) / ~$3.14/hr (3yr). NC8as_T4_v3 → ~$0.49/hr (1yr) / ~$0.34/hr (3yr). ND96isr_H100_v5 → ~$63.40/hr (1yr; secondary-source estimate from Spheron, based on Azure list rates 21 May 2026). H200, MI300X, and GB200 have no published reserved rate — full on-demand only.
1.3 CPU VMs (control plane, data prep, orchestration)
SKU
vCPU
RAM GiB
PAYG $/hr
Spot $/hr
PAYG $/mo
Role
D2s_v5
2
8
~$0.096
~$0.02
~$70
small system node
D4s_v5
4
16
~$0.192
~$0.04
~$140
system node
D8s_v5
8
32
$0.384
$0.081
$280.32
system/user node
D16s_v5
16
64
~$0.768
~$0.16
~$561
system/orchestration
D32s_v5
32
128
~$1.536
~$0.32
~$1,121
heavy orchestration
D8as_v5 (AMD)
8
32
$0.344
$0.073
$251
cheaper GP node
E4s_v5
4
32
~$0.252
~$0.05
~$184
data prep (mem)
E8s_v5
8
64
~$0.504
~$0.10
~$368
tokenization/dataprep
E16s_v5
16
128
~$1.008
~$0.21
~$736
large dataprep
E32s_v5
32
256
~$2.016
~$0.42
~$1,472
large dataprep
E64s_v5
64
512
~$4.032
~$0.84
~$2,943
very large dataprep
F8s_v2
8
16
~$0.338
~$0.07
~$247
compute-opt
F16s_v2
16
32
~$0.676
~$0.14
~$493
compute-opt
F32s_v2
32
64
~$1.352
~$0.28
~$987
compute-opt
B2s (burstable)
2
4
~$0.042
n/a
~$30
cheap always-on utility
Dpsv6 / Epsv6 (ARM, Cobalt 100)
varies
—
~10–15% cheaper than Dsv5/Esv5
yes
cheap ARM utility/CI nodes
(CPU figures marked “~” are derived from the confirmed D8s_v5 anchor of $0.384/hr and standard family scaling; validate exact E/F/ARM rates in the Pricing Calculator. ARM Cobalt 100 is materially cheaper for stateless utility/CI workloads and worth adopting for non-GPU nodes.)
1.4 Quota, scarcity & deprecation flags
All N-series (GPU) SKUs are quota-gated. New subscriptions start at zero N-series vCPU quota; you must file per-region, per-family vCPU quota requests via the portal. Approval ranges from a few hours to several days depending on demand and account history.
Scarcity: H100/H200/GB200 capacity is constrained; ND96isr_H100_v5 is most consistently available in East US, South Central US, and West Europe, and scarce in Southeast Asia and UK South. ND A100 v4 remains widely available (27 regions). ND96asr_v4 (8× A100 40GB) is available in only ~5 regions (East US, West US 2, South Central US, Italy North, West Europe).
Deprecation posture: Direct new training clusters to H100/H200; keep A100 for inference and cost-sensitive fine-tuning. Do not build new long-lived commitments on the A100 SXM (NDv4) generation.
1.5 Regional price comparison — 8× H100 (ND96isr_H100_v5) as anchor
Region
8× H100 $/hr
Variance vs East US
East US
~$88.49
baseline
East US 2
~$88.49
0%
South Central US
~$88.49
0%
West US 3
~$94.68
+7%
North Europe
~$105.30
+19%
West Europe
~$105.30
+19%
Sweden Central
~$100–105
+13–19%
Southeast Asia
~$115.04
+30%
(Some trackers list ND96isr_H100_v5 East US at $98.32/hr; the $88.49 figure is from a March 2026 tracker. The relative regional spread of 20–33% is the reliable planning takeaway — put non-latency-sensitive training in the cheapest US region.)
PART 2 — AKS CLUSTER COSTS
2.1 Control plane tiers
Tier
$/cluster/hr
$/mo
SLA
Use
Free
$0
$0
No financially-backed SLA (best-effort ~99.5%)
dev/test, <10 nodes
Standard
$0.10
~$73
99.95% (with Availability Zones) / 99.9%
production default; scales to 5,000 nodes
Premium
$0.60
~$438
99.95% + Long-Term Support (2 yr per version)
regulated/long-lived
2.2 Itemized billable components (everything, not just VMs)
System node pool: min 1 node; production recommendation 3 nodes across AZs of D8s_v5-class for HA of system pods.
Azure NetApp Files: Standard $0.1475/GiB/mo, Premium $0.2942/GiB/mo, Ultra $0.3927/GiB/mo (per Microsoft Learn ANF cost model; region-dependent). Reserved capacity ~34% off.
Azure Managed Lustre: priced per TiB provisioned (Durable-Premium-40/125/250/500 MB/s-per-TiB tiers), billed on provisioned capacity regardless of use, minimum cluster size applies; per-GB list not publicly retrievable — get from the Pricing Calculator. Only justified when multi-node training I/O is the bottleneck.
Guidance:Blob + BlobFuse2 for datasets/model weights (cheapest, use for HF cache); Azure Files Premium for shared checkpoints on small/medium clusters; NetApp Files / Managed Lustre only for multi-node distributed-training I/O at scale.
Blob Storage: Hot $0.018/GB/mo, Cool $0.01, Cold $0.0036, Archive $0.00099; transactions per 10k: Hot write ~$0.065 / read ~$0.005; Cool write ~$0.13; Archive read ~$5.50/10k + $0.022/GB retrieval (don’t put hot model weights in Archive).
Networking: Standard Load Balancer ~$0.025/hr + data-processed; NAT Gateway ~$0.045/hr + $0.045/GB; Application Gateway/AGIC (WAF_v2) ~$0.36/hr + capacity units; Private Endpoints ~$0.01/hr each; VNet peering ~$0.01/GB each direction; internet egress $0.087/GB after 100 GB free (Zone 1: US/EU/UK); Zone 2 (Asia-Pacific) $0.12/GB, Zone 3 (S. America/Africa/ME) $0.181/GB. Cross-Availability-Zone traffic within a region is free.
ACR: Basic $0.167/day ($5/mo, 10 GiB), Standard $0.667/day ($20/mo, 100 GiB), Premium $1.667/day ($50/mo, 500 GiB, geo-replication + Private Link); overage $0.10/GB/mo; geo-replication ~$50/mo per extra region. vLLM/Axolotl images are 10–25 GB, so Standard is the practical floor and Premium is realistic once you have several images plus retention.
Log Analytics / Container Insights:$2.30/GB Analytics Logs (after 5 GB free/month per billing account), Basic Logs $0.50/GB, Auxiliary $0.05/GB; retention free 31 days then ~$0.10/GB/mo (interactive) or ~$0.02/GB/mo (long-term). Commitment tier from 100 GB/day saves ~30%. Realistic GPU cluster: 5–30 GB/day → $350–$2,000/mo if unmanaged. This is the most common budget surprise.
Key Vault: ~$0.03/10k operations; Managed Identity: free; Defender for Containers: ~$7/vCPU/mo (billed per core — meaningful on large GPU nodes with 40–96 vCPU).
Azure Backup for AKS: backup instance ~$5/instance/mo + snapshot storage.
2.3 Three reference AKS clusters (itemized monthly, PAYG)
SMALL / DEV — Free control plane, scale-to-zero GPU:
Line item
Config
$/mo
Control plane
Free
$0
System nodes
2× D4s_v5
~$280
GPU node
1× NC8as_T4_v3 (scale-to-zero, ~8h/day ≈ 240h)
~$180
OS disks
3× P10
~$59
Blob (datasets)
200 GB Hot
~$4
ACR
Standard
~$20
Monitoring
~3 GB/day
~$130
Networking/LB
minimal
~$30
Total
~$703/mo
MEDIUM / PRODUCTION — Standard control plane, autoscaling, private networking:
Line item
Config
$/mo (PAYG)
Control plane
Standard
$73
System nodes
3× D8s_v5
~$841
GPU inference
1× NC24ads_A100_v4 always-on
$2,681
GPU training (burst)
1× NC24ads_A100_v4 Spot, ~150h/mo
~$102
OS + data disks
4× P15 + 1 TB P30
~$287
Shared storage
1 TB Azure Files Premium
~$164
Blob
2 TB Hot
~$37
ACR
Premium
~$50
Load Balancer + NAT + Private Endpoints
~$150
Egress
~2 TB
~$174
Monitoring
~10 GB/day
~$690
Key Vault + Defender
~$120
Total
~$5,369/mo
LARGE / SCALE — multi-region, full observability:
Line item
Config
$/mo (PAYG)
Control plane
Standard ×2 regions
~$146
System nodes
3× D16s_v5 ×2 regions
~$3,366
GPU inference pool
4× NC40ads_H100_v5 always-on
~$20,382
GPU training
2× ND96isr_H100_v5 Spot, ~200h/mo
~$7,268
Storage (disks + 10 TB NetApp Premium)
~$3,500
Blob
20 TB Hot
~$369
ACR
Premium + 1 geo-replica
~$100
Networking + egress
~10 TB egress
~$1,200
Monitoring
~40 GB/day
~$2,760
Security/Backup
~$500
Total
~$40,000–$45,000/mo
2.4 Same three clusters — PAYG vs Spot vs 1-Yr Reserved (compute portion)
Cluster
PAYG/mo
Spot (training + burst)
1-Yr Reserved (baseline)
Small/Dev
~$700
~$550
~$620
Medium/Prod
~$5,370
~$3,600
~$3,900 (RI on inference node)
Large/Scale
~$42,000
~$26,000
~$28,000–$30,000
PART 3 — ALTERNATIVES TO AKS
Option
Billing
Best for
Watch-outs
AKS + GPU node pools (baseline)
VM-hours + platform
Full control, multi-adapter vLLM, always-on + burst
You operate everything
Azure Container Apps serverless GPU (GA)
Per-second, scale-to-zero, T4 & A100
Spiky/low-duty inference
Initially West US 3 / Australia East / Sweden Central (later +11 regions); cold starts; per-second GPU rate not surfaced in the calculator
Azure ML compute clusters + managed online endpoints
VM-hours (no AzureML surcharge on the VM — you pay list VM price) + low-priority/Spot support
Managed training jobs, MLOps
Less control than AKS; online endpoint is always-on billed by VM
Azure Batch
VM-hours + low-priority VMs (deep discount)
Fire-and-forget training jobs
Not for serving
IaaS VMs + CycleCloud / SkyPilot
VM-hours
HPC-style multi-node training, multi-cloud burst
You build orchestration
Azure AI Foundry serverless (MaaS)
Per-token
BUY option; pilots, low/variable volume
No infra control; per-token adds up at scale
Azure OpenAI PTU
$1.00/PTU/hr; ~$260/PTU/mo reserved; ~$2,652/PTU/yr; min 15 PTU (Global)
Predictable high-volume OpenAI models (GPT-4o/5)
Not applicable to your open-source domain adapters
Foundry serverless per-1M-token prices: GPT-4o $2.50 in / $10.00 out (wrvishnu.com, Azure AI Foundry Pricing 2026); Phi-4-mini $0.07 in / $0.23 out (~35–40× cheaper than GPT-4o); DeepSeek R1 $1.35 / $5.40 and DeepSeek V3 $1.14 / $4.56 (Future AGI, verified 2 Jun 2026). Llama 3.3 70B: ~$0.59 in / $0.79 out per wrvishnu.com (May 2026 Global Standard) — note a pricing discrepancy: LiteLLM/Future AGI lists Llama 3.3 70B at ~$0.71/$0.71 (verified 2 Jun 2026). Use ~$0.6–0.8/1M as the planning band and confirm live.
Break-even (self-host vs serverless): An always-on NC24ads_A100_v4 (1× A100 80GB) costs $3.67/hr PAYG ($2,681/mo) or ~$1,206/mo on a 3-yr reservation. Serving an 8B model, one A100 can sustain roughly 1,000–5,000 output tok/s at good batch efficiency. Even at a conservative 1,000 tok/s sustained, that is 2.6B output tokens/month. At Llama-70B-class serverless output pricing ($0.79/1M) the equivalent serverless output spend would be ~$2,000+/month — so self-hosting on a reserved A100 becomes cheaper once you sustain roughly 30–55M tokens/day (≈1–1.6B tokens/month) at >40–50% GPU utilization. Below that, serverless wins on both cost and operational simplicity; above it, reserved AKS wins and the advantage grows with volume. Validate with a load test on your actual adapter — throughput depends heavily on sequence length, batch size, and quantization.
PART 4 — HOW TO BUILD A PLATFORM THAT SCALES GPU CAPACITY
Cluster Autoscaler (CAS) vs Node Autoprovisioning (NAP / Karpenter for Azure): CAS is fully supported and reliable — it scales predefined node pools on CPU/memory pressure. NAP (built on Karpenter) went GA in AKS mid-July 2025 (docs refreshed Sep/Oct 2025; public preview since Dec 2023). NAP picks the optimal VM SKU per pending pod via selectors (karpenter.azure.com/sku-gpu-name, sku-gpu-count, sku-gpu-manufacturer), consolidates underused nodes, and falls back across SKUs/zones during capacity shortages. Use NAP for GPU pools because it survives regional H100/H200 scarcity far better than fixed pools. Limits: no Windows/IPv6, you can’t stop a NAP-enabled cluster, requires Azure CNI Overlay/Cilium, and you still manage the GPU device plugin.
KEDA for event-driven scaling of inference — scale vLLM replicas on queue depth or custom vLLM metrics (pending-request count, batch fullness).
Scale-to-zero GPU pools (min-count 0): realistic cold start = node provision (3–5 min on Azure, slower than EC2) + image pull (10–25 GB vLLM image) + model load (tens of seconds to minutes). Mitigate with pre-pulled node images, ACR artifact caching, and a persistent HF-cache volume.
Spot node pools: set max price, eviction policy, and handle the <30 s eviction notice; checkpoint training every N steps to Blob/Azure Files so evictions are cheap. Spot discounts run 57–82% on GPU SKUs. Do not put always-on inference on Spot.
GPU sharing / partitioning: NVIDIA MIG on A100/H100 (partition one GPU into up to 7 instances), and time-slicing, via the NVIDIA GPU Operator (more control, DCGM metrics) or the simpler AKS GPU device plugin.
Multi-LoRA serving (directly relevant to your multi-species/feed-formulation plan): vLLM serves many LoRA adapters on one base model on a single GPU, loaded dynamically and routed per request. This is the single highest-leverage cost move — one A100/H100 can serve dozens of domain adapters instead of dedicating a node per adapter, driving marginal cost-per-adapter to near zero.
Quota management: per-region, per-family vCPU quotas. Request increases early, design multi-region fallback for scarce H100/H200, and use Capacity Reservations to guarantee scarce GPU availability for production.
Commitment blend (standard pattern):RI or Savings Plan for the always-on inference baseline; Spot for training and burst; PAYG only for unpredictable overflow.
Caching to cut cold starts: ACR artifact cache + Azure Container Storage + pre-pulled images + persistent PV for the HuggingFace cache.
Observability & cost attribution: namespace/label-based allocation via OpenCost/Kubecost on AKS + Azure Cost Management; DCGM exporter for GPU utilization. GPU utilization is the #1 cost lever — track it obsessively.
FinOps: idle-GPU detection, right-sizing, and scheduled shutdown of dev clusters (nights/weekends can roughly halve dev cost).
Pivot 2 — Scenario × Commitment Model (compute, monthly)
Scenario
PAYG
Spot
1-Yr RI
3-Yr RI
Max saving
Pilot
$700
$550 (−21%)
$620 (−11%)
$560 (−20%)
~21%
Departmental
$5,370
$3,600 (−33%)
$3,900 (−27%)
$3,100 (−42%)
~42%
Enterprise
$42,000
$26,000 (−38%)
$30,000 (−29%)
$23,000 (−45%)
~45%
Pivot 3 — Workload Type × Cost Driver × Lever
Workload
Cost driver
Optimization lever
Training (periodic)
GPU-hours × run frequency
Spot + checkpointing; NAP; run in cheapest US region
Inference (always-on)
GPU idle time
RI/Savings Plan + vLLM multi-LoRA + KEDA autoscale
Data prep
E-series CPU-hours
Spot E-series; ARM (Cobalt 100) nodes
Dev/Test
Idle clusters
Scheduled shutdown; scale-to-zero; Free control plane
Pivot 4 — Model Size × GPU need × monthly serve cost
Model
Fine-tune (LoRA/QLoRA)
Full fine-tune
Serve (vLLM)
Serve $/mo (PAYG / 3-yr RI)
3B
1× T4 16GB / A10
1× A100 40–80GB
1× T4
$549 / ~$247
8B
1× T4/A10 (QLoRA) or 1× A100 80GB
2× A100 80GB
1× A100 80GB or A10
$2,681 / ~$1,206
32B
1× A100 80GB (QLoRA)
4–8× A100/H100
1–2× A100 80GB or 1× H100
$2,681–$5,362 / ~$1,206–$2,400
70B
2× A100 80GB (QLoRA)
8× A100/H100 (ND)
2× A100 80GB or 2× H100
$5,362–$10,191 / ~$2,400–$4,600
(Your stated target — an 8B domain adapter with LoRA/QLoRA/DPO/GRPO — fits comfortably on a single A100 80GB for both training and serving, which is why the departmental plan centers on one such node.)
Pivot 5 — Build vs Buy (8B/70B-class, monthly)
Sustained volume
Foundry serverless (per-token)
Self-host AKS (1× A100, 3-yr RI)
Azure OpenAI PTU (15 PTU min)
Verdict
Low (~5M tok/day)
~$300–600
~$1,206 (+ ops)
~$3,900
Buy (serverless)
Medium (~40M tok/day)
~$2,000–4,000
~$1,206 (+ ops)
~$3,900
≈ Break-even → Build
High (~200M tok/day)
~$10,000–20,000
~$1,206–2,681 (+ scale)
~$3,900–7,800
Build (self-host)
Pivot 6 — Cost per business unit (the metrics leadership cares about)
Metric
Approx value
Cost per 1M tokens (self-host A100, 3-yr RI, ~50% util)
~$0.20–0.50
Cost per 1M tokens (Foundry serverless, ~70B)
~$0.59–0.79
Cost per query (~500 tok avg)
~$0.0001–0.0004
Cost per fine-tuning run (8B LoRA, ~6h on 1× A100 Spot)
~$4–22
Cost per adapter per month (vLLM multi-LoRA, shared GPU)
near-zero marginal (amortized on shared node)
Pivot 7 — 3-Year TCO (Departmental Production example)
Component
Year 1
Year 2
Year 3
3-yr total
Infrastructure (3-yr RI blend)
~$47,000
~$47,000
~$47,000
~$141,000
Engineering FTE (2 eng @ ~$225k loaded)
$450,000
$450,000
$450,000
~$1,350,000
Total
~$497,000
~$497,000
~$497,000
~$1.49M
FTE cost dwarfs infrastructure — the platform decision should optimize engineer time, not just VM price. Prefer managed options (NAP, Container Apps serverless GPU, Azure ML) where they save engineer-hours.
VM cost is fixed; cost-per-token is 4× worse at 20% than 80% util
Queries double (within node headroom)
$0 extra; beyond headroom, +1 node ($1,206–2,681)
Move region East US → West/North Europe
+~19% on GPU
Move region East US → Southeast Asia
+~30% on GPU
Spot vs on-demand (training)
−57% to −82%
PAYG → 3-yr RI (inference)
−~55%
EA/CSP negotiated discount
−10% to −30% on top of everything
Key Takeaways for Leadership
GPU utilization is the single biggest cost lever. An idle A100 costs the same ~$2,681/month as a busy one. Every optimization (multi-LoRA, autoscaling, right-sizing) exists to push utilization up.
Buy before you build. Below ~30–55M tokens/day, Azure AI Foundry serverless per-token APIs are cheaper and far simpler than running your own always-on GPU. Only self-host once sustained volume clears that bar.
One A100 80GB node serves your entire multi-adapter plan. vLLM multi-LoRA lets a single reserved A100 (~$1,206/mo on a 3-yr reservation) host dozens of species/feed-formulation adapters — cost-per-adapter approaches zero.
Blend commitments: reserve the baseline, Spot the bursts. 3-year reservations cut ~55% off always-on inference; Spot cuts 57–82% off training. Combined, they roughly halve a naïve PAYG bill.
Watch monitoring and egress — the silent budget killers. Log Analytics at $2.30/GB and egress at $0.087/GB can quietly add thousands per month; cap log ingestion and keep traffic in-region.
Engineering headcount, not hardware, is the biggest 3-year number (~$1.35M FTE vs ~$141k infra in the departmental case). Optimize for engineer productivity with managed AKS features (NAP is GA), not just cheap VMs.
You have negotiating leverage. As a large enterprise, apply your EA/CSP discount (10–30% off list) before committing, and get exact GPU reserved-instance quotes from the Azure Portal — public trackers do not publish them.
Caveats
GPU reserved-instance dollar rates are not publicly published (Vantage/CloudPrice show “N/A”). The ~35% 1-year rule is confirmed; the ~55% 3-year figure is an industry estimate, not Microsoft-confirmed for GPU SKUs. H200, MI300X, and GB200 have no Spot or RI discount today — full on-demand only.
Several CPU (E/F/ARM) and A10 sub-SKU prices are derived/estimated from anchor SKUs and marked “~”; validate before budgeting.
Azure Managed Lustre and Container Apps serverless-GPU per-unit prices could not be extracted from public pages — obtain them from the Azure Pricing Calculator/portal.
ND96isr_H100_v5 East US is quoted as both ~$88.49/hr and $98.32/hr by different trackers; the ~20–33% regional spread is the reliable planning figure.
Llama 3.3 70B serverless pricing shows a source discrepancy ($0.59/$0.79 vs $0.71/$0.71) — treat ~$0.6–0.8/1M as a band and confirm live.
Token-throughput and break-even figures are engineering estimates dependent on model, sequence length, batch size, and quantization — validate with a load test on your actual adapter.
All prices exclude your negotiated EA/CSP discount and Azure Hybrid Benefit, and must be re-validated in the Azure Pricing Calculator before budget commitment.
Read the companion reference:Post-Training Techniques and Training Platforms — LoRA, QLoRA, FFT, DPO, and GRPO across Axolotl, Oumi, LLaMA-Factory, Unsloth, TRL, torchtune, NeMo, LLM Foundry, Ludwig, and PEFT.