

Azure OpenAI
Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything.
Updated on:
Azure OpenAI pricing: paying for capacity instead of consumption
Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything. Alongside the usual per-token billing, Microsoft sells Provisioned Throughput Units that reserve dedicated capacity and bill hourly whether or not a request arrives. Choosing between the four deployment types is the actual pricing decision, and it turns on traffic shape rather than volume.
Key takeaways
The same model carries four billing modes: standard per-token, priority per-token, provisioned per PTU-hour, and discounted batch.
Provisioned deployments bill from the moment they are created until they are deleted, regardless of tokens consumed.
Hourly provisioned pricing runs $1 per PTU on Global, $1.10 on Data Zone and $2 on Regional, with one-month and one-year reservations cutting that substantially.
Reservations are locked to a deployment type, so Global, Data Zone and Regional each need their own.
Azure OpenAI pricing in 2026
Deployment type | Billing | Latency SLA | Built for
|
|---|---|---|---|
Standard | Per token | None | Development, testing, variable traffic |
Priority processing | Per token at priority rate | Defined target per model | Latency-sensitive work without a commitment |
Provisioned | Per PTU per hour, or via reservation | Defined target per model | Mission-critical, high-scale, predictable load |
Batch | Per token at a discounted rate | None | Bulk asynchronous processing |
What Azure actually meters
Three of the four deployment types meter tokens in the familiar way, separated only by rate: standard, a priority tier that costs more and carries a latency target, and batch that costs less and returns results asynchronously.
The fourth meters capacity. A Provisioned Throughput Unit is a slice of dedicated model processing throughput held exclusively for your deployment. Microsoft is explicit that a provisioned deployment holds that capacity whether or not requests are being made, and bills at an hourly rate per PTU deployed regardless of the number of tokens consumed. The meter starts when the deployment is created and stops when it is deleted.
Each model publishes its own PTU-to-tokens-per-minute ratio and its own minimum PTU count, so the smallest viable deployment differs by model. PTU quota is a separate concept from capacity: quota is a policy ceiling on how many PTUs you may deploy per subscription, per region and per deployment type, and it carries no cost of its own. Having quota does not guarantee capacity is available in the region.
How reservations work
Provisioned deployments bill two ways. Hourly is the flexible mode, useful for benchmarking a model or scaling up for a short event. Microsoft's own documentation advises against treating hourly as a scaling strategy, for two stated reasons: capacity may not be available when you try to scale back up, and sustained hourly billing at high utilisation typically costs more than a reservation.
Reservations commit to one month or one year and bill matching usage at a discounted rate instead of the hourly one. Three constraints shape how they are bought:
Reservations are not interchangeable across deployment types. Global, Data Zone and Regional each need their own, and a Global reservation will not cover Data Zone usage.
Global reservations are not region-specific. One Global reservation can cover deployments across several regions provided you reserved enough units.
Exchanges reset the term. You can swap region, deployment type, term or payment option, but the clock restarts.
Payment runs upfront or monthly.
What happens when you hit the limit
Nothing runs out, because nothing is allocated. Standard and batch deployments bill every token at list rate with no included allowance to exceed. Provisioned deployments have the opposite problem: the capacity is fixed, so exceeding it means requests queue or fail rather than costing more.
The real constraint is quota. If you need more PTUs than your subscription allows in a region, you request a quota increase, and separately confirm the region actually has capacity to give.
How Azure OpenAI pricing has changed
Date | Milestone | Source
|
|---|---|---|
18 Nov 2025 | Azure AI Foundry becomes Microsoft Foundry at Ignite. Docs move to /azure/foundry/ and rename PTUs to Foundry Provisioned Throughput | Vendor |
14 Aug 2024 | Provisioned Reservations added in one-month and one-year terms, quoted at up to 82% and 85% below the hourly rate | Vendor |
14 Aug 2024 | Self-service provisioned deployments introduced at a flat $2 per PTU-hour, replacing sales-negotiated capacity | Vendor |
Customer
Sentiment Highlights
The cheapest option is one of the GPT models, at a minimum of $10k/month
juliangoldsmith, on provisioned throughput units, Hacker News, September 2025
"it was just a farce to get you using provisioned throughput."
7thpower, Hacker News, April 2026
Explore other providers

AWS Bedrock
Enterprise LLM
AWS Bedrock pricing charges per token by default, then offers three other ways to buy the same model.

Claude
Chatbot
Claude pricing sells seats from $0 to $200 a month, and each buys an allowance Anthropic describes only as a multiple of the tier below.

Braintrust
Data Platform
Braintrust pricing charges nothing per span.
How much does Azure OpenAI cost?
What is a PTU in Azure OpenAI?
Do Azure OpenAI reservations cover every deployment type?
Is provisioned throughput cheaper than pay-as-you-go?
























