

AWS Bedrock
AWS Bedrock pricing charges per token by default, then offers three other ways to buy the same model.
Updated on:
AWS Bedrock pricing: four tiers, two provisioned modes, one undocumented unit
AWS Bedrock pricing charges per token by default, then offers three other ways to buy the same model. A service_tier parameter per request sets both the rate and the queue position. Reserved capacity sits outside that, billed monthly against tokens-per-minute you commit to in advance. The older Provisioned Throughput mode still exists and bills hourly per Model Unit.
Key takeaways
Bedrock exposes four service tiers on one API call: Reserved, Priority, Standard and Flex, selected with a service_tier parameter set to reserved, priority, default or flex.
Reserved capacity meters tokens-per-minute, not compute slices, and you size input and output TPM separately with minimums of 100,000 input TPM and 10,000 output TPM.
Batch inference runs at 50% below on-demand rates on selected models, and geographic and in-region routing costs 10% more than global cross-Region inference on current models.
AWS publishes hourly Provisioned Throughput rates for a handful of older Cohere and Meta models and tells everyone else to call their account team.
AWS Bedrock pricing in 2026
Mode | Billing unit | Commitment | Published rate
|
|---|---|---|---|
Standard | Per token, per model | None | Yes, per model per Region |
Priority | Per token at a premium | None | Yes, per model |
Flex | Per token at a discount | None | Yes, per model |
Batch | Per token, 50% off on-demand | None | Yes, per model |
Reserved | Per hour per 1K input and 1K output TPM, billed monthly | 1 or 3 months | Yes, per model per Region |
Provisioned | Per hour per Model Unit | None, 1 or 6 months | Partial. Cohere Command $49.50 / $39.60 / $23.77 |
Model | Input /1M | Output /1M | Batch in/out |
--- | --- | --- | --- |
Claude Opus 5 | $5.50 | $27.50 | $2.75 / $13.75 |
Claude Sonnet 5 | $2.20 | $11.00 | n/a |
Source: aws.amazon.com/bedrock/pricing, read 23 September 2026.
What AWS Bedrock pricing actually meters
Three different units, depending on how you buy.
On demand, the unit is the token, priced per million, per model, per Region and per tier. Priority costs more and jumps the queue. Flex costs less and accepts longer processing. Standard is the default when the parameter is absent. Your on-demand quota is shared across all three, so switching tier changes price and latency but not headroom.
Reserved meters throughput instead. You commit to a number of input tokens-per-minute and a number of output tokens-per-minute, priced per hour per 1,000 TPM and billed monthly. AWS sets floors of 100,000 input TPM and 10,000 output TPM, and warns that your consumption counts both InputTokenCount and CacheWriteInputTokens, so prompt caching inflates the reservation you need.
Provisioned Throughput meters a Model Unit. The docs define an MU as a level of input and output tokens per minute for one model, then decline to say what that level is, directing you to your AWS account manager. You cannot size a Provisioned Throughput purchase from public documentation. AWS now also sells it by Tokens, with no dated announcement.
The Reserved Tier block publishes rates per Region: $0.33 per hour per 1K input TPM and $1.65 per 1K output TPM on Claude Opus 4.6 at one month, $0.297 and $1.485 at three. Provisioned Throughput still reads "please reach out to your account team". Non-GA Claude Mythos access is gated separately and routes to an Anthropic account manager.
What happens when you hit the limit
Reserved capacity degrades rather than fails. When traffic exceeds what you reserved, Bedrock overflows to the Standard tier automatically and bills those tokens at on-demand rates, which keeps the application up and makes the bill variable at exactly the moment you bought a reservation to avoid variance. The tier targets 99.5% uptime for model response.
On-demand tiers have no spend ceiling. Throttling is governed by token quotas, not budget, so hitting a limit produces a throttling error rather than a charge stop. Provisioned Throughput inverts that: capacity is fixed and billing continues until you delete it. Reserved reservations only stop billing when your AWS account manager removes them, not from the console.
How AWS Bedrock pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
11 August 2026 | IAM principal cost allocation extended to the bedrock-mantle endpoint | Vendor |
31 December 2025 | Model support list for Priority and Flex updated | Vendor |
26 November 2025 | Reserved service tier added | Vendor |
18 November 2025 | Priority and Flex service tiers added for on-demand inference | Vendor |
16 July 2025 | Custom models deployable for on-demand per-token inference without Provisioned Throughput | Vendor |
21 August 2024 | Batch inference reaches general availability | Vendor |
29 March 2024 | Provisioned Throughput becomes purchasable for base models with no commitment | Vendor |
Customer
Sentiment Highlights
Llama 3.3 Instruct 70B is 5-10x cheaper than Sonnet 5
Bedrock user on model selection, Hacker News, July 2026
Explore other providers

Claude API
API
Claude API pricing charges per million tokens, and it splits them five ways.

Azure OpenAI
Enterprise LLM
Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything.

Canva
Workspace Platform
Canva pricing charges per person per month, includes AI in the seat, then caps it with a shared monthly allowance that resets on your billing date.
How much does AWS Bedrock cost?
What is a Model Unit in Bedrock?
Does Bedrock have a free tier?
Is cross-Region inference more expensive?
























