AWS Bedrock Pricing (2026): Models, Tiers & Real Examples

When building generative AI applications, AWS Bedrock is one of the most powerful services out there. It’s fully managed, so you can focus on your app while AWS handles the scaling behind the scenes. And it gives you models from Anthropic, OpenAI, Meta, Amazon, Mistral, and others through a single API.
But pricing? That’s where things get complicated. What used to be “pay per token” has grown into six inference tiers, model-by-model rates, caching discounts, regional routing, and a growing list of add-ons that show up as separate lines on your bill.
We first wrote this guide in 2024. Since then, AWS has added Priority, Flex, and Reserved tiers, prompt caching, global cross-Region pricing, and a new free tier. So we’ve rewritten it for 2026 with current prices and fresh examples.
In this guide, we’ll cover:
- Frequently asked questions
- What AWS Bedrock is
- How the six inference options are priced
- What the major models cost per million tokens
- A realistic monthly cost example (and how to cut it)
- How to estimate and track what you’re actually spending
- A few cost optimization strategies
What is AWS Bedrock?
First things first: AWS Bedrock is a fully-managed service designed to simplify the development and scaling of generative AI applications. Bedrock provides access to a range of high-performing foundational models (FMs) from top AI companies–all conveniently through a single API.
How Amazon Bedrock pricing works: the six inference options
Every Bedrock bill starts with one question: which inference option did the request use? AWS now offers six, and the same model can cost half as much or 75% more depending on which you pick.
| Option | Price vs Standard | Commitment | Good for |
|---|---|---|---|
| Standard | baseline | None | Most synchronous traffic; the default |
| Priority | +75% | None | Customer-facing requests where latency matters most |
| Flex | −50% | None | Work that can tolerate slower, variable response times |
| Batch | −50% | None; async via S3, up to 24h | Offline summarization, classification, evaluation |
| Reserved | fixed monthly | 1 or 3 months | Predictable, mission-critical traffic; 99.5% availability target |
| Provisioned Throughput | fixed hourly | None, 1 or 6 months | Custom models and workloads that need dedicated capacity |
Let’s walk through each one. All examples below use Claude Sonnet 4.6 in US East (N. Virginia) at August 2026 prices, unless otherwise noted.
Standard (on-demand)
Standard is super straightforward: you pay for the input and output tokens your model chews through (or per image, or per second of video or audio). No commitment, no strings attached. This is where nearly everyone starts and where most production traffic stays.
Say you’re building a customer support chatbot. Here are our assumptions:
- Traffic: 10,000 conversations per day
- Input: 1,500 tokens per conversation (system prompt plus chat history), at $3.00 per 1M input tokens
- Output: 400 tokens per conversation, at $15.00 per 1M output tokens
And the cost breakdown:
- Input cost: 10,000 × 1,500 = 15M tokens/day × $3.00 = $45/day
- Output cost: 10,000 × 400 = 4M tokens/day × $15.00 = $60/day
- Total daily cost: $105/day
- Total monthly cost: $105 × 30 days = about $3,150/month
Notice that output is the bigger line even though there are far fewer output tokens. That’s true on every model: output tokens cost three to six times more than input tokens, so the length of the answer, not the prompt, usually drives the bill.
Hold onto this $3,150 number. We’ll use it to show what each of the other options does to the same workload.
Priority and Flex
AWS added these two in November 2025, and they’re the easiest lever most teams haven’t pulled yet.
- Priority costs 75% more than Standard and puts your requests at the front of the queue. Use it for the one or two paths where a slow response costs you a customer.
- Flex costs 50% less than Standard and lets AWS process your request when capacity is available. Responses may be slower and more variable, which is fine for internal tools and background jobs.
Both share the same on-demand quota as Standard, and not every model supports them, so check the model’s page first.
Our chatbot on Flex: about $1,575/month, half the Standard price. If some of that traffic is an internal helpdesk rather than customers, moving just that slice to Flex is free money.
Batch
With batch inference, you submit a set of prompts as a single JSON Lines file in S3, get results back in anywhere from minutes to 24 hours, and pay 50% of the Standard price. Who doesn’t love saving money?
A few caveats: no tool calling, no multi-turn conversations, and not every model supports it. Refer to the AWS docs for the current model list.
Suppose you want to summarize yesterday’s 10,000 support conversations overnight, each one 2,000 tokens in and 200 tokens out:
- Input cost: 20M tokens × $1.50 (batch rate) = $30
- Output cost: 2M tokens × $7.50 (batch rate) = $15
- Total: $45 per nightly batch, about $1,350/month
If you were doing that same job in real time on Standard, it’d be twice that. Batch is the cheapest way to run a frontier model on anything that doesn’t need an answer right now.
Reserved Tier
Reserved launched in November 2025. You reserve input and output tokens-per-minute capacity separately, for one or three months, at a fixed monthly price. Traffic above your reservation overflows to Standard automatically, and AWS targets 99.5% availability for reserved traffic.
It started with Claude Sonnet 4.5 and the model list has been growing. Reserved makes sense once you have a few months of steady production traffic to size it against. Before that, you’re guessing, and you pay for the reservation whether you use it or not.
Provisioned Throughput
Provisioned Throughput is dedicated model capacity, measured in model units and billed hourly whether you use it or not. You can go no-commitment, 1-month, or 6-month; the longer the commitment, the better the rate.
Here’s the thing in 2026: Provisioned Throughput is mostly for custom and fine-tuned models, which can’t run on the on-demand tiers at all. AWS now quotes most Provisioned Throughput pricing per model, often through your account team, so treat it as a conversation rather than a line on the price list. If you’re considering it for a standard model, price Reserved first. It’s usually the better fit.
Prompt caching and cross-Region pricing
Two modifiers apply on top of whichever tier you pick.
Prompt caching stores a stable prefix, like your system prompt or a long reference document, so repeated requests pay a cache-read rate (typically around 10% of the input price) instead of the full input price. Writing to the cache costs a premium, often 1.25× the input price, so caching pays off once a prefix gets reused a few times within the cache window.
Back to our chatbot. If 1,200 of those 1,500 input tokens are a system prompt that never changes, caching drops input from $45/day to roughly $16/day. Total: about $2,280/month, down from $3,150, with no change to the model or the answers.
Cross-Region inference profiles come in two kinds. Geographic profiles (US, EU, and so on) route requests within a boundary at Standard prices and keep your data residency intact. Global profiles route anywhere and cost roughly 10% less on supported models. Either way, the source Region sets the price.
Other charges to expect
A few more things can show up on a Bedrock bill:
- Guardrails, Knowledge Bases, Data Automation, Flows, and AgentCore: Each is metered separately. They’re out of scope for this guide, but if you’re running agents, expect them to be a meaningful share of the bill. One user request can trigger several model calls plus retrieval, memory, and tool charges.
- Model customization: Fine-tuning and distillation are billed per training hour or training token, plus $1.95 per model per month for storage. The resulting model needs Provisioned Throughput or custom-model-unit pricing to run.
- Model evaluation: You pay for the inference of the model under test plus judge-model tokens; human evaluation is $0.21 per completed task.
What do Bedrock models cost per million tokens?
The tier decides the multiplier; the model decides the base price. Here are the models most teams actually choose between, at Standard tier, US pricing, as of August 2026. Batch and Flex are 50% off these rates where supported. Always check the AWS pricing page for your Region.
| Model | Input per 1M tokens | Output per 1M tokens | Typical use |
|---|---|---|---|
| Amazon Nova Micro | $0.035 | $0.14 | Classification, routing, cheap text tasks |
| Amazon Nova Lite | $0.06 | $0.24 | Multimodal at low cost |
| Amazon Nova Pro | $0.80 | $3.20 | Balanced general-purpose |
| Claude Haiku 4.5 | $1.00 | $5.00 | Fast, high-volume assistants |
| Claude Sonnet 4.6 | $3.00 | $15.00 | Production default for most teams |
| Claude Sonnet 5 | $2.00* | $10.00* | Newest balanced model |
| Claude Opus 4.8 | $5.00 | $25.00 | Deep reasoning, complex agents |
| OpenAI GPT-5.6 Terra | $2.75 | $16.50 | Balanced OpenAI option (in-Region only) |
| OpenAI gpt-oss-120b | $0.15 | $0.60 | Open-weight general-purpose |
| Meta Llama 4 Scout | $0.17 | $0.36 | Open-weight, cost-sensitive workloads |
*Claude Sonnet 5 global launch pricing through August 31, 2026; $3.00 / $15.00 after that.
Two things jump out. First, the gap between the cheapest and the most expensive model is over 100×. Second, our $3,150/month chatbot would cost about $300/month on Nova Pro, if Nova Pro passes your quality tests. Picking the right model for each task matters more than any discount AWS offers.
How to estimate and track what Bedrock actually costs you
The examples above are the template for estimating. Here’s how to do it for real, and then how to keep an eye on the number once traffic is flowing.
AWS Pricing Calculator
Start with the AWS Pricing Calculator for a rough estimate before you even have an account. You enter expected input and output tokens per month per model and it produces a number. Its Bedrock coverage has historically lagged for third-party models, so for Anthropic or OpenAI models, grab the per-million rates from the pricing table above and run the math yourself. It’s three multiplications.
AWS Cost Explorer
Once you’re running, Cost Explorer is the fastest way to see where the money goes. Filter Service to “Amazon Bedrock”, then group by:
- Usage type, which encodes the model and Region, or
- API operation, which separates input tokens, output tokens, cache reads, and batch.
Use monthly granularity for the trend and daily when something spikes. Cost Explorer also forecasts up to a year ahead, which gets useful once you have a few months of history.
Tags, inference profiles, and Budgets
Bedrock invocations can be attributed to an application inference profile and, since April 2026, to the IAM principal that made the call. That’s what makes team-level chargeback possible. Tag your inference profiles by team and environment, then set an AWS Budget with an alert at 80% of what you expect.
One more: model invocation logging is off by default. Turn it on if you want per-call token counts to reconcile against the bill.
If you’d rather not build the dashboards yourself, CloudForecast’s daily report breaks Bedrock spend out by model, tag, and account and emails it to the people who need to see it.
Cost Optimization Strategies for AWS Bedrock
Before we dive into cost-saving tips, let’s quickly highlight a few challenges that can make AWS Bedrock pricing tricky to manage:
- Dynamic usage patterns: Your workloads can fluctuate, causing unpredictable costs.
- Complexity of cost structures: There are several pricing models and different rates for different foundation models (FMs), which can be overwhelming.
- Visibility into costs: Without proper monitoring, it’s tough to see exactly where your money is going.
The good news is you can tackle all of these issues with a few simple strategies.
Monitoring Your Workloads
Want to optimize your Bedrock costs? Start by keeping an eye on your usage. AWS provides great tools like AWS CloudWatch, which lets you monitor real-time metrics such as token usage and model activity. Here are some essential features of CloudWatch to take advantage of:
- Custom dashboards: Track the key metrics that matter the most to your application. This can be input/output tokens, or model performance.
- Alarms: Set a few usage thresholds such that CloudWatch will notify you before things get too expensive.
Additionally, you can utilize AWS CloudTrail to log API calls. This tool gives you a handy audit trail that shows who is using what resources, and how. Stay on top of your workloads, and those surprise charges won’t have a chance to sneak up on you.
Optimizing Model Usage
Bedrock offers a wide selection of FMs. While it’s tempting to go for the “best” or the “cheapest”, the real key is finding the most cost-effective option for your specific use case.
Not all models are created equal. Some are total overkill for certain applications. For example, if you’re working on a simple text classification task, there’s no need to use a complex, expensive model. Instead, choose a less complex FM that’s cheaper and better suited to your needs. This saves you both processing and token costs, without sacrificing too much on the quality your application requires.
By matching your model choice to the complexity of your task, you can get the job done efficiently and avoid unnecessary expenses.
Using Flex, Reserved, and Provisioned Throughput
For most teams, the biggest easy win is moving anything that isn’t customer-facing to the Flex tier for a 50% discount. Internal tools, background enrichment, and evaluation runs rarely need Standard-tier latency.
For larger, steady workloads, the Reserved tier gives you dedicated tokens-per-minute capacity at a fixed monthly price, with overflow to Standard. Provisioned Throughput still exists, but in 2026 it’s mostly for custom and fine-tuned models.
Either way, the only way to know for sure whether a reservation is worth it is by monitoring your usage for a few months first. That gives you the data to size it right instead of guessing.
Caching Your Prompts
If your requests share a long, stable prefix, like a system prompt, a tool definition, or a reference document, turn on prompt caching. Cached reads cost around 10% of the normal input price. Put the stable content first and the user’s message last so the cache actually hits. Track cache reads versus writes in Cost Explorer; if writes dominate, your prefix isn’t stable enough to cache.
Batch Processing for Cost Efficiency
Again for larger datasets, setting up batch processing can be a smart way to reduce costs in AWS Bedrock. Rather than processing data on-demand and in real-time, batch processing allows you to submit large datasets as a single input file, which can be more cost-effective.
This is best seen in use cases like sentiment analysis or language translation, where you can process everything as a batch, rather than analyzing each text individually. By bundling multiple requests into one big batch, you can cut down on token usage and streamline your workflows.
AWS Tags and Cost Categories for Cost Visibility
To get a clearer picture of AWS Bedrock expenses, start by implementing resource tags and organizing them with AWS Cost Categories. Tagging allows you to track costs by project, team, or environment, ensuring accountability across your organization.
Leveraging AWS Cost Categories, you can then group these tagged resources into customized spending buckets, aligning costs with business objectives. Together, these strategies provide transparency and control, empowering you to make informed decisions and optimize your AWS Bedrock usage effectively.
We wrote up two helpful guides on tagging resources and also cost categories:
- AWS Cost Categories: A Better Way to Organize Costs
- AWS Tags Best Practices and AWS Tagging Strategies
Conclusion
AWS Bedrock is a powerhouse for building and scaling generative AI applications, but the pricing has more moving parts every year. The good news is that the big levers are simple: pick the cheapest model that passes your quality bar, cache your prompts, move anything that isn’t customer-facing to Flex or Batch, and only reserve capacity once you’ve measured a few months of real traffic. Do those four things and the same workload can cost a quarter of what it did on day one.
Need a little extra help? Check out tools like CloudForecast, which makes tracking your Bedrock (and overall AWS) spending a breeze. CloudForecast can help you proactively catch those sneaky cost spikes early to keep your operation running smoothly.
Frequently asked questions about AWS Bedrock pricing
How much does AWS Bedrock cost?
There’s no flat fee. You pay per token (or per image or second of media) for the model you use. Prices range from about $0.035 per million input tokens for Amazon Nova Micro to $25 per million output tokens for Claude Opus 4.8. A support chatbot handling 10,000 conversations a day on Claude Sonnet 4.6 costs roughly $3,150 a month at Standard rates, and about half that with caching and Flex.
Does Amazon Bedrock have a free tier?
Not a model-specific one, but new AWS accounts created since July 2025 get $100 in credits at sign-up and can earn up to $100 more by trying services including Bedrock. The credits expire after six months or when used up, and they apply to Bedrock usage. There’s no ongoing free monthly allowance of tokens.
How much does Claude cost on AWS Bedrock?
As of August 2026, at Standard tier: Claude Haiku 4.5 is $1 input / $5 output per million tokens, Claude Sonnet 4.6 is $3 / $15, Claude Sonnet 5 is $2 / $10 on global launch pricing through August 31, 2026 (then $3 / $15), and Claude Opus 4.8 is $5 / $25. Batch and Flex are half price where supported.
Is there an AWS Bedrock pricing calculator?
Yes. The AWS Pricing Calculator has a Bedrock section where you enter monthly input and output tokens per model. Coverage of third-party models is uneven, so for those, multiply your expected tokens by the per-million rates on the Bedrock pricing page. The examples in this guide show the math.
What is Bedrock provisioned throughput?
Provisioned Throughput is dedicated model capacity, sold in model units and billed by the hour with no-commitment, 1-month, or 6-month terms. You pay for the capacity whether you use it or not. It’s required for custom and fine-tuned models; for standard models, the newer Reserved tier usually fits better.
Is Bedrock more expensive than calling Anthropic or OpenAI directly?
For most models, Bedrock’s Standard per-token price matches the provider’s own API price. What you get for it is running inside your AWS account: IAM, VPC endpoints, CloudTrail, consolidated billing, and spend that counts toward an EDP commitment. Some in-Region endpoints carry a roughly 10% premium over global pricing.
What are Bedrock’s token limits?
Limits are quotas, not prices. Each model and Region has a tokens-per-minute and requests-per-minute quota on the on-demand tiers, visible in Service Quotas, and you can request increases. Reserved and Provisioned Throughput give you capacity outside those quotas. Some models also apply a burndown multiplier to quota consumption, which is separate from what you’re billed.
Blog
More from CloudForecast
Cloud Cost Management is Easy With CloudForecast
We would love to learn more about the problems you are facing around AWS and Azure cost. Connect with us directly and we’ll schedule a time to chat!
Start Free Trial