AWS Bedrock: What It Is and How to Run Claude in Production with Governance (and the Bill in Reais)
AWS writes great documentation on how to turn on AWS Bedrock. Console, "Enable model access," three clicks, first call comes back. Lovely.
Nobody writes about the rest.
Nobody writes about the difference between CloudTrail and model invocation logging — and why the former won't save you in an audit. Nobody writes that the São Paulo region exists in Bedrock but doesn't give you data residency. And nobody writes about what shows up on the bill at the end of the month, converted to reais, with IOF on top.
That's what this post is about. We'll look at what AWS Bedrock actually is, wire Claude into it, lock down the four governance layers any serious company is going to demand from you, and run the numbers in reais for an internal agent running for real.
TL;DR
- What it is: AWS Bedrock is AWS's managed service that serves foundation models (including Anthropic's Claude) over an API, inside your AWS account, with native IAM, VPC, CloudTrail and billing.
- Stack: Python +
anthropic[bedrock], IAM, CloudWatch/S3, Bedrock Guardrails. - Cost/access: same per-token price as Anthropic's direct API. Regional endpoints cost 10% more. Batch is 50% cheaper.
- Useful links: Claude in Amazon Bedrock and Bedrock Pricing.
AWS Bedrock: what it is, in practice
AWS Bedrock is a managed service that exposes foundation models from several vendors — Anthropic, Meta, Mistral, Amazon, Cohere — behind a single API, inside your AWS account.
That's the definition. What it means in practice is more interesting: when you call Claude through Bedrock, the call never leaves your AWS perimeter. Authentication is SigV4 with IAM. Traffic can go over PrivateLink. The charges land on the same bill as RDS and S3. And, according to Anthropic's documentation, the inference infrastructure runs with zero operator access — nobody at Anthropic has access to it.
What Bedrock is not: a transparent proxy to Anthropic's API. It's a product operated by AWS, with its own release cycle, its own subset of features and its own model IDs. Treating the two as interchangeable is the most expensive mistake you can make in a migration — and we'll see exactly where it hurts.
One thing has changed and is worth noting: the current Claude integration on Bedrock serves the Messages API at the endpoint https://bedrock-mantle.{region}.api.aws/anthropic/v1/messages. It's the same request format as the first-party API. The old integration (InvokeModel and Converse, with those versioned ARN-style IDs like anthropic.claude-sonnet-4-5-20250929-v1:0) still exists, but it's the legacy path. New code goes on Mantle.
The context: why route Claude through Bedrock
The right question isn't "which one is cheaper." Per token, the price is the same: Haiku 4.5 at $1/$5, Sonnet at $3/$15, Opus at $5/$25 per million input/output tokens — identical on both paths.
The right question is: who has to sign off on it?
If the answer is "me," use Anthropic's direct API. You get structured outputs, the Files API, server-side tools, Managed Agents and access to new features on launch day.
If the answer involves a security team, a DPO, an auditor or a contract with a data residency clause, Bedrock wins for one reason only: it speaks the language those people already speak. There's no "AI vendor approval process" — there's an IAM policy, a CloudTrail log and a line on the AWS bill that finance already reconciles every month.
This is the kind of decision that separates a prototype from a product: it's not about the model being good, it's about the architecture getting past the committee. That's exactly the angle we break down live in Clã Beer and Code — AI systems built in front of everyone, with code running and the bill open on screen.
Prerequisites
- An AWS account with model access enabled for the model you're going to use (Bedrock console → Model access). Claude Fable 5, Opus 4.8, Sonnet 5, Opus 4.7 and Haiku 4.5 are open to any Bedrock customer.
- AWS credentials resolved through the default chain (env vars,
~/.aws/config, SSO, ECS task role, IMDS). - Python 3.11+ with
pip install -U "anthropic[bedrock]". - IAM permission for
bedrock-mantle:CreateInferenceon the ARNs of the allowed models.
How to use AWS Bedrock: the first call
Step 1: the client and the model ID
Anthropic's SDK has a dedicated client for Bedrock. It resolves credentials and region using AWS's standard precedence and signs with SigV4.
from anthropic import AnthropicBedrockMantle
client = AnthropicBedrockMantle(aws_region="us-east-1")
message = client.messages.create(
model="anthropic.claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(next(b.text for b in message.content if b.type == "text"))
Notice the anthropic. in front of the model ID. That prefix is mandatory on Bedrock and doesn't exist on the direct API. If you copied a snippet from the first-party docs and sent model="claude-opus-5", you get a 404. It's the first rock in the road, and the one that eats the most time from people who "already know how to use Claude."
If you prefer cURL, the shape is the same — only the signing changes:
curl https://bedrock-mantle.us-east-1.api.aws/anthropic/v1/messages \
--aws-sigv4 "aws:amz:us-east-1:bedrock-mantle" \
--user "$AWS_ACCESS_KEY_ID:$AWS_SECRET_ACCESS_KEY" \
-H "x-amz-security-token: $AWS_SESSION_TOKEN" \
-H "content-type: application/json" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "anthropic.claude-opus-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Claude"}]
}'
Step 2: pick the endpoint (and find out São Paulo doesn't solve residency)
This is where the trap nobody tells you about lives.
Bedrock offers two types of endpoint: Global, which routes dynamically across all available regions, and Regional, which resolves to a single region — the path for anyone with a data residency requirement. The regional endpoint costs 10% more than the global one.
Now look at the region table in the documentation: sa-east-1 (South America, São Paulo) shows up with exactly one supported endpoint type: Global.
In other words: you can call Claude from São Paulo, but inference will be routed globally. There's no "SA" geographic profile the way there is for US, EU, JP and AU. If your customer's contract says "the data doesn't leave Brazil," Claude on Bedrock doesn't meet that requirement today — and finding that out after legal has approved the architecture is an expensive problem.
For anyone with regional residency (EU, US, JP, AU), the path is the geographic inference profile — those IDs prefixed with us., eu., apac.. The AWS documentation is explicit about the nuance: the request stays inside the geography, but "your input prompts and output results may move outside of your source region" during cross-region routing.
Step 3: the 403 that looks like a permissions problem and isn't
If your organization uses SCPs to block unused regions — and most large companies do — the geographic inference profile will break.
Here's why: a us. profile routes across us-east-1, us-east-2 and us-west-2. If the SCP only allows us-east-1, the call fails when the router picks another region. The error comes back as a 403 and looks like missing permission on the model. It isn't: it's the SCP blocking the destination region.
The correct policy allows the profile and the foundation model in every destination region:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "GrantInferenceProfileAccess",
"Effect": "Allow",
"Action": "bedrock:InvokeModel",
"Resource": [
"arn:aws:bedrock:us-east-1:<ACCOUNT_ID>:inference-profile/us.anthropic.claude-sonnet-4-5-20250929-v1:0"
]
},
{
"Sid": "GrantModelAccessInAllDestinations",
"Effect": "Allow",
"Action": "bedrock:InvokeModel",
"Resource": [
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-5-20250929-v1:0",
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-sonnet-4-5-20250929-v1:0",
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-sonnet-4-5-20250929-v1:0"
],
"Condition": {
"StringEquals": {
"bedrock:InferenceProfileArn": "arn:aws:bedrock:us-east-1:<ACCOUNT_ID>:inference-profile/us.anthropic.claude-sonnet-4-5-20250929-v1:0"
}
}
}
]
}
The bedrock:InferenceProfileArn condition is the detail worth its weight in gold: it opens the destination regions only when the call comes through the approved profile. Without it, you've opened three regions to any invocation.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãGovernance: the four layers
Governance on Bedrock isn't a feature, it's the composition of four independent controls. Fail at any one of them and the others are worthless.
1. Who can call what (IAM + SCP)
The basics: an identity policy restricting bedrock-mantle:CreateInference to the ARNs of the approved models. The less basic: using IAM to enforce the guardrail. You can deny any inference request that doesn't come with an approved guardrail identifier. Without that, the guardrail is optional in practice — and what's optional gets ignored under deadline.
2. What the model can say (Guardrails)
Bedrock Guardrails applies consistent filters on input and output: denied topics, content filters, PII redaction, grounding checks against a source. It's the layer that makes sense when the output goes to the end customer, not to your terminal.
The architecture point: a guardrail is not a substitute for a well-written prompt, it's the safety net for when the prompt fails. Treat it as a circuit breaker, not as business logic.
3. Where the traffic goes (VPC + PrivateLink)
A VPC endpoint for Bedrock, a policy on the endpoint restricting models, and the call never touches the public internet. Combined with layer 1, you have a closed network path.
4. What got recorded (and the CloudTrail gotcha)
This is the one that creates the most false sense of security.
CloudTrail records the call as a management event: who called, when, from where, which model. Metadata, not content. It's on by default. And it does not show what was asked of the model or what the model answered.
What stores the prompt and completion is model invocation logging, which writes to S3, CloudWatch Logs or both. And it's opt-in — it ships turned off.
The practical result: a whole lot of companies are in production thinking they have an AI audit trail because "CloudTrail is on." They don't. They have an access log. The day someone asks "what customer data did this agent see in March?", the answer is going to be silence.
Turning it on is one call:
aws bedrock put-model-invocation-logging-configuration \
--logging-config '{
"cloudWatchConfig": {
"logGroupName": "/aws/bedrock/invocations",
"roleArn": "arn:aws:iam::<ACCOUNT_ID>:role/BedrockLoggingRole"
},
"s3Config": {
"bucketName": "meu-bucket-bedrock-logs",
"keyPrefix": "invocations/"
},
"textDataDeliveryEnabled": true
}'
CloudWatch for real-time alarms and metrics; S3 for long retention and replication. Anthropic recommends keeping at least 30 days of rolling logs. And remember: you now have user prompts sitting in a bucket. KMS encryption, a restricted access policy and an expiration lifecycle stop being optional.
Bedrock pricing: the bill in reais
Let's run the numbers on a concrete case, not a "hello world."
Scenario: an internal support agent. 10,000 requests per month. Each one sends 8,000 input tokens (system prompt + 6 documents retrieved via RAG) and returns 600 tokens. Sonnet-class model, at $3 / $15 per million.
With no optimization at all:
- Input: 10,000 × 8,000 = 80M tokens → 80 × $3 = $240
- Output: 10,000 × 600 = 6M tokens → 6 × $15 = $90
- Total: $330/month
At R$ 5.22 to the dollar (August 2026 exchange rate), that's R$ 1,722 per month. Before IOF, before your card's FX spread.
With prompt caching: the system prompt and the base documents add up to ~7,000 of the 8,000 input tokens, and they're stable across requests. Bedrock supports prompt caching for Claude with two TTLs: 5 minutes (writes at 1.25x, reads at 0.1x) and 1 hour (writes at 2.0x, reads at 0.1x).
- Cache reads: 70M × $3 × 0.1 = $21
- Variable input: 10M × $3 = $30
- Cache writes (~300 in the month): 2.1M × $3 × 1.25 ≈ $8
- Output: $90
- Total: ≈ $149/month → R$ 778
A 55% cut from changing where you put the cache breakpoint. Not a single line of business logic changed.
And remember that tokens are just one line item. The real bill for an agent in production has five more buckets beyond inference, and I opened up the full ledger in how much an agent in production costs in 2026.
Three warnings about these numbers:
The cache has a break-even point. A write costs more than normal input. Below a ~30% hit rate, each write costs more than the reads save. If your prefix changes on every request, you're paying 1.25x to use nothing. Measure cache_read_input_tokens in the response — if it comes back zero on consecutive requests with the same prefix, there's a silent invalidator in your prompt (timestamp, UUID, JSON without stable ordering).
Regional endpoints cost 10% more. If data residency forces you onto regional, add 10% on top of everything. In the example, $149 becomes $164.
Batch cuts 50% — but it's not Anthropic's batch. Bedrock has its own batch inference (CreateModelInvocationJob): you drop a JSONL in S3, it processes asynchronously within 24h and writes the results back to S3, at half the on-demand price. Anthropic's Message Batches API is not available on Bedrock. They're different APIs, with different integrations. If your pipeline tolerates hours of latency — ticket classification, dataset enrichment, overnight summary generation — half the bill is on the table.
Limitations and things to watch
What you lose by choosing Bedrock over the direct API, according to the official documentation:
- Structured outputs aren't supported. If your code depends on
output_config.formatto guarantee valid JSON, you're going back to tool use withstrict: trueor to validation with retry. - Server-side tools don't run: web search, web fetch, code execution, advisor. Everything that depended on Anthropic executing the tool becomes your responsibility.
- Files API and URL sources are out. Attachments go as base64 in the request.
- MCP connector, Agent Skills and Managed Agents are out. An agent on Bedrock is a loop you write and host.
- Server-side fallback (the
fallbacksparameter) doesn't exist. Use the client-side pattern.
What stays: the full Messages API, prompt caching, thinking, tool use (including bash, computer use, memory, text editor) and citations.
On quota: the default is 2 million input tokens per minute, and you can request up to 4 million without additional approval from Anthropic. Requests-per-minute limits are imposed by AWS — that ceiling gets adjusted through AWS support, not through Anthropic.
FAQ
Why am I getting a 404 on the model ID?
You probably left out the anthropic. prefix. On Bedrock it's anthropic.claude-opus-5, not claude-opus-5. If the prefix is right, check that model access is enabled for that model in your account, in the Bedrock console.
403 even with the correct IAM policy. Now what? If you're using a geographic inference profile, check the organization's SCPs. The profile routes across several regions and the SCP has to allow all the destination regions, not just the source one. Blocking any destination region breaks cross-region inference entirely.
Does CloudTrail give me an AI audit trail? No. CloudTrail records call metadata. To store the prompt and response you need to enable model invocation logging, which is off by default.
Can I keep the data in Brazil?
Today, not through Bedrock. The sa-east-1 region only offers the Global endpoint for Claude, and there's no geographic profile for South America. If residency in Brazil is a contractual requirement, this architecture doesn't work.
Is it worth migrating from the direct API to Bedrock? Only if you need the AWS perimeter or consolidated billing. The per-token price is the same, and you lose a meaningful set of features. Migrate for governance reasons, never for cost. If your question really is about cost, the comparison that matters is a different one: local LLM vs API, with a 90-day spreadsheet.
Wrapping up
Bedrock doesn't make Claude cheaper or smarter. It changes who answers the hard questions: where the data went, who authorized it, what got recorded and how much it cost.
If your context has those questions, the job isn't turning the service on — it's locking down the four layers. IAM that restricts and enforces the guardrail. Guardrail as a circuit breaker. Private network. And model invocation logging turned on, because the layer everyone forgets is precisely the one the auditor will ask for first.
The natural next step is evals: you now have the invocation logs in S3. That's a dataset of real cases waiting to become an evaluation suite. A model in production without evals is faith, not engineering.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã