Oracle Generative AI: Which Models OCI Has and When to Choose It Over Bedrock
Everybody compares OpenAI with Anthropic. Some people compare Bedrock with Vertex. And almost nobody in Brazil writes about the third LLM infrastructure provider, the one that already has a São Paulo region, a catalog of 40+ models, and a billing model that exists nowhere else: Oracle Generative AI.
This isn't hype. It's that Oracle's documentation is dense, poorly indexed, and written for enterprise architects, not devs. The result is a service that runs Grok, Gemini, Llama, Cohere, and gpt-oss behind the same OpenAI-compatible API, and that practically nobody evaluates before signing the Bedrock contract.
In this post you'll see the real OCI Generative AI model catalog, how per-character billing works (and why it changes the math in Portuguese), how to call the service with the OpenAI SDK in five lines, and the objective criteria for deciding between OCI and Bedrock.
TL;DR
- What it is: Oracle Cloud's managed service for chat, embeddings, rerank, and agents, with models from Cohere, Meta, xAI, Google, and OpenAI.
- Stack/Models: Grok 4.3 (1M context), Gemini 2.5 Pro/Flash, Llama 4 Maverick and Scout, Cohere Command A, gpt-oss-120b/20b, Cohere Embed 4 and Rerank 4.
- Cost/Access: on-demand billed per character (1 transaction = 1 character) or dedicated cluster billed per unit-hour, with a 744 unit-hour minimum.
- API: OpenAI-compatible endpoint at
https://inference.generativeai.<região>.oci.oraclecloud.com/openai/v1. - Useful link: official OCI Generative AI documentation.
The context: why look at Oracle Generative AI now
For two years the answer to "where do I run my LLM in production" was trivial: OpenAI directly, or Bedrock if the company required a cloud contract. OCI stayed out of the conversation because the catalog was small — basically Cohere and Meta — and Oracle sold the service to people who already had Fusion, NetSuite, or an Oracle Database estate to justify it.
That changed quietly. Today the pretrained model catalog includes xAI Grok 4.3 with a 1 million token context window, the Google Gemini 2.5 family (Pro, Flash, and Flash-Lite), Llama 4 Maverick and Scout, Cohere Command A with 256,000 tokens, and OpenAI's open models, gpt-oss-120b and gpt-oss-20b. In December 2025 Model Import landed, which lets you bring your own weights. In June 2026 the list of importable models expanded to cover MiniMax, Mistral, Moonshot's Kimi, and Z.ai's GLM.
And there's the detail that matters to anyone building product in Brazil: OCI Generative AI is available in the Brazil East (São Paulo) region, with 25 models served from inside the country. In-country inference isn't legal nitpicking — it's the difference between the compliance team approving the project in two weeks or in two quarters.
Choosing where to run the model has become an architecture decision, not a matter of taste: it changes price, latency, data jurisdiction, and the length of the rope tying you to the vendor. That kind of decision, with the numbers on the table and the code running, is what we pick apart every week in the Clã Beer and Code, and there are people in there building in Python, Go, and .NET — with AI in the mix, the vendor stopped being the problem; knowing how to compare is what became the job.
Which models OCI Generative AI has today
The catalog splits into four tasks: chat, embeddings, rerank, and voice. The main active models:
| Model ID | Provider | Context | Note |
|---|---|---|---|
xai.grok-4.3 |
xAI | 1,000,000 | largest window in the catalog |
xai.grok-4.20-multi-agent |
xAI | — | multi-agent variant |
google.gemini-2.5-pro |
— | heavy reasoning | |
google.gemini-2.5-flash |
— | cost/latency | |
cohere.command-a-03-2025 |
Cohere | 256,000 | enterprise generalist |
cohere.command-a-reasoning |
Cohere | 256,000 | reasoning |
cohere.command-a-vision |
Cohere | — | multimodal |
meta.llama-4-maverick-17b-128e-instruct-fp8 |
Meta | — | MoE, open weights |
meta.llama-4-scout-17b-16e-instruct |
Meta | — | smaller MoE |
openai.gpt-oss-120b |
OpenAI | — | open weights |
For RAG, what matters is the other two groups. For embeddings the active model is cohere.embed-v4.0, multimodal (text and image) — the entire Embed 3 family, including the multilingual variants, is marked as deprecated. For rerank, it's cohere.rerank-4 in the Pro and Fast variants, with Rerank 3.5 also deprecated.
Notice the pattern: the catalog turns over fast and Oracle deprecates without ceremony. The xAI model list currently has more IDs marked as deprecated (Grok 3, 3-mini, 3-fast, 4, 4-fast, 4.1-fast, code-fast-1) than active ones. If you pin a model ID in your .env and forget about it, one day the call comes back with an error. Treat the model ID as a versioned dependency, with a deprecation alert on the team's board.
What it does not have, and it's the loudest absence: no Anthropic models. No Claude, in any variant.
Per-character pricing: the catch that changes the math
Here's the most concrete technical difference between OCI and everyone else. Bedrock, Vertex, and OpenAI bill per token. OCI Generative AI bills per transaction, and the cost calculation documentation is explicit about what that means:
"One transaction equals to one character."
One transaction is one character. On-demand pricing is published per 10,000 transactions, meaning per 10,000 characters. A few order-of-magnitude examples: Llama 4 Scout and Maverick at around $0.0018 per 10,000 transactions, larger Cohere models at $0.0156, and embedding models around $0.001.
Why does this matter in practice? Two reasons.
First: the math becomes deterministic. You don't need a tokenizer to estimate cost. mb_strlen() does the job.
// OCI: 1 transação = 1 caractere. Estimativa exata, sem tokenizer.
$prompt = file_get_contents(storage_path('app/contrato.txt'));
$chars = mb_strlen($prompt);
$precoPor10k = 0.0156; // USD / 10.000 transações
$custoUsd = ($chars / 10_000) * $precoPor10k;
// 40.000 caracteres => 40.000/10.000 * 0.0156 = US$ 0,0624
Try doing that with per-token billing and you'll import tiktoken, find out Llama's tokenizer is a different one, and still get the estimate wrong.
Second: Portuguese is expensive in tokens and neutral in characters. Tokenizers are trained mostly on English. A Portuguese word breaks into more tokens than its English equivalent — accents, long suffixes, conjugation. On a model billed per token, the same text translated costs more in Portuguese. With per-character billing, that language tax simply doesn't exist: cost tracks the length of the text, not the luck of the tokenizer.
For a Brazilian product that processes high volumes of Portuguese text — customer service, contracts, clinical records, support tickets — this stops being a curiosity and becomes a spreadsheet line. One warning, though: if your bill is already out of whack, switching providers is the second step. The first is finding the token leaks in what you already run — migrating waste to another cloud only changes whose name is on the invoice.
Hands-on: calling OCI with the OpenAI SDK
Oracle published OpenAI-compatible endpoints. You don't rewrite your integration: you swap base_url, authentication, and the model name.
Step 1: install the auth helper
Authentication is not a Bearer token. It's OCI IAM request signing, and that's what the oci-genai-auth package handles.
pip install openai oci-genai-auth
Step 2: point the client at the region
import httpx
from openai import OpenAI
from oci_genai_auth import OciSessionAuth
client = OpenAI(
base_url="https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1",
api_key="not-used", # a auth real vai no http_client
http_client=httpx.Client(auth=OciSessionAuth(profile_name="DEFAULT")),
)
resp = client.chat.completions.create(
model="xai.grok-4.3",
messages=[{"role": "user", "content": "Explique embeddings em uma frase."}],
)
print(resp.choices[0].message.content)
Three things to get right here. The api_key is ignored, but the OpenAI SDK requires the key to be filled in — put any string there. The profile_name points to a profile in your ~/.oci/config, so oci setup config needs to have been run first. And the region in the base_url defines the catalog: asking for xai.grok-4.3 on a São Paulo endpoint returns a model-not-found error, not a fallback.
Step 3: choose the region with your eyes open
This is the part the documentation hides well and that takes projects down in staging. Availability by region is not uniform:
- US Midwest (us-chicago-1): about 41 models, the full catalog, including Grok and Gemini.
- Brazil East (sa-saopaulo-1): about 25 models, no Grok and no Gemini, and with a good chunk of the active ones available only on a dedicated cluster (Command A Reasoning, Embed 4, Rerank 4 Pro, Llama 4).
Translation: data in Brazil and Grok 4.3 are, today, mutually exclusive requirements. Find that out during architecture design, not on go-live day.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãOn-demand or dedicated cluster?
Two ways to serve. On-demand is pay-as-you-go per character, no commitment — it's where every project should start. A dedicated cluster gives you reserved GPU, predictable throughput, fine-tuning, and hosting for your own model.
The number that decides it is the minimum commitment. The dedicated cluster billing documentation sets 744 unit-hours per hosting cluster. That's 31 calendar days: you rent the whole month, whether you use it or not. With a large Cohere model at around $24 per unit-hour, the full-month math easily tops $17,000 per unit — before any inference.
Fine-tuning is cheaper to experiment with: the minimum is 1 unit-hour per job, although some models require at least 2 units. And one exception that's worth gold: a model imported via Model Import does not carry the 744 unit-hour commitment.
Rule of thumb: only move to dedicated when your monthly on-demand spend is already close to the cost of the cluster, or when the requirement is tenancy isolation. Outside those two cases, dedicated is burning budget to buy predictability you don't need yet.
When to choose OCI and when to stay on Bedrock
No picking sides. The criterion is the use case.
Stay on Bedrock when:
- You need Claude. It's the simplest tiebreaker in this post. OCI has no Anthropic models, and Bedrock has the full Claude family — the governance path and the cost in reais for that scenario are detailed in AWS Bedrock: what it is and how to run Claude in production. If your product's quality depends on it, the discussion is over.
- Your data already lives in S3. Keeping inference next to the data avoids an export pipeline, egress costs, and a whole security surface.
- You want the widest third-party catalog. Claude, Llama, Mistral, Cohere, AI21, and Titan under a single API.
- Your team already operates AWS IAM, VPC, and billing. A new provider is an operations learning curve, not just an API one.
Look at OCI Generative AI when:
- Data residency in Brazil is a hard requirement. Inference in São Paulo, inside the same tenancy as the rest of the application.
- Your volume is long-form text in Portuguese. Per-character billing eliminates the language's tokenization tax.
- The company is already an Oracle shop. Fusion, NetSuite, Autonomous Database. Running inference next to the transactional data kills the same problem S3 kills on AWS, just from the other side.
- You want Grok or Gemini under an enterprise cloud contract. Grok 4.3 with 1M context and Gemini 2.5 under a single Oracle contract is a combination Bedrock doesn't offer.
- You're going to host your own weights. Model Import without the 744 unit-hour floor is a concrete cost advantage.
And then there's the answer nobody likes to give: both. The OpenAI-compatible endpoint makes it cheap to keep a thin abstraction layer and route by task — Claude on Bedrock for what demands reasoning, Command A on OCI São Paulo for the Portuguese volume with sensitive data.
Limitations and things to watch
No Anthropic. Repeating it because it's the point that most often knocks OCI off the shortlist. There's no Claude, and no public timeline for it.
Aggressive model churn. The volume of deprecated IDs in the xAI catalog is high. A model ID in production needs monitoring, not faith.
Brutal asymmetry between regions. 41 models in Chicago versus 25 in São Paulo, and several of the good ones only on dedicated. Validate availability by region before locking in the architecture.
Authentication is cloud auth, not an API key. Request signing via OCI IAM, with ~/.oci/config, a profile, and a key. In containers and CI that costs configuration — plan for instance principal or workload identity from the start.
Fragmented documentation. Model, region, price, and API live on four pages that don't reference each other well. That's exactly why this post exists.
744 unit-hour commitment. A full month of cluster, billed regardless of usage. A common mistake for people who spin up a cluster "just to test."
Quick FAQ
Do I need to rewrite my code to migrate from OpenAI to OCI?
No, if you use the official OpenAI SDK. Swap base_url, inject OCI authentication into the http_client, and adjust the model name. Tool calls and streaming follow the same format. Integration through the old Assistants API is another story — the equivalent path on OCI is the Responses API.
How do I calculate cost if billing is per character?
Count the request's characters with mb_strlen() (or equivalent) and divide by 10,000 to get the number of transaction blocks. No tokenizer, no estimating. Details in the official cost calculation.
Can I use Llama 4 on-demand in São Paulo?
No. In Brazil East, Llama 4 Maverick and Scout show up as available only on a dedicated cluster. For on-demand Llama in the region, the path is meta.llama-3.3-70b-instruct.
Is it worth it if you're not an Oracle customer? It's worth evaluating in two scenarios: data residency in Brazil and high volumes of Portuguese text. Outside of those, the integration advantage favors those already in the Oracle camp, and Bedrock usually wins on catalog and ecosystem.
Conclusion
Oracle Generative AI is not the best LLM provider on the market, and that's not what's at stake. It's a legitimate provider, with a competitive catalog, an OpenAI-compatible API, a region in Brazil, and a per-character billing model that solves a real problem for anyone processing Portuguese at scale. It stayed off the radar because the documentation is bad, not because the product is.
What comes next is predictable: with Model Import, the Responses API, and the OpenAI-compatible endpoint, providers are converging on the same surface. When the API is a commodity, the decision moves for good to where it always should have been — where your data lives, how much a character costs, and which jurisdiction you accept.
Now that you have the criteria, the useful exercise is to take your most expensive production prompt, measure it in characters, and compare it with last month's token bill. The math answers better than any benchmark.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã