~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / openrouter-what-it-is-how-to-use-it $
Tutorials

OpenRouter: What It Is, How to Use It, and When It's Worth It

LS Lucas Souza · · 9 min read
OpenRouter: What It Is, How to Use It, and When It's Worth It

There's a pattern that's been repeating on OpenRouter for about two years: an ownerless model shows up in the list, free, named after an animal, and starts eating everyone's traffic. On April 28, 2026, it was Owl Alpha. It ran in disguise for two months, climbed to the top of the platform by call volume, and only on June 30 did Meituan own up to it: it was LongCat-2.0, 1.6 trillion parameters in an MoE — we told that whole story here. If you tried that model, you've already used OpenRouter without stopping to ask what OpenRouter actually is.

The short answer: it's a single endpoint, compatible with the OpenAI API, that talks to hundreds of models from dozens of different providers. One key, one base URL, one payload format. Switching models becomes switching a string.

In this post you'll see what it does under the hood, the calls that matter (curl, Python, and PHP), how to control routing and cost, and — the part almost nobody writes about — when this extra layer is not worth it.

TL;DR

  • What it is: a unified LLM gateway. One OpenAI-compatible endpoint for 400+ models and providers, with automatic fallback between them.
  • Stack/Models: any OpenAI client (Python, Node, PHP, Go) pointed at https://openrouter.ai/api/v1.
  • Cost/Access: model pricing is passed through with no markup; the fee is on credit purchases (5.5% via Stripe, $0.80 minimum; 5% in crypto). There are :free models with request limits.
  • Useful link: openrouter.ai/docs/quickstart

What OpenRouter is, in practice

Before it, supporting three models in production meant three SDKs, three error formats, three auth schemes, and three billing contracts. You wrote one adapter per provider and prayed none of them changed the response shape.

OpenRouter swaps that for one indirection. You speak OpenAI-compatible to it; it speaks each provider's native dialect underneath. The model slug carries all the information in autor/modelo format — anthropic/claude-sonnet-4.5, openai/gpt-latest, deepseek/deepseek-chat. Switching labs means editing one line in your .env.

The interesting part isn't the translation, it's the routing. The same open source model is usually hosted by several providers, with different pricing, latency, and quantization. By default, OpenRouter load-balances by price: it drops anyone with a recent outage and picks among the rest at random, weighted by the inverse square of the price — a provider at $1 per million tokens is 9x more likely to be picked than one at $3, according to the docs. A provider goes down, it tries the next one without you ever knowing.

That's an architecture decision, not a matter of taste: you're outsourcing inference failover to a third party. It's the kind of call we settle with running code and numbers on the table, live every week, in Clã Beer and Code. It's paid, it's a subscription, and it's exactly the environment this post describes.

Prerequisites

  • An account at openrouter.ai and a key in OPENROUTER_API_KEY.
  • Purchased credit (or none, if you're just going to play with the :free models).
  • Any OpenAI-compatible client: openai in Python/Node, Laravel's Http in PHP, or plain curl.
  • Awareness that base_url is configurable. That is literally the trick.

Hands-on

Step 1: the first call

No new SDK. Just point your existing client somewhere else.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-4.5",
    "messages": [{"role": "user", "content": "Explique embeddings em 2 frases."}]
  }'

In Python, the OpenAI SDK handles everything:

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

resp = client.chat.completions.create(
    model="deepseek/deepseek-chat",
    messages=[{"role": "user", "content": "Explique embeddings em 2 frases."}],
)
print(resp.choices[0].message.content)

Swapped deepseek/deepseek-chat for google/gemini-3-pro? Same call, same parsed response, different invoice. That's the product.

Step 2: model fallback

This is where the real difference from calling the provider directly begins. You declare a list of models: if the first one fails (rate limit, content filter, provider down), the request slides to the next one without throwing an exception in your code.

{
  "model": "anthropic/claude-sonnet-4.5",
  "models": ["openai/gpt-latest", "google/gemini-3-pro"],
  "messages": [{"role": "user", "content": "..."}]
}

The response includes a model field with whoever actually served it. Log that field. Without it, you don't know which model generated the output your customer complained about — and that wrecks any serious eval.

You can go one level down and control the provider:

{
  "model": "meta-llama/llama-4-70b-instruct",
  "provider": {
    "sort": "throughput",
    "ignore": ["Provedor-X"],
    "data_collection": "deny",
    "max_price": { "prompt": 1.0, "completion": 3.0 }
  }
}

data_collection: "deny" only routes to providers that don't retain your prompt. max_price sets a cap per million tokens. sort accepts price, throughput, or latency — and turns off price-based load balancing, going in deterministic order instead.

Step 3: pricing and catalog via the API

The /api/v1/models endpoint returns the entire catalog with pricing, context length, and supported parameters. You can build a budget-based model picker without hardcoding a price table that goes stale in two weeks:

curl -s https://openrouter.ai/api/v1/models \
  | jq -r '.data[] | select(.context_length >= 200000)
           | [.id, .pricing.prompt, .pricing.completion] | @tsv' \
  | sort -k2 -n | head

Step 4: in Laravel

No package, no new abstraction:

$response = Http::withToken(config('services.openrouter.key'))
    ->post('https://openrouter.ai/api/v1/chat/completions', [
        'model'  => 'anthropic/claude-sonnet-4.5',
        'models' => ['openai/gpt-latest'],
        'messages' => [
            ['role' => 'system', 'content' => 'Você responde em PT-BR, direto.'],
            ['role' => 'user', 'content' => $pergunta],
        ],
    ])->throw()->json();

Log::info('llm.call', [
    'model_usado' => $response['model'],
    'tokens' => $response['usage']['total_tokens'] ?? null,
]);

If you already use Prism or the Vercel AI SDK, both accept an OpenAI-compatible driver with a custom baseUrl. Same story.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

When it's worth it (and when it isn't)

It's worth it when:

  • You need to compare models with the same code, without rewriting an adapter for every eval.
  • Your product has routes with different costs: a cheap model to classify, an expensive one to write. It's one of the levers we break down in how to cut AI API costs.
  • You want access to a Chinese or open source model hosted by third parties without opening accounts in five places.
  • Availability matters more than control: fallback between providers comes for free.

It's not worth it when:

  • You use a single model and don't plan to switch. Then it's one more network hop and one more point of failure between you and Anthropic/OpenAI.
  • You depend on a specific cutting-edge feature of the provider (prompt caching with its own rules, batch API, latency tiers, beta features). A gateway's common denominator is always smaller than the native API.
  • You have a contract or compliance requirement that demands a direct relationship with the provider. In that case, look at BYOK: you plug in your own key and use only the routing, with a $25K/month allowance on the pay-as-you-go plan and 5% of the equivalent cost above that.

Limitations and things to watch

The same slug can run on providers with different quantization. An int4 and an fp8 of the same model don't deliver the same quality — and if you haven't pinned provider.only or quantizations, yesterday's eval may not reproduce today. This is the most annoying and most silent bug in the layer.

Routing also leaks into determinism: with models and fallback on, two identical calls can be served by different models. Great for uptime, terrible for debugging a regression. In an evaluation pipeline, turn fallback off.

Privacy: by default OpenRouter logs neither prompt nor completion — only metadata (timestamp, model, token count). There's an opt-in that trades logging for a 1% discount. Before turning that on in a product with customer data, read the whole contract. And remember the final provider has its own policy: that's what data_collection: "deny" is for.

Finally, latency. Every gateway adds a hop. For interactive chat it's noise; for a pipeline with dozens of chained calls, measure before assuming it's irrelevant.

FAQ

Is OpenRouter more expensive than going straight to the provider? Per token, no: pricing is passed through with no markup. The fee is on credit purchases, 5.5% by card ($0.80 minimum) or 5% in crypto. At volume, that's the price of failover.

Are the :free models good enough for production? No. With no credit on the account the limit is 50 requests per day; with $10 or more, it goes up to 1,000 per day. It's a test bench, not infrastructure.

How do I know which model answered when I use openrouter/auto? From the model field in the response. The Auto Router classifies the task and picks based on the platform's spend share over the last 7 days, filtered by the cost_tier you set. If you don't log that field, you lose traceability.

Can I use my own OpenAI key behind it? Yes, via BYOK. You keep the commercial relationship with the provider and use OpenRouter only as a routing and observability layer.

The point

OpenRouter isn't magic. It's a well-executed indirection: one endpoint, one format, and the decision of which model serves the request becoming configuration instead of a deploy.

The real win isn't saving on SDKs. It's that the cost of switching models drops so far that you start testing for real — and then you find out, with numbers, that the $0.75-per-million model handles 80% of your routes. That's how a model with no name, no owner, and no release became the platform's top model before any launch post existed.

The next step is to stop picking models on intuition: instrument your calls, store the model field, and let the data decide.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing