Mistral Patented Code-Based Tool Calling. And Your Agent Is Caught in the Middle
On June 30, 2026, the USPTO granted Mistral AI patent US 12,670,045 B1. Title: "Code implemented tool calls". It went unnoticed for six weeks. Then, on August 10, the document landed on Hacker News and racked up 223 points in a few hours.
What the patent describes: the LLM writes a block of code that calls the tools, the server runs that code in a sandbox, pauses when it hits a tool call, sends the call to the client to execute, gets the result back and picks up where it left off. If that sounds familiar, it's because it is. It's Anthropic's programmatic tool calling. It's Cloudflare's Code Mode. It's CodeAct, which hit arXiv in February 2024.
I read the patent at the source. All 20 claims, the full description, the list of references the examiner considered. Not the HN summary. This post is what you can actually say with the document in hand: what Mistral's tool calling patent really covers, where it stops, what the prior art is and what the real risk is for anyone building agents. Spoiler: it's smaller than the headline suggests. But it's not zero.
TL;DR
- What it is: patent US 12,670,045 B1, "Code implemented tool calls", granted to Mistral AI on June 30, 2026 (filed March 4, 2026, application 19/557,103, inventor Gabriel Vergnaud).
- The mechanism: LLM generates a code block (TypeScript, in the description) → server runs it in a sandbox → pauses at the pending tool call → client executes the tool → sandbox resumes, substituting in the result → only the final output goes back to the model.
- Obvious prior art: CodeAct (Feb 2024), smolagents, Cloudflare's Code Mode (Sep 2025), Anthropic's code execution with MCP and programmatic tool calling (Nov 2025).
- Risk for you: territorial and low. A US patent applies in the US. In Brazil, a "computer program as such" isn't even considered an invention.
- Primary source: the USPTO Official Gazette and the full text with all 20 claims.
What Mistral's tool calling patent says, word for word
Independent claim number 1, straight from the USPTO gazette:
A method, comprising: receiving, at a server, a user request for execution of one or more tool calls; generating, by a large language model (LLM), a code block in a programming language, the code block configured to encapsulate the one or more tool calls; executing, by the server, the code block in a sandbox; in response to obtaining a pending tool call, pausing execution of the code block; transmitting the pending tool call to a client for execution; receiving, from the client, a first result of the pending tool call; resuming execution of the code block and substituting the first result of the pending tool call for the pending tool call in the code block; and returning a second result of the executed code block to the LLM.
Breaking that into eight steps, because that's how a claim works — every comma is a lock:
- A server receives a request from the user.
- An LLM generates a code block in a programming language.
- That block encapsulates one or more tool calls.
- The server executes the block in a sandbox.
- When it hits a pending tool call, it pauses.
- It transmits the call to the client to execute.
- It receives the result and resumes, substituting the value for the call.
- It returns the final result to the LLM.
In the description, Mistral gets specific: the block is generated "in a programming language, such as TypeScript", parallelism comes from Promise.all and sequencing from await. The sandbox is described as "resumable".
In practice, the code the model emits looks like this:
// Não é JSON de tool call. É código que chama as tools como funções.
const funcionarios = await listarFuncionarios({ time: "engenharia" });
const estouros = await Promise.all(
funcionarios.map(async (f) => {
const despesas = await buscarDespesas({ id: f.id, trimestre: "Q3" });
const total = despesas.reduce((soma, d) => soma + d.valor, 0);
return total > f.limite ? { nome: f.nome, total } : null;
})
);
// Só isso volta pro contexto do modelo. As 20 respostas cruas, não.
return estouros.filter(Boolean);
Each await buscarDespesas(...) is a pending tool call: the sandbox freezes, the call travels to the client, the client runs it against its own database, sends the result back, the sandbox unfreezes. Twenty lookups, zero round-trips through the model.
This isn't exotic. It's the industry standard
Here's the detail that makes this post exist. Open Anthropic's programmatic tool calling documentation and read the "How programmatic tool calling works" section:
Claude writes Python code that invokes the tool as a function (...) When a tool function is called, code execution pauses and the API returns a
tool_useblock (...) You provide the tool result, and code execution continues (intermediate results are not loaded into Claude's context window).
Pause. Hand off to the client. Client executes. Resume. Intermediate results stay out of the context. That's all of claim 1, plus dependent claim 6 ("the LLM receives the second result and does not receive the first result"), written in the prose of public documentation.
And it's not just Anthropic. Cloudflare published Code Mode on September 26, 2025, more than five months before the filing. The post's thesis is literally the patent's title: "Convert the MCP tools into a TypeScript API, and then ask an LLM to write code that calls that API." Same language, same sandbox (V8 isolate), same conversion of schema into types.
The numbers behind this have already been published by the people who built it. Anthropic showed a drop from 150K to 2K tokens in a Google Drive → Salesforce flow, and an average of 43,588 down to 27,297 tokens — 37% — on complex research tasks. This isn't a nicety: it's the difference between an agent that fits the budget and one that doesn't.
In other words: the patent doesn't describe an invention locked away in a French lab. It describes the mechanism you probably already have running in production, if your agent is even remotely modern. Running this pattern safely going forward takes understanding the mechanism, not the headline — and that's the kind of teardown we do every week, live, at Beer And Code.
Where claim 1 stops
You don't read a patent by its headline. You read it by the all-elements rule: to infringe a claim, your implementation has to have every one of its elements. Miss one, and the claim doesn't reach you.
So let's be specific about what falls outside:
- Does your agent execute the tool in the same process as the sandbox? Claim 1 requires "transmitting the pending tool call to a client for execution". If there's no separate client doing the executing, it doesn't read on you.
- No sandbox? If the code runs directly in your worker, with no isolation, step 4 doesn't match.
- The code doesn't pause and resume? If you generate the block, resolve everything in one go and only then assemble the result, steps 5 through 7 don't exist.
- The result doesn't go back to the LLM? If the block is the final answer delivered to the user with no further model turn, step 8 doesn't match.
- Ran it locally, in a dev agent on your own machine? Private, non-commercial use isn't the arena patents live in.
The dependent claims are where the technically interesting part lives — and also the part most vulnerable to prior art:
- Claims 4, 5, 10 and 14, 15, 20: resumption by means of an "evaluation stack" that re-executes the block from the beginning, capturing and replaying the results of non-deterministic operations, wrapping those operations in a function that stores the initial value.
- Claim 8: translating tool definitions from JSON Schema into type definitions in the programming language, before the model generates the code.
- Claim 7: parallel with a parallelism construct, sequential with
await. - Claim 6: the model receives only the final result, never the intermediate ones.
Look at claims 4/5/10. That's deterministic replay:
// Claim 10: envolver a operação não determinística numa função
// que guarda o valor inicial, para o replay devolver o mesmo valor.
const agora = capturar(() => Date.now());
const id = capturar(() => crypto.randomUUID());
Hold on to that idea. It's coming back in the next section.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãPrior art: the list any dev can put together in ten minutes
- CodeAct — Executable Code Actions Elicit Better LLM Agents, February 1, 2024. The thesis: instead of the agent emitting JSON, it generates executable Python as a unified action space. Two years before the filing.
- smolagents, Hugging Face. The entire library is built on top of
CodeAgent: the agent writes code, the code runs in a sandbox (E2B, Docker, Modal), the tools are Python functions. - Cloudflare's Code Mode, September 26, 2025. Converts MCP into a TypeScript API and asks the LLM to write code. It covers the spirit of claim 8 with a date that predates the filing.
- Code execution with MCP (Nov 4, 2025) and Advanced tool use (Nov 24, 2025), from Anthropic. Pause, client result, resume, context preserved. All public more than three months before the filing.
- Deterministic replay: claims 4, 5 and 10 describe, under a different name, what Temporal has been doing since 2019 (and Cadence before it): re-executing the function from the start, feeding it the recorded history instead of redoing the work, and capturing non-deterministic operations so the replay produces the same value. Before that, event sourcing had already been doing this for two decades.
Now the detail that hurts. The patent's front page lists five patent documents as references considered — and no non-patent literature. No arXiv. No engineering blogs. No API documentation. The examiner assessed novelty against the body of knowledge that least represents this specific field, while the state of the art in agents lives precisely in what was left out: engineering posts, arXiv papers and GitHub repos.
That doesn't make the patent invalid on its own. It makes it fragile.
The real risk for anyone building agents in Brazil
Let's separate panic from engineering. I'm not a lawyer, and nothing here is legal advice — but three facts are checkable and they change the conversation.
One: patents are territorial. US 12,670,045 B1 is a United States patent. It gives Mistral the right to stop third parties from exploiting that invention on American soil. If you run a SaaS in São Paulo, serving Brazilian customers, on Brazilian infra, this patent doesn't reach you. If you sell into the US, that's a different conversation — and that's when it's worth talking to an actual IP lawyer.
Two: in Brazil this probably wouldn't even be granted. Lei 9.279/96, art. 10, item V, is explicit: "computer programs as such" are not considered inventions. The INPI does examine computer-implemented inventions, but it requires a technical effect beyond the processing itself. An orchestration of API calls in a sandbox has a pretty narrow path here.
Three: in Europe it's harder, not impossible. The HN meme is that "software patents don't exist in Europe". Half true. The EPO excludes computer programs "as such", but grants computer-implemented inventions when there's a further technical effect. Reducing memory usage and network calls is the kind of argument that has already gotten through over there. Don't count on automatic invalidity.
And the most important point for reading this move: a big software patent is rarely a weapon, it's a shield. Mistral is worth a fraction of its American rivals. A patent portfolio is what gives you something to put on the table when someone bigger shows up demanding a license. This is an IP cold war, not a startup hunt. The uncomfortable scenario isn't Mistral suing you in 2027 — it's the patent getting sold in 2030 to someone who litigates for a living.
If someone really wants to knock it down, the paths exist: post-grant review in the nine months after the grant (a window that closes around the end of March 2027), inter partes review after that, ex parte reexamination at any time, and the patent-ineligible subject matter defense under the Alice test in any eventual litigation. Prior art dated before March 4, 2026 is the one thing there's no shortage of.
What changes in your architecture
Nothing. And that's the part that needs to be said in so many words, because the wrong reaction to this kind of news is a dev abandoning a good pattern out of fear of a headline.
What changes is the hygiene around it:
- Keep using the pattern. If your agent calls twelve tools per task, writing code that orchestrates the calls is still the right architectural decision, on tokens, on latency and on composition. We've already torn the pattern down with code in Programmatic tool calling: let the agent write the code instead of calling tool by tool and compared cost and latency in Programmatic Tool Calling: why running your tools in code is the future of the agent.
- Publish what you invent. Defensive publication is cheap: a dated engineering post, a public repository, a versioned ADR. It becomes prior art. It's the most accessible mechanism a small team has against this game.
- Document the dates. If your product was already running this pattern before March 2026, have the commit, the changelog and the dated post. That's an asset.
- If you sell in the US, take it to legal. Not to panic: to get a freedom-to-operate analysis done by someone who knows how to read a claim.
Quick FAQ
Do I need to stop using programmatic tool calling or Code Mode? No. You use these features as a client of an Anthropic or Cloudflare API — the infringement risk, if there were any, would be theirs, not yours. And both published the mechanism before the patent was filed.
My agent runs the tool in the same process as the sandbox. Does that infringe? Claim 1 requires transmitting the pending call to a client for execution and receiving the result back. If there's no such round trip between server and client, an element is missing and the claim doesn't reach your implementation. That goes for all 20 claims, because the dependent ones inherit the elements of the independent one.
Does this affect people using MCP? Indirectly. Claim 8 (translating tool definitions from JSON Schema into language types) describes what practically every MCP-to-code bridge does. It's also exactly what Cloudflare published in September 2025, which gives you dated prior art.
Why was the patent granted so fast? Filed March 4, 2026, granted June 30, 2026: under four months. That kind of timeline usually points to accelerated examination. Combine that with a front page with no non-patent literature, and you can imagine how much of the real state of the art made it into the analysis.
The takeaway
Mistral's patent on code-based tool calling is a technically competent document describing a pattern the entire industry had already published. Claim 1 is the architecture of programmatic tool calling. Claim 8 is Code Mode. Claims 4, 5 and 10 are deterministic workflow replay under a new name. And the prosecution record cites five patents, no papers, no engineering posts.
What this really signals isn't a threat to your stack. It's that the "everyone publishes everything on the engineering blog" phase is being overtaken by the phase where the same patterns turn into legal assets. The two will coexist for a while, and anyone building agents needs to know how to read both.
The practical conclusion is boring and correct: keep building. Just start dating what you build.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã