Ox Alpha Was GLM-5.3-Flash: Z.ai Confirmed It, Opened the Weights Under MIT — and the 80% Benchmark Is Still Fake
Ox Alpha was GLM-5.3-Flash. Z.ai confirmed it, the weights are on Hugging Face under an MIT license, and the model that ran free for a week now costs $0.075 per million input tokens.
Five days ago, when nobody had claimed authorship, we wrote here that the tokenizer evidence pointed to Zhipu — 95 probes out of 95, zero mean error against the GLM-5 vocabulary. We wrote "high confidence, not confirmed."
Now it's confirmed.
In this post you'll see how the 80% benchmark fell apart, the $0.00 method that traced the authorship before the announcement, and what GLM-5.3-Flash actually delivers now that it has a name, a price and open weights.
Update, August 26, 2026: what changed since publication
This post went out on August 25 with the model still anonymous. In less than 24 hours, three things happened:
Z.ai confirmed it. On Wednesday, the company told Bloomberg that Ox Alpha is a new iteration of the GLM series and that the weights would drop that night.
The model got a name and a price. stealth/ox-alpha left the OpenRouter catalog and reappeared as z-ai/glm-5.3-flash. Free access is over: it's now $0.075 per million input tokens and $0.25 for output, with a 50% discount announced through September 9. Cache reads at $0.015.
The weights opened up. zai-org/GLM-5.3-Flash was published on Hugging Face on August 26 under an MIT license — the most permissive there is, commercial use allowed with no acceptable-use clause.
And the Z.ai release notes finally gave the architecture numbers: 320 billion total parameters, 18 billion active, combining linear and sparse attention, with native vision.
The rest of the post still stands — including the number we measured. The order in which things happened is worth reading.
TL;DR
- What it is: GLM-5.3-Flash, a multimodal reasoning model from Z.ai (Zhipu), focused on code and long agentic work. It circulated anonymously as "Ox Alpha" on OpenRouter from August 20 to August 26.
- Architecture: 320B total parameters, 18B active. Combined linear + sparse attention. Native vision.
- Specs: 1,048,576 tokens of context, up to 131,072 output tokens. Accepts text, image and video. Function calling and JSON output.
- Cost: $0.075 input and $0.25 output per million, with 50% off through September 9. It was $0.00 during the stealth phase.
- Weights: open on Hugging Face under an MIT license.
- Performance: ~63% on the full DeepSWE (113 tasks) — not the 80% from the headline, which came from a sample of ten.
The benchmark that went viral (and why it never backed up the headline)
The headline was born from a DeepSWE test. In the initial sample, the numbers were these:
| Model | DeepSWE (10-task sample) |
|---|---|
| Ox Alpha | 80% |
| Claude Fable 5 | 65% |
| GLM-5.3 | 62% |
| Grok 4.6 | 62% |
| GPT-5.6 Sol | 52% |
Looks decisive. Except DeepSWE has 113 tasks. Ten tasks is less than 9% of the benchmark.
With a sample that size, each task is worth 10 percentage points. Getting one more or one fewer right moves the result from 80% to 70% or 90%. That's not measurement, it's noise with decimal places.
That's what happened when people ran the full set: two independent runs of the 113 tasks landed near 63% — practically tied with GPT-5.6 Sol, not 15 points ahead of Claude Fable 5.
That number survived the confirmation. In the discussions that followed Z.ai's announcement, 63% on DeepSWE remained the figure people cited — not 80%.
There's a second, quieter trap in the tables that made the rounds: several of them compare Ox Alpha's 80% on DeepSWE against the ~96% that Claude Opus 5 and GPT-5.6 Sol score on SWE-bench Verified. Those are different benchmarks, with different difficulty and methodology. Putting both in the same column isn't analysis — it's a spreadsheet error turning into a headline.
The difference between swallowing a chart and reading a chart is the same difference between using AI and building with AI. The second one is what we train every week in the Clã Beer and Code.
None of this means GLM-5.3-Flash is bad. 63% on DeepSWE is a frontier model with 1M context, costing seven and a half cents per million tokens, with MIT-licensed open weights. It's just very different from what the headline promised.
The tokenizer is a fingerprint
This is the part that deserves attention from anyone who likes engineering — and that, with Z.ai's confirmation, became a closed case.
No lab was claiming Ox Alpha. But someone figured out its lineage with no access to any weights, using a detail that practically every API exposes: the token count on the bill.
The method is elegant. Every chat API returns usage.prompt_tokens. If you measure the count for a base string and then the count for the same base plus a probe, the chat template overhead cancels out in the subtraction. What's left is exactly how many tokens that model spent to represent your probe.
tokens(base + probe) - tokens(base) = tokens do probe naquele tokenizador
Do that with dozens of probes — different languages, code, Unicode edge cases — and you have a signature. A tokenizer is a trained vocabulary: two different labs practically never produce identical counts across the whole battery.
The result of the published survey was 95 probes against 14 candidate vocabularies, in 126 API calls, costing $0.00 — because the model was free.
| Vocabulary tested | Exact matches | Mean error |
|---|---|---|
| GLM-4.5 (Zhipu) | 84 / 95 | 1.08 |
| Best non-GLM (Llama-3.1) | 46 / 95 | 3.83 |
| GLM-5 (public vocabulary) | 95 / 95 | 0.00 |
Ninety-five out of ninety-five, with zero mean error. Add to that the video encoder behavior matching GLM-5V-Turbo and the GLM-style audio rejection, and the conclusion got hard to dodge: GLM family, probably an unannounced multimodal variant.
That's exactly what it was. The model card published by Z.ai declares the glm5_next architecture and the image-text-to-text task — multimodal, just as the tokenizer suggested.
If the name Zhipu sounds familiar, it's because the GLM line has already shown up here in another technical forensics story.
The caveats — which the method itself states
Strong evidence isn't a confession. The survey lists its own limits, and it's worth recording which of them the confirmation resolved and which are still standing:
Vocabulary identifies lineage, not operator. The tokenizer tells you which family the model descends from. It doesn't prove who is serving the API, or that the company authorized the release. Here Z.ai's confirmation closed the gap — but only in this case.
It's a snapshot, not a movie. A change of chat template or backend invalidates earlier measurements. The reading holds for the day it was taken. Still true.
The method can be defeated. All the provider would have to do is corrupt the usage numbers to scramble everything. What holds it together is the incentive: billing transparency is the one piece of credibility an anonymous provider can't give up. Still true.
The honest phrasing was "high confidence, not confirmed." Five days later, the inference became fact. That doesn't turn the method into an oracle — it turns it into evidence that held up when tested against reality, which is the most an empirical method can offer.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãWhat GLM-5.3-Flash delivers in practice
The specs are official now, and they're on the model card on OpenRouter:
- Architecture: 320B total parameters, 18B active per token. Linear attention combined with sparse.
- Context: 1,048,576 input tokens, up to 131,072 output tokens.
- Modalities: text, image and video in; text out. Native vision, with a stated focus on reading interfaces.
- Tools: function calling via
tools/tool_choiceand JSON output viaresponse_format(no schema enforcement). - Price: $0.075 per million input, $0.25 output, $0.015 cache read. 50% discount announced through September 9.
- Operations (measured while still in the stealth phase): 100% uptime and 98.94% success over 72 hours, P50 latency of 4.34s and about 27 tokens per second.
Eighteen billion active parameters out of 320B total is the explanation for the price. You pay for the compute of a small model and get the knowledge of a big one — it's the same architectural bet Qwen3.8-Flash-Next made with 6B active out of 125B total, published the same week.
Twenty-seven tokens per second is decent, not impressive. For long agentic work — the use case it declares — throughput matters less than stability and context window, and it delivers on both.
Watch the model name in the call: the old identifier stealth/ox-alpha was removed from the catalog and now returns an error. The correct one is z-ai/glm-5.3-flash:
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3-flash",
"messages": [{"role": "user", "content": "Refatore esta função e explique o porquê."}]
}'
If you've never used the platform, there's a complete guide here.
GLM-5.3-Flash's weights opened under MIT: the detail that almost slipped by
The license is the most underrated part of the announcement.
zai-org/GLM-5.3-Flash went up on Hugging Face on August 26 under MIT. It's not "open weights with an acceptable-use clause," it's not a community license with a user cap, it's not a ban on training a competing model. It's MIT: use, modify, redistribute, sell, without asking anyone's permission.
For comparison, a good chunk of the heavyweight open models come with restrictions stitched into the license. A frontier model under MIT is rare.
The community reacted at the expected pace: within a few hours of publication there were already GGUF quantizations, MLX versions in 4, 6 and 8 bits, and NVFP4 builds on Hugging Face. If you run locally, there's probably already a variant in the format you use.
What this doesn't solve: 320B total parameters are still 320B to load. The 18B active save compute per token, not memory. Aggressive quantization is the path to running this outside a serious server — and the VRAM math deserves a post of its own.
We've seen this movie before
A model with no owner, an animal name, zero price, climbing the rankings before any official announcement. It's a pattern, not an accident.
In April 2026 it was Owl Alpha: it ran two months in disguise, hit the top of OpenRouter in call volume, and only in June did Meituan own up that it was LongCat-2.0. We told the whole story here.
The mechanics are always the same, and it's good product engineering: a lab needs real traffic, with real prompts, to evaluate a model before launch. Zero price buys volume fast. Anonymity removes brand bias — nobody praises it because it's from Anthropic or trashes it because it's Chinese. And the public ranking becomes a free eval, run by thousands of devs.
Ox Alpha reached second place on OpenCode in three days, and spent six days live before becoming GLM-5.3-Flash. As a test campaign, it worked perfectly.
The real price of "free"
The window closed. Free access ran from August 20 to August 26 — six days — and ended with the model becoming a paid product and the old identifier vanishing from the catalog, with no changelog. Exactly what you'd expect from a preview.
If your code still points to stealth/ox-alpha, it's broken right now. Switch to z-ai/glm-5.3-flash or declare a fallback with the models array in the request.
And there's the lesson from the free phase, which applies to the next stealth model that shows up: you weren't paying in dollars, you were paying in prompts. All the code sent to a model with no declared owner goes to a provider that didn't identify itself, under a retention policy you didn't read, in a country you can't name. For playing around with public code, who cares. For the client's repository, it was a security decision made by omission — and only afterward did anyone learn who the prompts went to.
FAQ
Is GLM-5.3-Flash better than Claude Fable 5 for code? Based on the available data, no. It ties with GPT-5.6 Sol on the full DeepSWE (~63%). The 80% from the headline came from a sample of ten tasks and didn't hold up.
Does Ox Alpha still exist?
No. The stealth/ox-alpha identifier was removed from OpenRouter on August 26 and replaced by z-ai/glm-5.3-flash. It's the same model, with a name, a price and weights.
How much does GLM-5.3-Flash cost? $0.075 per million input tokens and $0.25 for output, with a 50% discount announced through September 9. Cache reads at $0.015.
Can you use it commercially? Now you can. The weights are under an MIT license, and the model has an identified provider with a published price. During the stealth phase the answer was no — with no identified lab, there were no terms of use, no SLA and no guarantee of continuity.
Can you run it locally? The weights are on Hugging Face and the community has already published GGUF, MLX and NVFP4 quantizations. But that's 320B total parameters to load into memory: the 18B active reduce compute per token, not the size of the model.
Why would a lab launch this way? To get eval at scale, with real traffic and no brand bias, before putting the brand's reputation on the line with an official launch.
The point
Two lessons here, and the second is worth more than the first.
The first: always ask for the sample size. Ten tasks became a headline that went around the world, and all it took was running the whole benchmark for the number to drop 17 points. Six days later, with the model named and confirmed, 63% is still the real number. If a chart doesn't say how many tasks were run, on which benchmark and with what methodology, it isn't evidence — it's decoration.
The second: the most reliable information in this episode didn't come from any announcement. It came from someone subtracting token counts from an API response, five days before any company opened its mouth. While the headline argued over whether the model was good, a $0.00 method was already answering whose it was — and got it right.
That's the difference between consuming AI news and doing AI engineering. One waits for the announcement. The other measures.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã