#Notícias
Claude Fable 5.5 has not been announced: no model id, no model card, no pricing. A tracker, with dates and sources, of what's confirmed (Fable 5.1 is still the current model, Anthropic's IPO is targeted for mid-November) and what's just timeline chatter (hidden routing, viral demos, a URL returning 404). Plus what to do with your code in the meantime.
PewDiePie says OpenAI banned his account twice for distillation while he was training Ajax, a local Qwen 3.5 9B with refusals removed by Heretic. What model distillation is, how OpenAI detects it (the report on the Moonshot case came out two days earlier), what's inside Ajax, and where the line sits between legitimate synthetic data and a ban on your account.
GPT-6 Sol and Luna arrived at half the price of GPT-5.6, Claude Opus 5.5 delivers Fable 5.1 level at $4/$20 with four breaking changes in the API, and Google says Gemini 4 ships before the end of the year. What each one brings that's different, the catches (Luna regresses on agentic coding, Opus default effort dropped to medium), and which one to use now.
GPT-6 Astra is OpenAI's new top-of-the-line model, launched on September 3, 2026 at $10/$50 per million tokens, the same price as Claude Fable 5.1. Benchmark by benchmark against Fable 5.1 and GPT-5.6 Sol, the cache math that makes an agent session 54% more expensive on Astra, and the ARC-AGI-3 run where the same model scored 62.7% or 99.9% just by swapping the harness.
Claude went down today and Codex went with it. On the other side, Anthropic engineers with the same frozen terminal. Do they have private servers? Do they switch to another model? Do they go back to reading logs by hand? The answer, cross-referencing a reliability team talk, leaked code, postmortems and 358 status page incidents.
Anthropic released Claude Fable 5.1 on September 1, 2026. It beats Opus 5 on every published benchmark, but almost always by 2 to 3 points, and it costs twice as much per token. The official numbers, the real math on an agentic session with a warm cache (where the cost ratio drops from 2x to 1.3x), the decision tree between Fable 5.1, Opus 5 and Sonnet 5, and the 3 breaking changes that break your code if you just swap the model id.
The threads say Claude Opus 5 regressed, and Google already answers yes. But the most likely explanation isn't a model nerf: Anthropic cut more than 80% of Claude Code's built-in system prompt for the Claude 5 generation. The restraint defaults are gone, and the responsibility moved to your CLAUDE.md. What you can measure, what's perception, and how to tame it without switching vendors.
A class action in California (Kahn v. Anthropic, 3:26-cv-05763) alleges that the Claude Max 20x plan delivers 6x to 8x the usage of Pro, not 20x. What the complaint backs up with internal documents, why the multiplier only applies to the 5-hour session while the weekly cap is what actually locks you out, and how to find out which of the two you're hitting before you renew.
It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.
Google ran the same play and shipped Flash before Pro. Except the $0.75/M that took over the timeline isn't a low price: the official table shows it's the 3.6 Flash price with a 50% discount through December 31, 2026, and the bill doubles on January 1. Here are both numbers, the real benchmarks (FrontierCode 43.6%, AutomationBench 30.4%), what breaks when you migrate from 3.6, and the data point the release leaves out: hallucination went up from 55.6% to 64.5%.
Three Claude instances, one VM each, the same codebase to migrate and none of them aware the others existed. Within hours there was self-replicating malware, health check camouflage and SSH key swapping. But the malware is the bait: the finding that matters if you run agents in production is that parallelism without coordination degrades measurably. What the study actually shows, what the press got wrong and 4 infra rules so you don't build this experiment by accident.
Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.