News
Releases, news, and trends from the world of software development.
In July, Bloomberg reported Altman's briefing to the US government and the internet read it as an imminent GPT-6 launch. This tracker logged, with a date and a source, what was confirmed and what was speculation until the model shipped on September 3, 2026 as GPT-6 Astra. It stays up as a timeline; the numbers, the pricing and the comparison with Fable 5.1 and GPT-5.6 Sol are in the comparison post.
Fable 5 helped mathematician Levent Alpöge produce a counterexample to the Jacobian Conjecture, open since 1939. What the conjecture is without the jargon, what a counterexample is, why checking it is easy this time, and what this says about AI doing frontier mathematics.
Alibaba dropped Qwen 3.8 Max, a 2.4-trillion-parameter MoE it claims is the second-best model in the world, behind only Fable 5. With no public benchmark at launch, the only independent test scored it 80/100 against Kimi K3's 83. Here: what it is, what it costs (Token Plan from $6 to $68), how to access it via API, and where each model in the Qwen 3.8 family fits.
Kimi K3 costs $3 input and $15 output per million tokens on the API (the chat at kimi.com is free), 1.7x cheaper than Claude Opus 4.8, not the "5x" your timeline is claiming. We did the honest pricing math, separated the verifiable benchmarks from the hype, and show where it beats Claude (and where it doesn't).
GPT-5.6 has been in Codex since July 9, across all three tiers. Ultra mode coordinates four subagents that cooperate during the task and buys 3 points on Terminal-Bench (88.8% to 91.9%), at a much higher token cost. And CodeRabbit's independent test kills the reflex of picking the cheap tier: Terra burned 2.6x more tokens than Sol to pass 40.7% of tasks versus 63.7%. Price per token is not price per task.
OpenAI removed the Codex 5-hour limit without warning and Anthropic answered within hours, extending the 50% bump to Claude Code's weekly limit through July 19. What exactly changed, what is still capped, and how to decide which subscription is worth it right now, with the numbers in hand.
Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.
What's the difference between GPT-5.6 Sol, Terra and Luna? Same generation, three tiers: Sol is the agentic top end ($5/$30 per 1M tokens), Terra is the middle ground ($2/$12) and Luna is the cheap, fast one ($0.20/$1.20). With the pricing in effect since the July 30 cut, here's when each one is worth it and when Sol is just wasted budget.
Claude Code or Codex? The answer comes with numbers: Codex opens a 13-point lead on Terminal-Bench (82.7% vs 69.4%) and Claude Code opens a 10-point lead on SWE-bench Pro (69.2% vs 58.6%), which is the benchmark for real multi-file problems. On SWE-bench Verified they tie. Here is the verdict by scenario, the real cost per dev, and the criterion that matters more than quality: how much control you want during the task.
Fable 5 or Opus 4.8, which one should you use? A straight verdict by task type, with the 10 concrete situations where Fable gets it done and Opus 4.8 did it badly or not at all: migration at scale, code from a screenshot, long-running agents and reasoning over documents. Plus what each one costs and the cases where Opus 4.8 is still the right call.
There are three markets running in parallel for AI Engineers in Brazil, and each one has its own range, tax rate and negotiation criteria: CLT (salaried employment), PJ (contractor) and foreign contracts through an EOR. Here are the three 2026 ranges by level, with sources, plus the 1.5x ratio between CLT and PJ, the all-in cost a foreign employer actually pays for you, and the three specializations that move the number more than stack or tenure.