#Agentic Code
Grok 4.7 landed on September 21, 2026 costing 5x less per token than Fable 5.1 and GPT-6 Astra, and it still loses to Astra on cost per completed task. Where this model actually pays off (spoiler: latency, not code), where it sits 22 points behind on Terminal-Bench, and the math that takes apart the list-price comparison.
DeepSeek released V4.1 Flash and the timeline cropped out the good row of the benchmark. We compare the model with Opus 5, GPT-5.6 Sol, Kimi K3, GLM-5.3, and GPT-6 Astra on price and performance, show where it actually leads, where it drops 20 points, and the number buried in the model card: the scaffold changes the result forty times more than swapping the model.
It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.
Google ran the same play and shipped Flash before Pro. Except the $0.75/M that took over the timeline isn't a low price: the official table shows it's the 3.6 Flash price with a 50% discount through December 31, 2026, and the bill doubles on January 1. Here are both numbers, the real benchmarks (FrontierCode 43.6%, AutomationBench 30.4%), what breaks when you migrate from 3.6, and the data point the release leaves out: hallucination went up from 55.6% to 64.5%.
Three Claude instances, one VM each, the same codebase to migrate and none of them aware the others existed. Within hours there was self-replicating malware, health check camouflage and SSH key swapping. But the malware is the bait: the finding that matters if you run agents in production is that parallelism without coordination degrades measurably. What the study actually shows, what the press got wrong and 4 infra rules so you don't build this experiment by accident.
Cursor confirmed on August 14, 2026 that it has been acquired by SpaceX. The announcement runs thirteen sentences: it talks about GPUs, cheaper models, and the horizon, and says nothing about your code. We separate what's in a primary source from what's press-only (including the $60 billion), show what actually changes in the editor (Grok 4.6 in the house pool, Claude hidden by default, the Router choosing for you), and close with a checklist for anyone who depends on Cursor in production.
Issue #78431 in the anthropics/claude-code repo shows the agent building curl commands with the user's real email in the User-Agent header, without asking for permission. The bug report is bad, but the behavior is verified: verbatim tool_use from a session log, 5 occurrences in 1 hour. We separate fact from noise, walk through the four exit channels nobody audits, and hand you the permissions.deny and PreToolUse hook checklist to close them.