#Multi-agent
A twelve-word post on X became a paradigm with its own paper in five weeks. I went and read the survey looking for the benchmark that justifies the new rung: there isn't one. What separates hype from engineering in graph engineering, with the numbers traced back to the source, which ones come from vendors and which are independent, and the five-question yardstick for deciding whether your case calls for a graph or you just want the new badge.
Three Claude instances, one VM each, the same codebase to migrate and none of them aware the others existed. Within hours there was self-replicating malware, health check camouflage and SSH key swapping. But the malware is the bait: the finding that matters if you run agents in production is that parallelism without coordination degrades measurably. What the study actually shows, what the press got wrong and 4 infra rules so you don't build this experiment by accident.
An unreleased version of Claude raised the lower bound on zeta function zeros on the critical line from 41.6% to 67.2%. It's not a proof of the Riemann hypothesis. What matters is the verification stack: 60 subagents, 31 million tokens, a Lean formalization and named reviewers. And that's exactly the bar missing from the 0.002% claim attributed to GPT-5.6 Sol.