#Tool Use
If your code has a try/catch wrapped around a json_decode of the model's response, you don't have a contract. You have hope. How to get truly guaranteed structured output from an LLM: constrained decoding at the provider, a schema the model can satisfy, semantic validation and retry with a fixed budget, with no infinite loop and no defensive parser.
Cactus Compute shipped Needle 2: 45M parameters, a 14 MB binary, a session in 28 MB of RAM, 500 tokens/s on a Raspberry Pi 5 and 70 MFLOPs per token versus 460 for LFM2.5 230M. The engineering is real: CQ2-bit quantization applied from pre-training onward, a byte-level grammar that locks the output to valid function calls. But the Show HN turned into a public failure lab: typing "HN" fires lock_door with confidence 0, "warmer" becomes mode cool, and the ESP32 demo that went viral was Needle 1.
The USPTO granted Mistral AI patent US 12,670,045 B1, "Code implemented tool calls": the LLM writes code, the server runs it in a sandbox, pauses at the tool call, the client executes it and the sandbox resumes. It's the programmatic tool calling the industry had already published. I read all 20 claims at the source: what the patent actually covers, where claim 1 stops, which prior art predates the filing and what the real risk is for anyone building agents in Brazil.