~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / needle2-14mb-llm-esp32-tool-calling $
News

Needle2: A 14 MB LLM That Runs on an ESP32 and Does Tool Calling with 6.5x Less Compute

LS Lucas Souza · · 10 min read
Needle2: A 14 MB LLM That Runs on an ESP32 and Does Tool Calling with 6.5x Less Compute

An LLM that runs on an ESP32 sounds like a conference joke. A cheap microcontroller, a few megabytes of RAM, and someone wants to put a language model in there deciding which function to call. Except Cactus Compute shipped Needle 2, the Show HN hit the top of Hacker News with 524 points, and the number that opens the official page is honestly absurd: 45 million parameters in a 14 MB binary.

An entire session in 28 MB of RAM. 500 tokens per second of decode on a Raspberry Pi 5. Running on Meta Quest 3S, Apple Vision Pro and Android phones under $200.

And then comes the part the announcement doesn't tell you. On three of the four function calling benchmarks, Needle 2 ties or loses to a 230M model. Typing HN into the playground fires a lock_door call. And the microcontroller demo that went viral was Needle 1, not 2 — admitted by the author himself.

Let's separate the engineering (real and unprecedented) from the marketing (which oversold it). Everything here is VERIFIED against primary sources: official page, model card, repository and HN comments.

TL;DR

  • What it is: Needle 2, a 45M-parameter foundation model specialized in on-device tool calling, from Cactus Compute.
  • Stack: single 14 MB binary (weights embedded), 28 MB RAM ceiling per session, pip install cactus-needle, CLI + Python + WASM.
  • Cost/Access: open weights, proprietary training corpus. License differs between sources (details below).
  • Useful link: github.com/cactus-compute/needle — 4,023 stars, 298 forks, 30 open issues as of August 12, 2026.

The numbers that matter in an LLM that runs on an ESP32: 14 MB, 28 MB of RAM, 70 MFLOPs per token

Start with what's measurable. Of the 45M parameters, only 35M are matmul-active — the other 8M are engram tables read via gather, with zero arithmetic cost. That drops compute per token to 70 MFLOPs, according to Cactus's official table.

Compare: LFM2.5 230M burns 460 MFLOPs/token and FunctionGemma 270M burns 540 — 6.5x and 7.7x more. A dense transformer of the same shape would burn 164 MFLOPs. In other words: the savings don't come just from being small, they come from the architecture.

The architecture is 27 layers, 512 wide, in a Simple Attention Network: Hadamard MLP (a fixed Walsh-Hadamard transform with learned diagonals in place of dense projections), engram key-value memory, multi-lane hyper-connections with dot-product routing, GQA and a 256-token sliding KV window with the tools pinned as KV sinks.

Translated into product terms: the tools never leave the context, but the conversation does. After 256 tokens, the history evaporates. A design decision, not a bug — it just changes how you design the loop.

The claimed speeds are 500 tok/s decode and 800+ prefill on the Pi 5, 400–1,500 tok/s in VR and 300–700 on cheap phones. Caveat: all self-reported. The only independent number on HN came from a user running the WASM build in the browser at 310 tok/s. Different context, no reproduction of the Pi 5 numbers.

And if you're used to thinking of local LLMs as running a 30B on an RTX 3090, this is a different planet of engineering. Here the budget is the L2 cache of an entry-level phone.

Why 2-bit doesn't collapse here: quantization trained from the start, not applied afterward

This is the real technical point, and it's where most of the coverage gets it wrong.

Quantizing a 45M model down to 2 bits post-training collapses it. A small model has little redundancy, and rounding weights to 2 bits after training destroys representation. That's why almost nobody does it.

Cactus didn't do it post-hoc. CQ2-bit (Cactus Quants) is applied from pre-training onward and kept through post-training, covering weights, activations and KV cache, with an end-to-end int8 arithmetic path. It's QAT from token zero. The model never knew any other precision — it learned inside the 2-bit grid instead of being squeezed into it later.

Product impact: this is what makes the 14 MB binary possible with no separate weights file, and therefore a direct mmap from firmware. Anyone who has fought with model loading on edge knows half the pain is file logistics, not inference.

Worth remembering that Cactus isn't new to this — we've already covered their hybrid architecture, with a local model deciding when to escalate to the cloud. Needle 2 is the most extreme end of that same thesis. If you want the contrast with the traditional route, the guide on how to run a local LLM shows the other end of the scale.

The other piece of the engineering is the byte-level grammar compiler. It's built from the tool schemas you declare and skips up to 98% of the vocabulary projection on structural tokens. The result: the output is always a syntactically valid function call. It's not the model "learning JSON" — it's the decoder being forbidden from leaving the rails.

In practice:

import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])

Field constraints (needle.Field(gt=0, le=10000), pattern=..., max_length=...) go straight into the grammar, not into validation afterward. The model is physically incapable of emitting a value outside the range.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Where it beats models 5x to 70x bigger — and where it gets crushed

Here the narrative has to be honest, because the easy headline ("better than models 70x bigger") is partially false.

All the benchmarks use strict exact match: function name, call order and every argument have to match.

Where it wins:

Benchmark Needle 2 LFM2.5 230M FunctionGemma 270M
Seal-Tools In-Domain 32.6% 26.9% 16.3%
Seal-Tools Out-of-Domain 28.7% 17.0% 15.6%

Where it loses:

Benchmark Needle 2 LFM2.5 230M Others
BFCL v4 single-turn 42.6% 60.8% Apple FM 61.7% / FunctionGemma 46.1%
Google Mobile Actions 63.7% 69.1% FunctionGemma 64.0% / Apple FM 57.6%
DroidCall 17.0% — FunctionGemma 17.5%

On BFCL v4, the largest of the four at 3,641 rows, it lands in last place. Cactus itself admits it didn't train for general function calling and that the gaps are concentrated in Java/JavaScript. The well-formed rate is 93.4%.

The honest narrative isn't "it's better." It's: it competes head-to-head in the consumer-device domain while spending 6.5x fewer FLOPs, having seen ~120x fewer training tokens (115B of pre-training + 38B of post-training, versus LFM2.5, which saw roughly 120x that total). That's still impressive. It's just not the headline.

What it does NOT do: the limits the HN comments exposed

This is where the Show HN turned into a public lab. Every failure below is in the original comments and you can reproduce them in two minutes with needle playground.

1. False positive on random input. User Tiberium typed just HN and got a lock_door call on the front door, with the reasoning "User wants to lock the door". Others reported the same with potato and hungover.

2. Inverted semantics. "Make it a little warmer" returned mode: "cool", with explicit reasoning: "'warmer' implies need for cooling". Reported by dannyw and dbeardsl.

3. Brutal sensitivity to the tool description. User hathym showed the most instructive case:

# NÃO funciona: "calculate 1 + 1?" retorna "No calculator available"
@needle.tool
def add(a: float, b: float):
    "Add two numbers."
    return a + b

# FUNCIONA: mesma função, só a descrição mudou
@needle.tool
def add(a: float, b: float):
    "Calculate the sum of two numbers. Use for any arithmetic or math question."
    return a + b

Prompt engineering didn't go away. It moved into the schema's description field.

4. Confidence isn't calibrated — and that takes down the safety mechanism. Tiberium's lock_door came back with confidence: 0. The model called the tool while saying it had no confidence at all. Another user (pylotlight) saw confidence: 0 on correct answers. Henry Ndubuaku replied that this is exactly what confidence is for, "the model knows when it's wrong" — but curious_cat_163 countered that this is the classic out-of-distribution detection problem and the team hasn't explained the method. Without calibration, the gate that's supposed to escalate to the cloud filters nothing.

5. ESP32 is a promise, not a delivered demo. When asked, Ndubuaku admitted: "That demo was Needle 1 indeed and we are creating the guide for ESP32 now as we speak." And commenter TomatoCo raised the practical bottleneck: running from QSPI flash, the common case on the ESP32-S3, makes inference far slower than the RAM numbers suggest. The one that fits the 28 MB with room to spare is the ESP32-P4 with 32 MB of PSRAM.

6. License differs between primary sources. The official site and the Hugging Face model card say Apache 2.0. The repository's LICENSE file says MIT, and the GitHub API confirms spdx_id: MIT. They're probably different artifacts (weights vs. engine), but Cactus hasn't spelled it out. Both are permissive, so the practical impact is low. The inconsistency is real.

7. The official path to production goes through a paid API. The package ships synthetic data generation and LoRA fine-tuning:

export OPENROUTER_API_KEY=sk-or-...
needle generate-data --tools my_tools.json --num-samples 500 --output data.jsonl
needle finetune data.jsonl --epochs 3 --lora-rank 16 --lora-alpha 32
needle build checkpoints/needle2.pkl --lora checkpoints/needle_lora.pkl --out my_needle.cact --bits 2

Read that again: the "14 MB offline LLM" gets good after a cycle that depends on a big model in the cloud. That's not in the announcement.

Add the top-5 retrieval on top: large catalogs go through a head that renders only the five best tools per turn. If it gets that wrong, the model never even sees the right tool — and that layer doesn't show up in any benchmark.

When this replaces an API call — and when it's just a pretty demo

The right question isn't "is the model good?" It's "which metric matters in your loop?"

The benchmarks measure average accuracy with exact match. But the product Needle 2 promises to replace is the voice → action loop on a physical actuator, and there the metric that matters is false positive rate. A door that locks itself because you mumbled "potato" isn't 63.7% accuracy — it's an incident.

As commenter planb summed it up: if the hit rate is lower than old-Siri-style heuristics, and the rest fails unpredictably, the dumb "I didn't get that" is preferable. A silent, plausible failure is worse than an explicit one when the output drives hardware.

It replaces an API call when: the domain is closed and you fine-tuned on it; the number of tools is small and the descriptions have been tested; the output goes through human confirmation or is reversible (dimming a light, pausing music, adjusting volume); latency and privacy are worth more than coverage; and you have a cloud fallback for when retrieval finds nothing.

It's just a pretty demo when: you throw a generic tool catalog at it and expect generalization; the action is irreversible or has a cost (locking a door, transferring money, opening a gate); you rely on confidence as a safety gate without having calibrated anything; or you believe 14 MB replaces a big model at open-ended tool calling — it doesn't, BFCL v4 already showed that.

What's undeniably real: 70 MFLOPs/token versus 460, and CQ2-bit applied from pre-training onward. That's a genuine engineering contribution, not hype. "Unprecedented efficiency" and "ready to replace your API call" are different claims — and only the first one is proven.

Spin up the playground, try to break it before you try to use it, and treat every tool that drives hardware as code that needs confirmation. The interesting thing about Needle 2 isn't that it solves tool calling in 14 MB. It's that it showed, in public and in real time, exactly where that frontier still is.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing