~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #hardware $ grep

#Hardware

4 posts
01 #performance · #openai
GPT-5.6 Sol Ultrafast: 750 tokens/s on Cerebras, 11x faster than Fable 5

It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.

15 Aug · 10 min ›
02 #ai-agents · #noticias
Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training

Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.

15 Aug · 11 min ›
03 #mobile · #tool-use
Needle2: A 14 MB LLM That Runs on an ESP32 and Does Tool Calling with 6.5x Less Compute

Cactus Compute shipped Needle 2: 45M parameters, a 14 MB binary, a session in 28 MB of RAM, 500 tokens/s on a Raspberry Pi 5 and 70 MFLOPs per token versus 460 for LFM2.5 230M. The engineering is real: CQ2-bit quantization applied from pre-training onward, a byte-level grammar that locks the output to valid function calls. But the Show HN turned into a public failure lab: typing "HN" fires lock_door with confidence 0, "warmer" becomes mode cool, and the ESP32 demo that went viral was Needle 1.

12 Aug · 10 min ›
04 #ai-agents · #llm
Does Muse Glimmer 30B Really Fit on an RTX 3090? Meta Says One Thing, the People Who Tested It Say Another

The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.

11 Aug · 12 min ›
Meet the Clã Beer and Code
playing