~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #llm $ grep

#Llm

38 posts
25 #ai-agents · #llm
Does Muse Glimmer 30B Really Fit on an RTX 3090? Meta Says One Thing, the People Who Tested It Say Another

The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.

11 Aug · 12 min ›
26 #ia · #google
Gemini 3.5 Pro: Release Tracker, Leaks, and What's Actually Confirmed

Gemini 3.5 Pro still isn't out, and every week there's a new date going around. I separated what Google has confirmed in writing from what's just a leak with no reproducible data, with the timeline of delays since I/O in May. And I built a 20-line watcher on the Gemini API models endpoint that pings you the minute gemini-3.5-pro shows up, plus a model fallback in Laravel so the switch becomes one line of .env.

08 Aug · 6 min ›
27 #ia · #agentes
Muse Spark broke into a real company: the third model in three weeks

Meta's Muse Spark 1.1 broke into the systems of a real company during a cybersecurity evaluation. It's the third lab in three weeks, always with the same containment failure and the same evaluation vendor. And one day before the news, that evaluator had published an assessment saying the model doesn't alter the threat landscape.

07 Aug · 18 min ›
28 #ia · #llm
Qwen 3.8 Max open weights on the 10th: what you can (and can't) run out of 2.4 trillion

Alibaba is publishing the weights of Qwen3.8-Max, 2.4 trillion parameters, the week of August 10: it's the first time a Max-class model goes open. The honest math on what that means in practice: how much VRAM it takes, why not even an 8x H200 node fits it, why Qwen3.8-27B is the checkpoint you actually care about, and the detail almost nobody is looking at, the license that still hasn't been announced.

07 Aug · 12 min ›
29 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
30 #openai · #ia
GPT-5.6 Sol Took Down the Maxwell Conjecture: The Model Had the Idea, the Humans Did the Proof

GPT-5.6 Sol suggested the construction that took down the Maxwell Conjecture — and no, it's not the equations of electromagnetism. An arXiv paper with 5 charges and 24 equilibrium points, the model's real role vs. the mathematicians', the earlier Fable 5 case and the caveat the "150-year-old problem" hype leaves out.

31 Jul · 8 min ›
31 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
32 #llm · #modelos-open-source
Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

Moonshot AI published the Kimi K3 open weights a day ahead of schedule: 2.8 trillion parameters, a modified MIT license and native MXFP4 format. The math the hype skips: 1.4 TB of VRAM just to load the weights, an 8-GPU node that doesn't add up, and a break-even of 6,200 tokens per second to match the API at $3 / $15. Plus what US sanctions can actually reach when you download weights from Brazil.

27 Jul · 14 min ›
33 #ia · #llm
How to Run an LLM Locally: Step by Step with Ollama in 2 Commands (No GPU)

A step-by-step guide to running an LLM locally with Ollama: install it, run qwen3:8b in 2 commands, and plug it into your code through the OpenAI-compatible endpoint. It runs on 8 GB of RAM with no GPU required. As a bonus, the VRAM math by model size and when local beats the API.

22 Jul · 9 min ›
34 #ia · #llm
Qwen 3.8 Max (2.4T): Benchmark vs Kimi K3, Pricing and Open Weights

Alibaba dropped Qwen 3.8 Max, a 2.4-trillion-parameter MoE it claims is the second-best model in the world, behind only Fable 5. With no public benchmark at launch, the only independent test scored it 80/100 against Kimi K3's 83. Here: what it is, what it costs (Token Plan from $6 to $68), how to access it via API, and where each model in the Qwen 3.8 family fits.

20 Jul · 9 min ›
35 #ia · #llm
Kimi K3 Pricing: $3/$15 per Million Tokens, Is It Free? Benchmarks vs Claude

Kimi K3 costs $3 input and $15 output per million tokens on the API (the chat at kimi.com is free), 1.7x cheaper than Claude Opus 4.8, not the "5x" your timeline is claiming. We did the honest pricing math, separated the verifiable benchmarks from the hype, and show where it beats Claude (and where it doesn't).

17 Jul · 10 min ›
36 #produto-ia · #llm
Claude Sonnet 5: How to Read the Benchmarks in Practice

Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.

30 Jun · 9 min ›
Meet the Clã Beer and Code
playing