~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #custos $ grep

#Custos

20 posts
13 #ia · #rag
Fine-Tuning, RAG, or Prompt: Which One Your Problem Actually Calls For

The decision most people get wrong by overshooting. Objective criteria for choosing between prompt, RAG, and fine tuning, what each path costs with real numbers, the three rare cases where training a model wins, and the five-minute test that separates a knowledge problem from a retrieval problem. Plus the news that changes the math: OpenAI is shutting down its fine-tuning platform.

31 Aug · 14 min ›
14 #performance · #openai
GPT-5.6 Sol Ultrafast: 750 tokens/s on Cerebras, 11x faster than Fable 5

It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.

15 Aug · 10 min ›
15 #google · #ai-agents
Gemini 3.7 Flash Is Here: 43.6% on FrontierCode and $0.75/M, with Flash Ahead of Pro Again

Google ran the same play and shipped Flash before Pro. Except the $0.75/M that took over the timeline isn't a low price: the official table shows it's the 3.6 Flash price with a 50% discount through December 31, 2026, and the bill doubles on January 1. Here are both numbers, the real benchmarks (FrontierCode 43.6%, AutomationBench 30.4%), what breaks when you migrate from 3.6, and the data point the release leaves out: hallucination went up from 55.6% to 64.5%.

15 Aug · 11 min ›
16 #ai-agents · #noticias
Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training

Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.

15 Aug · 11 min ›
17 #ai-agents · #produtividade
Cursor SpaceX: What the Acquisition Changes in Your Editor (and What Nobody Confirmed)

Cursor confirmed on August 14, 2026 that it has been acquired by SpaceX. The announcement runs thirteen sentences: it talks about GPUs, cheaper models, and the horizon, and says nothing about your code. We separate what's in a primary source from what's press-only (including the $60 billion), show what actually changes in the editor (Grok 4.6 in the house pool, Claude hidden by default, the Router choosing for you), and close with a checklist for anyone who depends on Cursor in production.

15 Aug · 11 min ›
18 #llm · #custos
DeepSeek V4 Pro 0813: 87.9 on Terminal Bench Beats Opus 4.8, and Your Bill Is Going Up

DeepSeek published DeepSeek-V4-Pro-0813 on its official pricing page, and the official V4 Pro release has finally dropped the preview label. The reported numbers, what's fact and what still has no public document, the comparison with Flash 0731 and Opus 4.8, and the notice on the pricing page itself that DeepSeek is going to raise prices soon.

12 Aug · 11 min ›
19 #cache · #contexto
How to Save Tokens in Claude Code, Codex, and Cursor: Stop Breaking the Cache

You send a one-line question and /usage reports a whole day's worth of consumption. Saving tokens in a coding assistant has nothing to do with prompt size: it's about prefix caching. How it works in Claude Code, Codex, and Cursor, the seven actions that invalidate it without you noticing, how to measure it with cache_read vs cache_creation, and eight levers to stretch the session.

06 Aug · 16 min ›
20 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
Meet the Clã Beer and Code
playing