~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #deepseek $ grep

#DeepSeek

3 posts
01 #harness · #llm
DeepSeek V4.1 Flash vs Opus 5, Sol, and K3: 41x Cheaper and the Benchmark Nobody Quotes

DeepSeek released V4.1 Flash and the timeline cropped out the good row of the benchmark. We compare the model with Opus 5, GPT-5.6 Sol, Kimi K3, GLM-5.3, and GPT-6 Astra on price and performance, show where it actually leads, where it drops 20 points, and the number buried in the model card: the scaffold changes the result forty times more than swapping the model.

10 Sep · 11 min ›
02 #llm · #custos
DeepSeek V4 Pro 0813: 87.9 on Terminal Bench Beats Opus 4.8, and Your Bill Is Going Up

DeepSeek published DeepSeek-V4-Pro-0813 on its official pricing page, and the official V4 Pro release has finally dropped the preview label. The reported numbers, what's fact and what still has no public document, the comparison with Flash 0731 and Opus 4.8, and the notice on the pricing page itself that DeepSeek is going to raise prices soon.

12 Aug · 11 min ›
03 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
Meet the Clã Beer and Code
playing