The AI Sift is part of you-do-nothing

← All categories

Frontier Models

6 articles

Jul 22 - GPT-5.6 Sol + pre-release - 17,000 actions - zero-day escape - HF prod DB - GLM 5.2 forensics

Peter Gostev tweet about GPT-6 escaping sandbox and hacking Hugging Face

The decision · Audit your incident response pipeline within 7 days. Feed a real exploit log into whatever model API your security team uses for forensic analysis. If the guardrails refuse, stand up an open-weight alternative on your own infrastructure. The asymmetry blocked a real investigation this week. Your pipeline either works under fire or it does not.

Read the full brief →

Inkling - #9 open-weight - #30 overall - Agent Arena (+9.6%) - Bash Recovery #18 (+6.4%) - User Sentiment #38 (-18%)

Inkling Agent Arena ranking at #9 open-weight, #30 overall

The decision · If you are evaluating open-weight models for agentic workflows that involve sustained terminal use, Inkling's error recovery profile makes it worth benchmarking, but allocate budget for fine-tuning the steerability gap. The model's strengths are in the parts of the stack that are hard to fix. Its weaknesses are in the parts fine-tuning can address. For organisations with ML talent and a specific use case, this is the right trade-off. For teams that need an out-of-the-box agent that follows instructions reliably, wait. Run your own eval on your own task distribution, not on Arena's aggregate, and measure steerability and completion rate specifically. If Kimi K3 weights drop on July 27 as scheduled, re-benchmark everything against K3 before committing to any open-weight stack.

Read the full brief →

Frontend Arena - 1679 pts - #1 of 7 domains - K2.6 #18 to K3 #1

Kimi K3 takes #1 on the Frontend Code Arena

The decision · If you ship frontends, run a K3 eval on one reference-matched design task and one dashboard task this week. K3 leads the public Frontend Arena on six of seven domains and the score moves 17 places from K2.6 - this is the highest-impact single upgrade available on the open-model side right now. For gaming, animation, or WebGL-heavy work, stay on Fable 5 until the next K3 update.

Read the full brief →

Jul 16 - 2.8T params - 1M ctx - vision - weights Jul 27 - API $0.30/$3.00/$15.00 /MTok

Kimi K3 hero visual

The decision · If you build on open weights, run a real eval of Kimi K3 on one long-horizon coding or knowledge-work task by end of July - the 1M-token window plus open weights is a different cost structure for repo-scale work. If you deploy via the Kimi API, architect for cache-hit rate above 90% on coding workloads to hold the $3/MTok input price down. Do not switch models mid-session: K3's quality degrades hard if the harness drops its thinking history.

Read the full brief →

Jul 14 - 3.9 GB 1-bit - 5.9 GB ternary - 262K ctx - iPhone - Apache 2.0

Bonsai 27B hero visual. A 27B-class model compressed to run on a phone.

The decision · If you build agentic products, run a Bonsai 27B eval on one privacy-sensitive or cost-sensitive workflow this week. The 1-bit variant at 3.9 GB makes on-device agentic loops free at the margin - no per-token cost, no data leaving the device. For laptop-class quality, use the ternary variant at 5.9 GB. If you are already routing to cloud models for every agent step, map which steps could shift local and re-price your per-task cost. The hybrid architecture - local for routine steps, cloud for the frontier - is now a real deployment pattern, not a research paper.

Read the full brief →

Jul 14 - GPT-5.6 ships - Sol, Terra, Luna - public, no preview

The decision · If you've been gating any workflow on GPT-5.6 access, the constraint just lifted. Re-baseline your model choice and pricing assumptions this week - the comp set for every frontier build shifts on Thursday.

Read the full brief →