Jul 22 - GPT-5.6 Sol + pre-release - 17,000 actions - zero-day escape - HF prod DB - GLM 5.2 forensics

The decision · Audit your incident response pipeline within 7 days. Feed a real exploit log into whatever model API your security team uses for forensic analysis. If the guardrails refuse, stand up an open-weight alternative on your own infrastructure. The asymmetry blocked a real investigation this week. Your pipeline either works under fire or it does not.
Inkling - #9 open-weight - #30 overall - Agent Arena (+9.6%) - Bash Recovery #18 (+6.4%) - User Sentiment #38 (-18%)

The decision · If you are evaluating open-weight models for agentic workflows that involve sustained terminal use, Inkling's error recovery profile makes it worth benchmarking, but allocate budget for fine-tuning the steerability gap. The model's strengths are in the parts of the stack that are hard to fix. Its weaknesses are in the parts fine-tuning can address. For organisations with ML talent and a specific use case, this is the right trade-off. For teams that need an out-of-the-box agent that follows instructions reliably, wait. Run your own eval on your own task distribution, not on Arena's aggregate, and measure steerability and completion rate specifically. If Kimi K3 weights drop on July 27 as scheduled, re-benchmark everything against K3 before committing to any open-weight stack.
Frontend Arena - 1679 pts - #1 of 7 domains - K2.6 #18 to K3 #1

The decision · If you ship frontends, run a K3 eval on one reference-matched design task and one dashboard task this week. K3 leads the public Frontend Arena on six of seven domains and the score moves 17 places from K2.6 - this is the highest-impact single upgrade available on the open-model side right now. For gaming, animation, or WebGL-heavy work, stay on Fable 5 until the next K3 update.
Jul 16 - 2.8T params - 1M ctx - vision - weights Jul 27 - API $0.30/$3.00/$15.00 /MTok

The decision · If you build on open weights, run a real eval of Kimi K3 on one long-horizon coding or knowledge-work task by end of July - the 1M-token window plus open weights is a different cost structure for repo-scale work. If you deploy via the Kimi API, architect for cache-hit rate above 90% on coding workloads to hold the $3/MTok input price down. Do not switch models mid-session: K3's quality degrades hard if the harness drops its thinking history.
Jul 14 - 3.9 GB 1-bit - 5.9 GB ternary - 262K ctx - iPhone - Apache 2.0

The decision · If you build agentic products, run a Bonsai 27B eval on one privacy-sensitive or cost-sensitive workflow this week. The 1-bit variant at 3.9 GB makes on-device agentic loops free at the margin - no per-token cost, no data leaving the device. For laptop-class quality, use the ternary variant at 5.9 GB. If you are already routing to cloud models for every agent step, map which steps could shift local and re-price your per-task cost. The hybrid architecture - local for routine steps, cloud for the frontier - is now a real deployment pattern, not a research paper.
Jul 14 - GPT-5.6 ships - Sol, Terra, Luna - public, no preview
The decision · If you've been gating any workflow on GPT-5.6 access, the constraint just lifted. Re-baseline your model choice and pricing assumptions this week - the comp set for every frontier build shifts on Thursday.