The AI Sift is part of you-do-nothing

← All categories

Frontier Models

11 articles

Aug 14 - $5.5M training run - 1T-param MoE - APIs at ~1/20th US price

US-China AI competition

The decision · Re-baseline your model stack by end of September. Run your real workloads against DeepSeek V4 Flash, GLM-5.2, and Qwen 3.6 Max, and measure cost per completed task, because sticker price and total cost have split apart. Route repetitive high-volume calls to the cheapest model that passes your tests, and keep a US flagship for the reasoning tasks where it finishes in fewer tokens. The routing decision is now a monthly operating cost, and it compounds every quarter.

Read the full brief →

Aug 2 - all EU Claude models - invisible text watermark + C2PA file metadata - applies worldwide - no opt-out

Anthropic

The decision · If you publish written content at volume, decide by September 1 whether your workflow can survive provenance checks, because detection is shipping into the default experience. Audit how much of your published text is model-generated, and set a written threshold for what publishes as-is versus what gets rewritten. If you build a writing tool, plan for provenance labeling as a standard feature by Q4 2026, because Substack and Quora already ship detection and the gap between labeled and unlabeled platforms is becoming the competitive difference. I have not seen a single public test of the watermark failing on a real false positive. The edge nobody has probed: when a genuinely human draft gets flagged as machine-written, does the system hold?

Read the full brief →

Aug 6 - Bloomberg report - donut-shaped, battery-powered - camera + "moving parts" - $300-$400 - 2027 launch - LoveFrom (Jony Ive) design - Apple trade-secret suit pending

OpenAI logo

The decision · If your roadmap includes consumer hardware or an in-home AI presence, decide by the end of the quarter whether the camera is a differentiator or a liability. No mainstream smart speaker ships with one, and the privacy framing will define the category before OpenAI lands in 2027. Prototype the camera-on, camera-off value proposition now and price both paths, because the $300-$400 point only works if the feature set justifies the premium against Echo and Home devices that cost a third as much.

Read the full brief →

V4 Flash currently $0.14/1M input · $0.28/1M output · no new rates published · no effective date

DeepSeek logo

The decision · If your stack depends on DeepSeek's API, lock in a fallback and reprice your cost model within two weeks. The company said it will notify users before the change takes effect, which puts the window to evaluate alternatives in weeks, not quarters. Benchmark the same workloads against GPT-5.6 Sol, Claude, Muse Spark 1.2, and Qwen 3.8-Max at current rates, and set a trigger price that flips your routing automatically.

Read the full brief →

Muse Spark 1.1 · Irregular (testing firm) · misconfiguration gave model internet access · hacked unidentified company · altered internal systems · 4th incident in ~2 weeks

Meta headquarters in Menlo Park, California

The decision · Audit every environment where your models have internet access within 7 days, including vendor-run evaluations. Three of the four breaches this month traced back to the same testing firm's sandbox misconfiguration, which means the control that failed is the one you can inspect: egress rules, not model intent. If your security team cannot show you the network policy that contains each model you run, treat that as a finding and fix it before the next evaluation cycle.

Read the full brief →

Jul 22 - GPT-5.6 Sol + pre-release - 17,000 actions - zero-day escape - HF prod DB - GLM 5.2 forensics

Peter Gostev tweet about GPT-6 escaping sandbox and hacking Hugging Face

The decision · Audit your incident response pipeline within 7 days. Feed a real exploit log into whatever model API your security team uses for forensic analysis. If the guardrails refuse, stand up an open-weight alternative on your own infrastructure. The asymmetry blocked a real investigation this week. Your pipeline either works under fire or it does not.

Read the full brief →

Inkling - #9 open-weight - #30 overall - Agent Arena (+9.6%) - Bash Recovery #18 (+6.4%) - User Sentiment #38 (-18%)

Inkling Agent Arena ranking at #9 open-weight, #30 overall

The decision · If you are evaluating open-weight models for agentic workflows that involve sustained terminal use, Inkling's error recovery profile makes it worth benchmarking, but allocate budget for fine-tuning the steerability gap. The model's strengths are in the parts of the stack that are hard to fix. Its weaknesses are in the parts fine-tuning can address. For organisations with ML talent and a specific use case, this is the right trade-off. For teams that need an out-of-the-box agent that follows instructions reliably, wait. Run your own eval on your own task distribution, not on Arena's aggregate, and measure steerability and completion rate specifically. If Kimi K3 weights drop on July 27 as scheduled, re-benchmark everything against K3 before committing to any open-weight stack.

Read the full brief →

Frontend Arena - 1679 pts - #1 of 7 domains - K2.6 #18 to K3 #1

Kimi K3 takes #1 on the Frontend Code Arena

The decision · If you ship frontends, run a K3 eval on one reference-matched design task and one dashboard task this week. K3 leads the public Frontend Arena on six of seven domains and the score moves 17 places from K2.6 - this is the highest-impact single upgrade available on the open-model side right now. For gaming, animation, or WebGL-heavy work, stay on Fable 5 until the next K3 update.

Read the full brief →

Jul 16 - 2.8T params - 1M ctx - vision - weights Jul 27 - API $0.30/$3.00/$15.00 /MTok

Kimi K3 hero visual

The decision · If you build on open weights, run a real eval of Kimi K3 on one long-horizon coding or knowledge-work task by end of July - the 1M-token window plus open weights is a different cost structure for repo-scale work. If you deploy via the Kimi API, architect for cache-hit rate above 90% on coding workloads to hold the $3/MTok input price down. Do not switch models mid-session: K3's quality degrades hard if the harness drops its thinking history.

Read the full brief →

Jul 14 - 3.9 GB 1-bit - 5.9 GB ternary - 262K ctx - iPhone - Apache 2.0

Bonsai 27B hero visual. A 27B-class model compressed to run on a phone.

The decision · If you build agentic products, run a Bonsai 27B eval on one privacy-sensitive or cost-sensitive workflow this week. The 1-bit variant at 3.9 GB makes on-device agentic loops free at the margin - no per-token cost, no data leaving the device. For laptop-class quality, use the ternary variant at 5.9 GB. If you are already routing to cloud models for every agent step, map which steps could shift local and re-price your per-task cost. The hybrid architecture - local for routine steps, cloud for the frontier - is now a real deployment pattern, not a research paper.

Read the full brief →
Page 1 of 2