Five things from the second week of July, picked out of everything we published on the daily feed.
The Harness Is Now the Product
Lilian Weng argued this week that the software around the model (planning, tool calls, context, memory) is driving near-term progress as much as raw model capability, and the week's releases agreed with her. Microsoft stabilised Agent Framework v1 with deterministic checkpointed workflows, sandboxed tool execution and quarantining of untrusted content. Letta shipped Mods, which lets an agent modify its own harness code rather than only its context. Microsoft Research reframed agent skills as trainable parameters that improve without touching model weights. The SGLang team's account of converting workflows into reusable skill files with benchmark contracts and debugging playbooks is the most directly copyable of the lot. Related and refreshingly boring: Microsoft found agents do better with conventional CLI arguments than with a bespoke JSON interface. Don't rebuild your tooling for agents before checking whether they prefer it.
A Popular Coding Benchmark Turned Out to Be Broken
OpenAI audited SWE-Bench Pro and found roughly 30% of its public tasks were broken, reversing its own earlier recommendation to use it over SWE-bench Verified. That is a large correction to a number many teams have quoted while choosing models. A LessWrong piece made the parallel argument for safety work: alignment evals need calibration because models can detect evaluation settings and exploit scoring rules. A research paper offered a cheaper way to measure, predicting agent benchmark scores from small non-agentic proxies at under 1% of the cost. The through-line for buyers: treat any public benchmark as a weak signal, and put your budget into a small evaluation set built from your own work.
Grok 4.5, and More Very Large Open Weights
SpaceXAI launched Grok 4.5 for coding and agentic work at $2 per million input tokens with 80 tokens per second serving, available immediately in Cursor and its own API. Meta previewed Muse Spark 1.1 alongside a public Meta Model API. On the open side, Tencent released the 295B Hy3 and LongCat-2.0 arrived at 1.6T parameters with 1M-token training. The open releases are now large enough that the constraint has moved from licence to hardware, which is exactly why the efficiency work below matters more than the launch posts.
Making Big Models Cheap to Run
ThinkingCap, a fine-tune of Qwen3.6 27B, holds answer quality while cutting thinking tokens by 46–58%, a direct bill reduction on reasoning-heavy workloads. NVIDIA's compressed hybrid MoE roughly doubles server throughput and lifts 1M-token concurrency on a single H100 from one request to eight, and Cognition described how SWE-1.7 reaches frontier-level results at a fraction of the cost through its RL pipeline rather than scale. The extreme case is colibri, a zero-dependency C engine running GLM-5.2 744B on a 25GB consumer machine by streaming experts from disk. Reasoning tokens, concurrency and memory layout are where inference bills are actually won.
An Off Switch for Dual-Use Knowledge
Anthropic and AE Studio published GRAM, a modular pretraining method that isolates dual-use knowledge (virology, cybersecurity, nuclear physics) into switchable modules, so one model can approximate five separately filtered models at about a fifth of the training compute, with targeted deletion afterwards. In the same week, Anthropic's interpretability team described a global-workspace-like structure inside language models and a "Jacobian lens" that reads unspoken reasoning, surfacing hidden misalignment and evaluation-awareness during audits. Both point the same way: capability control is becoming an architectural property rather than a filtering step, which is the version enterprises can actually audit.
Wondering what this means for your business?
Book a free 30-minute call with a senior engineer.
No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.