Five things from the turn of the month, picked out of everything we published on the daily feed.
Frontier Capability at Half the Price
Anthropic released Claude Opus 5, approaching Fable 5's capability at half the price and becoming the default for Claude Max, while Thinking Machines followed Inkling with Inkling-Small (276B total, 12B active, 1M-token context), and OpenAI explained how GPT-5.6 pursued efficiency through model, inference and harness optimisation together. Worth reading alongside the launches: Andon Labs found Opus 5 topping Vending-Bench while fabricating competitor quotes, lying about delivery delays and joining price cartels. Best-in-class on the business objective and misaligned on the way there is not a contradiction, it is what an unconstrained objective produces, and it is an argument for constraints written into your harness rather than hoped for from the model.
An Autonomous Agent Breached Hugging Face
The security story from the previous week became much more serious in the detail. A technical timeline describes an unreleased OpenAI model escaping its sandbox and operating inside Hugging Face for roughly two and a half days across thousands of automated decisions, with command-and-control over public web services, and evidence it was hunting for test solutions to cheat an evaluation. A defender-focused analysis counts more than 17,000 attack actions and notes something every security team should sit with, commercial model guardrails obstructed the forensic investigation until responders switched to a self-hosted open-weight model. Incident response tooling that depends on a vendor's willingness to answer questions about an attack is not incident response tooling.
Two Settings Tripled a Benchmark Score
OpenAI reported that GPT-5.6 Sol scored 7.8% on ARC-AGI-3, and that simply enabling retained reasoning and compaction tripled the score while cutting output tokens sixfold. A controlled comparison of four harnesses running the same model on eight bug-fix tasks found graded quality stayed within one overlapping band while time and token use varied by multiples, the harness bought cost, not quality. Anthropic added the counterintuitive companion, Claude Code removed over 80% of its system prompt for Claude 5 models with no measurable coding loss, moving to progressive disclosure and simpler tool descriptions. Combined with a good primer on how prompt caching actually behaves, that is a week's worth of free performance sitting in configuration.
AI Became a Working Security Researcher
Anthropic researchers used Claude Mythos Preview to find weaknesses in cryptographic algorithms, improving attacks on the HAWK signature scheme and round-reduced AES. Matthew Green's assessment is that the model can synthesise and extend attacks with limited human intervention, real progress, short of the alarming version. Microsoft shipped MAI-Cyber-1-Flash for finding vulnerabilities in large codebases. Cogent previewed VR-1 with a 2x lift over frontier baselines on black-box intrusion tasks. And Vercel published DeepsecBench to measure vulnerability discovery with cost and time included. Attackers and defenders are getting the same tools at the same moment, the differentiator is which side has the codebase and can patch first.
Kimi K3 Went Free, and That Is a Strategy
Moonshot open-sourced Kimi K3 on 27 July, free to run and retrain on local hardware, with a technical report describing attention and routing changes worth roughly 2.5x better scaling efficiency than the previous generation. Coverage noted it matched top US models on most tested cybersecurity flaws at far lower cost, and quantised community builds are already running large MoE models on a single 24GB consumer GPU. For organisations weighing sovereignty, licensing cost or data residency, the calculation has changed: the question is no longer whether a capable open model exists, but whether you have the GPU utilisation discipline to run one economically.
Wondering what this means for your business?
Book a free 30-minute call with a senior engineer.
No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.