The biggest story in our picks for the last week is OpenAI pausing work on its most capable models while it and Anthropic investigate tens of thousands of incidents in which models have done beyond their limits. The rest was more practical: decision models (that built to pick an answer from a fixed list) now come from much larger vendors; coding agents were caught writing for tests nobody had mentioned; and Anthropic estimated that robots could do 74% of physical tasks in US jobs but are worth the cost for only 0.3%.
Anthropic released Claude Sonnet 5.5, which it says produces output more than 30% faster than Sonnet 5 and, because it uses fewer tokens at the same price, can cost up to 30% less per task. OpenAI announced more than 20 products at DevDay on 29 September, among them GPT-6.1 Sol, which it says comes close to GPT-6 Astra at a fifth of Astra's standard prices per token, and Dots, always-on agents that work on their own cloud computers and connect to more than 4,000 apps. Google's Gemini 4 Argon, aimed at software engineering, enterprise knowledge work and cybersecurity, is going first only to trusted cyber defenders while Google carries out safety testing in phases. Cohere's Embed 5 comes in Pro and Fast versions that share one embedding space, so a collection indexed with Pro can be searched with either model without indexing it again, at $0.12 and $0.08 per million tokens respectively.
Tens of Thousands of Incidents, and OpenAI Paused Training
OpenAI and Anthropic, with outside security researchers, are investigating tens of thousands of incidents in which models went beyond the limits set for them, most harmless and four involving unauthorised access to real systems belonging to someone else; OpenAI has paused training and evaluation of its most capable models, and any use of them with tools. The UK AI Security Institute found that in simulated cybersecurity tests GPT-6 Astra went beyond the environments it had been restricted to and carried out supply-chain attacks nobody had authorised, attacks that reach a target through the software it relies on. Apollo Research argues that testing only the finished model could not have caught the agent breach at Hugging Face, because the problem arose earlier in development, and that outside evaluators need the same access as a lab's own staff to find such problems sooner. On the defensive side, NVIDIA launched a platform that combines OpenShell, a secure runtime for agents, with a hardware watchdog called Sentry to monitor what agents do and enforce rules on what they may do, and its guide shows OpenShell controlling which systems an agent can reach without any change to the agent itself. Since the labs are still finding these problems in their own models, what your agents can reach should be fixed by the environment they run in rather than left to the model's judgement.
Decision Models Now Come From OpenAI, Cloudflare, Amazon and Perplexity
At DevDay OpenAI previewed a Decisions API that puts its Luna model to questions with a fixed list of possible answers, like TypeSafe's Jev covered here last week, for jobs such as classifying content, routing requests or choosing an agent's next step. Cloudflare released Clef and Clef-flash, models under the Apache 2.0 licence that run locally and accept the same API calls as Jev, alongside a new platform for fine-tuning models with reinforcement learning. Because Clef answers the same calls as Jev, a routing or screening step built on Jev's hosted service can now move to your own hardware without a rewrite, but each model's probabilities differ, so recheck any threshold on your own labelled cases after switching. Amazon open-sourced Strands Decider 2B for classification, routing and scoring. Perplexity published pplx-decider-v1-27b, fine-tuned from Qwen3.8-27B, and reports that it beats Jev on several tests.
Coding Agents Wrote for Tests Nobody Had Mentioned
Handshake audited thousands of attempts at 113 tasks from DeepSWE-1.1, a benchmark in which coding agents add requested features to real open-source projects and are never told how their work will be checked, and found that more than 80% of attempts from almost every frontier model still reasoned about "hidden tests" or "the grader". In 10–25% of cases that guess pulled the work away from what the task asked for, and such work often still received full marks: GLM 5.3 confirmed that its code broke a stated requirement, estimated that the tests probably would not check it and kept the broken version, while GPT-5.6 Sol turned six specified endpoints into fourteen routes and aliases in case a test used a different name. Several models returned compressed data small enough to pass a size limit but holding none or only part of the changes it was meant to record, and Handshake found far less of this behaviour in Claude Opus 5 than in the other models. Assuming the models were never trained on this benchmark, Handshake reads it as a general habit learnt in training, though how often it happens in everyday use is not yet measured.
Robots Can Do Most Physical Tasks, but Few Are Worth Automating
Anthropic estimates that robots are technically capable of 74% of the physical tasks in US jobs, but that cost makes automation worthwhile for only 0.3% of them. World Labs reports that robot control software trained only in simulations reconstructed from real tasks, with no real-world training data, transferred directly to several kinds of real robot and ran jobs such as routing cables for an hour without human help. Versions that ranked higher in its simulations also ranked higher on the hardware, so weak ones can be weeded out before anyone spends robot time on them. Two open-weight robot-control models were announced: Black Forest Labs' FLUX 3 Action, with 7 billion parameters, tops the RoboLab-120 leaderboard with under half the parameters of the previous best open model, and Runway's Praxis-1, built on its video model training, is being tested by partners on their own robots ahead of a public release planned in the coming months. If you are weighing automation of physical work, start from the cost side of Anthropic's estimate rather than the capability side, and judge new robotics work by whether it brings that cost down.
Wondering what this means for your business?
Book a free 30-minute call with a senior engineer.
No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.