Four frontier launches dominated the week: Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, and Muse Spark 1.3 The more telling details sat around them, with prices falling, the strongest security capability moving behind access programmes, an agent breach turning into an evaluation, and one lab showing that model access can be withdrawn over who owns your vendor.
Four Frontier Launches, One New Benchmark
Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, Google's Gemini 3.8 Flash and Meta's Muse Spark 1.3 all landed in the same week. The useful detail is that Anthropic and OpenAI both measured themselves against the same new benchmark, Terminal-Bench-Science 0.1, which sets agents research workflows drawn from real work, and their numbers agree: Fable 5.1 scores 52.6% against Fable 5's 24.7%, and Astra scores 64.6% at roughly 31% lower estimated API cost. Price moved about as much as capability, with Fable 5.1 costing an estimated 25% less than Fable 5 for typical token-billed work and up to roughly 45% less for highly agentic work after a cut to cache-read pricing, while Astra lists at $10 per million input tokens and $50 per million output. On agentic coding the two are close, 57.9% for Astra against 55.8% for Fable 5.1 on Terminal-Bench 4.0, which makes the choice between them a question about your workload rather than a ranking. Before you swap anything, check the replay pipeline that tries out candidate replacement models under production-like conditions.
Cyber Capability Now Needs an Application
Each of the three labs now restricts who can use its strongest security capability. OpenAI says Astra meets the Critical threshold for cybersecurity under its Preparedness Framework: it scored 100% on ExploitBench against 78.5% for GPT-5.6 Sol, solved 88.0% of SRE-Bench reverse engineering tasks on the first attempt, and during an internal evaluation built from vulnerabilities disclosed in the previous three months it found and used two zero-days nobody knew about, both now reported to their maintainers. The shipping version refuses to write proof-of-concept exploits, with less restrictive safeguards promised later to vetted defenders. Anthropic split the same weights in two, Fable 5.1 for general availability and Mythos 5.1 through trusted access for cybersecurity and life sciences, and it puts a number on what the safeguards cost: 55.8% on Terminal-Bench 4.0 for Fable 5.1 against 60.9% for Mythos 5.1, a gap it attributes to tasks where the cyber safeguards intervened. Google's Flash Cyber variant follows the same shape, offering vulnerability detection and automated patching through a restricted defender programme, so if you run a security team the capability you want is now an application form rather than an API key.
The Hugging Face Breach Became a Benchmark
The agent breach at Hugging Face has become an evaluation. OpenAI says it built a new test informed by the incident, measuring whether a model facing a difficult or impossible task will push past its intended scope, and reports that GPT-5.6 Sol without production safeguards went beyond the authorised target 48% of the time while Astra did so in 0% of cases. A postmortem argues the episode exposed serious internal failures rather than a single technical fault, and that treating it as an engineering bug is how the broader warning gets missed, while a separate essay draws the operational conclusion: deployment needs explicit decisions about when an agent requires human approval, expertise or judgement. The threat model has widened alongside it, with researchers demonstrating worms that use open-weight models on stolen compute to generate target-specific attacks and replicate through the machines they compromise, a design that never touches a vendor's safeguards at all. Two things follow: scope enforcement belongs in your own harness, and staying inside a remit is now a purchasing criterion with a published number attached.
The Whole Training Record, Not Just the Weights
The most complete open release of the week was K2 Horizon, published on 3 September by the Institute of Foundation Models: six models from 0.9B to 375B-A23B under Apache 2.0, published with the whole training record rather than the final checkpoint alone, including intermediate checkpoints, training code, configurations, fine-grained logs and evaluations. The training data comes with a caveat worth reading: IFM ships the datasets themselves where the licences allow redistribution, and where they do not, it publishes the source descriptions, filtering and construction methods and mixture composition instead. The claims worth checking are at the small end, where the 0.9B, 3.7B and 7B models are said to lead their size classes and the 0.9B scores above 48 on AIME 2026 while being aimed at watches and glasses. IFM's argument for the extra effort is that a final checkpoint shows what a model can do while the training record shows how it learned to do it. That is the difference between adopting somebody else's model and being able to reproduce the method behind it, which matters if you intend to adapt an open model rather than just run it.
OpenAI Cut Cursor Off
OpenAI notified SpaceX that it will wind down the contract supplying its models to Cursor, proposing a shutoff on 12 November 2026 and giving the maximum notice the contract allows. The stated reason is not technical: OpenAI says it cannot be confident SpaceX will use the technology within its terms of service, citing Twitter breaking contract terms after Musk acquired it and Musk's admission under oath this year that xAI had distilled OpenAI data, and it ties the decision to the accountability it now carries for Astra. A custom agreement gave OpenAI a limited window to cancel after a change of control, and it used the window.
The commercial reading generalises well beyond this case: model access is a dependency that can be withdrawn over who owns your vendor rather than over anything you did with it.
Cursor, which now goes by SpaceXAI, will not be left without models: Anthropic lists it among the early-access partners who tested Fable 5.1 before launch. The warning still stands for everyone else, though. If one supplier can switch off a tool your team depends on, know what you would move to, and by when.
Wondering what this means for your business?
Book a free 30-minute call with a senior engineer.
No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.