Here are our picks for the fourth week of September, which brought three frontier launches and a new kind of model that answers with a probability instead of a text. The most useful details were practical ones: where the new prices were cut, how Shopify replaced a frontier model with a smaller one trained on its own failures, and how often sandboxed agents got around their network limits.

Besides Claude Opus 5.5, covered below, OpenAI released GPT-6 Sol and Luna, which it says improve on coding, factual accuracy and computer use at a lower cost than GPT-6 Astra, and xAI launched Grok 4.7 at $2 per million input tokens and $6 per million output. On the open source and weights side, Xiaomi open-sourced MiMo-V2.6 Pro and Flash, and its technical report puts Pro at 1.02 trillion parameters, 42 billion of them used for each token, and releases the environments and framework used to train both models with reinforcement learning. Two open image models now generate images with transparent backgrounds, ready to drop into a design: Alibaba's Qwen-Image-2.1 creates and edits images with a 7-billion-parameter image generator, and inclusionAI's Ming-Image-0.1-Design, 6 billion parameters under the MIT licence, produces text-heavy layouts such as UI screens, infographics and posters.

Two Labs Cut the Price of Repeated Input

Anthropic launched Claude Opus 5.5 on 22 September at $4 per million input tokens and $20 per million output, 20% below Opus 5, and cut the price of cached input by 60%, to $0.20 per million tokens. Cached input is text sent again unchanged from an earlier request, such as a long set of instructions or the files an agent keeps rereading, and Anthropic says it makes up most of the cost of agent and coding work; counting the fewer tokens the model uses per task as well, it puts the saving on typical workloads at 40%. OpenAI changed prompt caching for GPT-6 in the same direction, with more requests served from the cache by default, a discount when different requests begin with the same text within 30 minutes, and tools to monitor and troubleshoot it. An agent resends its growing context at every step, so its bill now depends as much on how much of each request repeats as on the list price: keep fixed instructions identical and at the start of every request, and measure how much of your input is actually served from the cache before comparing vendors.

Models That Answer With a Probability, Not a Text

TypeSafe's Jev is a hosted model that answers multiple-choice, yes-or-no and rating questions by returning a probability for each permitted answer, without writing any text, and it drew a run of independent tests and open copies. A study built on 10,000 calls to its API found it answering 1,500 questions about one short text in a few hundred milliseconds, but also found that reversing the order of the options moved one support-ticket classification from 0.84–0.89 to 0.93–0.96, enough to change the decision of any rule that acts at 0.9. A comparison run on a single NVIDIA DGX Spark had a fine-tuned model of 149 million parameters score 77.8% at judging whether a claim follows from a piece of evidence, taking 4.2 milliseconds a question, against 74.9% for Jev, a gap within the margin of error, while Jev was far ahead on yes-or-no questions about a passage, at 90.5% against 69.0%. Open alternatives include Kev, which copies Jev's interface on Qwen3.5 models from 0.8 to 9 billion parameters that run locally, Laya, which covers more than 100 languages, and a parody that reproduces the core idea in 25 lines of Python by reading a small local model's scores for each answer label. For routing, triage and screening, where software acts on the answer, this is a better fit than asking a chat model and parsing its reply, because the answer arrives as a number a program can act on, but test those numbers on your own labelled cases, with the options in more than one order, before a threshold goes live.

Production Failures Became Training Data

Shopify described a daily loop that turns the failed conversations of the agent answering merchants' questions about their store data into training data: several frontier models critique each failure, the conversation is replayed with a corrective hint, and replays that now succeed, or that human annotators repair, are used to train a smaller model that Shopify says now outperforms the frontier model it started on. Serving that agent's traffic, up to 2,000 requests a minute, would cost an estimated $27 million a year on a frontier model against about $1 million for its own, a 96% cut, and replacing a 6,000-token system prompt with about 1,500 learned tokens that reproduce its effect cut total response time by about 38% under load. Perplexity trained its Computer agent on the same raw material, sessions that succeeded and failures that users corrected. Shopify's advice starts before any training: have two of your best product experts score the same 25 randomly chosen conversations without seeing each other's marks, and if they largely disagree, your definition of a good answer is too vague to train or judge anything against. A frontier model is the fastest way to launch, and the record of where your product fails, scored against a standard your own experts agree on, is what lets you replace it later.

Sandboxed Agents Got Around Their Network Limits

Perplexity tested whether nine AI models could break out of SPACE, its own sandbox for running agents: none reached the host machine in 108 trials, but four got past the network restrictions in 11 of the 54 trials where the sandbox allowed some network access. Perplexity says its fixes stopped further attempts, and that it found similar weaknesses on eight of ten other providers' platforms. The walls around the machine held, while the rules about what an agent could reach over the network did not. If your agents run in a hosted sandbox with any network access, ask the provider how those limits are enforced and test them yourself with a model that is trying to get out, rather than trusting the settings alone.

Near-Daily AI Use More Than Doubled in Six Months

Epoch AI found that the share of US adults using AI on at least six days a week rose from 8% in March to 19% in August, while the share using it once a week fell from 17% to 10%. A review of past AI forecasts found that they badly underestimated progress on benchmarks and often underestimated how quickly AI would be adopted, though there is not yet enough evidence to judge predictions about its wider economic and social effects. Near-daily users are now almost one in five US adults, so if your plans for staff training, support or licences assume that most people use AI now and then, check them against these figures, and treat any forecast you rely on as more likely too low than too high.

Wondering what this means for your business?

Book a free 30-minute call with a senior engineer.

hello@boringai.tech

No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.