14 September 2026 · 5 minute read
OpenAI says an AI system resolved a Millennium Prize Problem while Claude formalised Fermat's Last Theorem, agent failures track step count rather than context length, misconfigured tests gave Claude real access, data drove most pretraining gains, and long-context latency differs sharply between GPT-5.6 and Claude 5.
Read article →
7 September 2026 · 5 minute read
Four frontier launches land in one week, frontier cyber capability moves behind access programmes, the Hugging Face breach becomes an evaluation, a six-model fleet opens its full training record, and OpenAI cuts Cursor off after the SpaceX acquisition.
Read article →
31 August 2026 · 5 minute read
Two labs open the weights of the same sparse-attention architecture on the same day, NVIDIA takes one model from 30% to 100% by changing only the harness, a scientific-workflow benchmark stops agents at 20.6%, the content layer becomes the agent attack surface, and Meta builds a training chip around the network.
Read article →
28 August 2026 · 10 minute read
A practical guide to improving Lovable-generated code with AGENTS.md, naming conventions, database and Edge Function boundaries, environment-variable rules, testing, and security review.
Read article →
24 August 2026 · 5 minute read
GLM-5.3 and Inkling keep the open frontier level with the closed labs, the 27B class becomes genuinely self-hostable on one consumer GPU, a headline benchmark score turns out to be gameable, refusal training comes off in minutes, and expertise turns out to be the bottleneck.
Read article →
17 August 2026 · 5 minute read
OpenAI pauses a model over cyber capability, researchers steal reasoning traces from production APIs, open weights get very cheap, and Claude moves a Riemann hypothesis bound.
Read article →
10 August 2026 · 5 minute read
Open weights reach parity on regulated tasks, most of your agent bill turns out not to be prompts, cloud agents write half of Cursor's merged PRs, and robots get whole-body control.
Read article →
3 August 2026 · 5 minute read
Claude Opus 5 halves the price of the frontier, an autonomous agent breaches Hugging Face, harness settings triple a benchmark score, and Kimi K3 goes free.
Read article →
27 July 2026 · 4 minute read
A security evaluation escapes into production systems, sparsity becomes the scaling story, smart routing saves 30%, and verification tightens around coding agents.
Read article →
20 July 2026 · 4 minute read
GPT-5.6 ships, orchestration design is measured at 41% cost savings, a popular coding CLI uploads whole repositories, and agent evaluation moves to long-horizon tasks.
Read article →
13 July 2026 · 4 minute read
The harness becomes the story, OpenAI finds 30% of a popular coding benchmark broken, Grok 4.5 ships, and Anthropic builds an off switch for dual-use knowledge.
Read article →
6 July 2026 · 4 minute read
GPT-5.6 and Claude Sonnet 5 arrive, long-horizon coding benchmarks stay brutal, model routing cuts costs by 40%, and custom models beat the frontier on proprietary work.
Read article →
29 June 2026 · 4 minute read
Agent memory becomes a system, production agent frameworks arrive, RL-tuned coding agents learn to game their tests, and research agents leak private documents.
Read article →
22 June 2026 · 4 minute read
Open coding models get serious, inference speed becomes a systems problem, agent evaluation grows an infrastructure layer, and value moves to what cannot be trained.
Read article →
15 June 2026 · 4 minute read
Claude Fable 5 arrives, invisible safeguards force a public reversal, Gemma 4 puts multimodal AI on a laptop, and Stanford finds coding agents are bad teammates.
Read article →
8 June 2026 · 4 minute read
Open-weight agent models arrive in force, Microsoft ships its own model family, the open-closed gap is measured at four months, and a Starlette flaw exposes the AI tooling stack.
Read article →
5 June 2026 · 8 minute read
A practical guide to AI transformation: start with business outcomes, redesign the workflow, prove value quickly, and build the capability to keep improving.
Read article →
Want to talk about AI?
Book a free 30-minute call with a senior engineer.
No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too. Not ready for a call? Send us a message instead.
Want to know when we publish something?