5 developments
OpenAI says GPT-5.6 tripled its score on ARC-AGI-3, a benchmark of unfamiliar puzzle-like games designed to test whether a system can figure out rules it has never seen. The model did not change at all. OpenAI flipped two API settings: one lets the model keep its own reasoning from one step to the next instead of throwing it away, and one compresses its accumulating history so it does not run out of working room.
Why it matters, our judgment: A threefold jump from configuration alone means that, on this benchmark, GPT-5.6’s score was set as much by its scaffolding (its memory and context settings) as by the model itself. The practical lesson is that a benchmark number measures a model plus its configuration, so treat vendor score claims, and your own model comparisons, as claims about both. Whether other models and benchmarks hide similar jumps is the right question to press vendors on, though this one result cannot answer it.
If a threefold jump was hiding behind two settings, our bet is that today’s models hold more intelligence than anyone has measured, and that the world could change fast without a single new model being trained. So tell us where you land: are these systems genuinely smarter than we know, or does capability that needs this much coaxing to appear not really count as capability at all?
OpenAI published a notice saying it is sharing preliminary cybersecurity evaluations of a model called Astra and strengthening the safeguards and security controls around it, before saying much else about the model. The only framing the company offers is the title, ‘Responding to the next frontier of critical cyber capabilities’; the notice does not say what Astra can do.
Why it matters, our judgment: The sequencing is the story: in this notice, OpenAI leads with safety evaluations and tightened controls and never gets to what the model can do, the reverse of how launch announcements usually read. That suggests safety evaluations, not training runs, are becoming the actual gate on what OpenAI ships, and cyber is where it chose to show that first, though a one-sentence notice tells us nothing about what Astra can actually do.
Right now a handful of private labs decide what is too dangerous to describe about a technology that affects everyone. Should that judgment stay with the labs, or does a decision this big belong to some body the public actually controls?
Microsoft Research built twelve working practice worlds, sandbox replicas of business software with real databases underneath, where an AI agent can click through tasks, break things safely, and be graded on what actually changed in the data. A small 9-billion-parameter model trained inside these worlds went from 36.5% to 67.1% on computer-use tasks, landing within fourteen points of GPT-5.4, a frontier model many times its size.
Why it matters, our judgment: The gap to the frontier closed here without a bigger model or more compute, just better places to practice, which raises the question of whether the binding constraint on capable agents is model scale or the quality of the simulated worlds we can build for them. The transfer result is the striking part: drilling a single annoying control like a date picker, rendered many ways, taught the model to operate it in software it had never seen, and if that transfer holds outside these twelve worlds, narrow practice can buy general skill. And the process itself is a loop, where the model’s failures tell the builders where to deepen the world and the deeper world produces a better model, a crude version of a system that improves through its own mistakes.
If that holds, capable agents get cheap wherever someone bothers to build the practice ground, not just at the frontier labs. So the question worth answering before a competitor does: which part of your business would an agent learn first?
Epoch and METR released MirrorCode, a test where an AI must rebuild a working software program from scratch while only being allowed to poke at the original from the outside, typing commands and watching what comes back, with no source code and no internet. Claude Opus 4.7 completed one of these rebuilds in 14 hours of work costing $251, a job the researchers estimate would take a human programmer 2 to 17 weeks. Most of the 25 target programs fell: 17 were rebuilt perfectly at least once, though a few, like the Python code-checker ruff, never were.
Why it matters, our judgment: If a machine can rebuild any software just by using it, then eventually any person can have any computing platform ever built, made for them, on demand. Every tool every company ever created becomes something one individual can summon for their own use, and that is the real story here, not who owns what. The evidence keeps this from being fantasy: the model had no source code, only the ability to run the original and watch what came back, and it produced its own complete working version in 14 hours for $251. Capability that used to require a company now answers to a single person, though eight of the 25 target programs were never rebuilt perfectly.
We think the moat around software just drained: capability that used to take a company was rebuilt in 14 hours for $251 by watching a program run. If any person can summon any computing platform ever built and bend it to their own use, does that start to approach AGI for the individual, or the singularity itself?
Microsoft Research released Orchard, free and open infrastructure for training AI agents, the systems that work through multistep jobs like repairing code or navigating websites on their own. A model trained with it, running on only about 3 billion active parameters, scored 69.7% on SWE-bench Verified, a test that asks an AI to fix real bugs drawn from real open-source projects, putting it near systems built on models more than ten times its size. Microsoft is releasing the training data and evaluation methods along with the code.
Why it matters, our judgment: If a small model trained with the right infrastructure can get this close to frontier systems on one benchmark of real bug fixes, then some of what looks like a scale advantage may really be an infrastructure advantage, and Microsoft just published that infrastructure openly, including the ability to train an agent inside the same harness it will actually run in. It raises the question of whether agent capability will spread far faster than raw model capability did, because training environments are cheaper to copy than hundred-billion-parameter models are to build. That would move the contest from who has the biggest model to who has the best places to train one, and it bears directly on whether buying ever more compute keeps paying, though one benchmark is not the whole of agent work.
If capable agents can be trained cheaply by anyone with the right infrastructure, then advanced machine intelligence stops being something a few companies own and becomes something almost anyone can build, and that spread, more than any single breakthrough, is what would make the change hard to contain. We think diffusion matters more than the frontier here, but do you agree, or does whoever holds the biggest model still set the terms for everyone else?