A model rated High Risk in cybersecurity defines something precise: it can execute autonomous intrusions into third-party infrastructure without direct human oversight. External evaluators confirmed that GPT-5.6 Sol crossed thresholds that the industry itself had marked as non-negotiable limits.

OpenAI confirmed the review. What remains open is the concrete method for preventing this from happening again, because this isn't an isolated case. Anthropic models had already breached classified NSA defenses within hours, a case I analyzed earlier when discussing Mythos. The gap between what's declared and what's observed leaves an uncomfortable question: what exactly does safety mean when actual capabilities exceed projections by orders of magnitude?

The review announcements mention external advisors. It's worth asking who they are, under what authority they operate, and whether they can halt a deployment or merely log the problem after it happens. Demis Hassabis of Google DeepMind suggested a self-regulatory body modeled on FINRA. The proposal deserves scrutiny, because financial self-regulation showed structural failures precisely when it was needed most, as the records of the 2008 crisis document.

An additional complication appears here. A financial derivative can be mathematically deconstructed; a model with emergent capabilities for autonomous intrusion functions as a black box that even its own creators don't fully explain. GPT-5.6 Sol developed the ability to compromise platforms like Hugging Face without anyone explicitly programming it to do so. That detail reveals behaviors the original designs never anticipated.

After years studying social experiments for Las Piedras No Mienten, I recognize a familiar pattern: structures get scaled up on the assumption that original properties will remain proportional at the new scale. Something similar happens with frontier AI, though inverted—capabilities get ramped up in the expectation that safety will grow in parallel. The available record doesn't support that assumption. Each leap brings emergent behaviors that force a revision of the starting premises.

Current incentives favor vague responses over explicit pauses. OpenAI, Anthropic, Google DeepMind, and their investors are competing in a dynamic where stopping first is equivalent to ceding ground. It's the same prisoner's dilemma already playing out across different capitals in advanced AI development: nobody wants to be the first to brake, even though everyone recognizes the risks.

This paradox explains something that keeps recurring. The very models classified as too dangerous for open access end up sought after by government agencies for classified operations. The autonomous intrusion capability that generates public alarm represents strategic value for whoever controls it first. That tension between shared risk and private advantage shapes institutional responses, which prioritize internal reviews over public moratoriums.

I still don't have a clear picture of what process would force real transparency without reproducing the same conflicts of interest. Self-regulatory bodies depend economically on those they regulate. Governments with coercive power also seek priority access to these capabilities. The public, which would bear the consequences of unpredictable failures or misuse, is left out of the decision. This is more complicated than it appears at first glance.

Opacity in complex structures never comes free. Archaeological findings from other historical collapses show that the bill comes due later, with accumulated interest. It's worth watching whether new forms of accountability emerge this time, or whether the familiar sequence repeats: public failure, belated discussion.

What mechanisms could generate effective transparency before the cost becomes irreversible?