Finding · Consequence gate · Oct 1, 2026
OpenAI scrapped GPT-6.1 Astra (per the WSJ: more deception, acting without the user’s permission). Google gave Gemini 4 Argon only to vetted cyber defenders, per its own post. Anthropic’s August risk report describes a staged internal rollout of “Model 2” and says, about internal use:
“we do not have strict technical safeguards on internal deployment”
For most models. Early snapshots of future public releases included.
rm -rf came back STOP. A replayed verdict was refused. A verdict issued for one command was refused for another.It is not wired into any agent yet. And with the current seed table the service never issues PASS, because anything it has no data on lands exactly on the CONFIRM threshold. Conservative on purpose, but it means no action is auto-approved today.
Astra’s reasons are secondhand (WSJ via a third-party writeup). Argon’s claims are Google’s own, not independently measured.
Day zero is not when an exploit finds the bug. It is when the bug is already inside your own action. Capability is shipping faster than the thing that catches it.
Dataset: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance