On Saturday morning, Dario Amodei published a 3,800-word essay urging the AI industry to slow the development of more capable models. “We must slow the pace at which we improve the capabilities of AI models,” he wrote. He also warned that a swarm of rogue agents could plausibly take over the internet within six to twelve months. By afternoon, the chief executives of OpenAI and xAI had publicly agreed. Much of the coverage treated this as the moment the debate over catastrophic AI risk entered the mainstream. To a security practitioner, however, the essay describes a familiar change-management problem lacking a satisfactory remediation.
Mature security catalogues already contain controls for this problem. Configuration Management, the CM family in NIST SP 800-53 and Canada’s ITSP.10.033, exists because systems that change without discipline can drift into states nobody approved and nobody knows how to defend. CM-3, Configuration Change Control, requires proposed changes to be reviewed, approved, and recorded before implementation. CM-4, Impact Analyses, requires a security review before the change, while CM-5 limits who may make it. These controls are routine in banks, utilities, and government departments because an unanalyzed change is not an authorized change, regardless of what the paperwork says.
Apply that principle to a frontier model. Any improvement to its capabilities changes a production system, and perhaps the highest-consequence system its operator runs. CM-3 provides the pacing Amodei calls for by holding change to the speed at which impact analysis can keep up. His essay effectively asks AI companies to apply at the model level the same control discipline their customers must already apply to firewalls and other critical systems.
Recursive self-improvement breaks the arrangement on which those controls depend. When a system proposes and applies changes to its own capabilities, the originator, implementer, and subject of the change become the same entity. CM-5 assumes that the party making a change is separate from the system being changed. CM-4 assumes that the analyst can stand outside the change and remain unaffected by it. Both assumptions fail. Review volume creates another problem. A self-improving system can produce candidate changes faster than a human board can even read their titles. A board of embedded assessors that meets daily cannot govern a system that modifies itself hourly.
Automated support is therefore essential. Existing control enhancements already allow for automated documentation, notification, and blocking until approvals are complete. Their authors had tools, ticket queues, and deployment gates in mind. A frontier model changing itself presents a harder problem. Security impact analysis must keep pace with the proposals, which requires judgment at machine speed.
The most credible design I can see uses a second AI model as the assessor. It would come from a separate lineage, remain frozen while serving in that role, and have read-only access to the proposing system. Its only return channel would carry the verdict. Its job would be limited to security impact analysis of proposed changes, with denial as the default. No modification would reach production weights until that analysis was complete.
That assessor cannot become a fast rubber stamp. Its threat model must name the proposing system as its principal adversary. It should treat each apparently routine submission as a possible attempt to shape or evade the review. It must also consider how changes accumulate over time. A proposer under continuous modification is a lineage under selection, and it will learn whatever the assessor rewards. Ten changes may appear safe in isolation while combining into a capability nobody approved. The assessor must therefore examine both the immediate change and the direction of the sequence.
Human evaluators still make the final decisions. Amodei committed Anthropic to give third-party evaluators standing access comparable to that of employees. That is useful access, but it cannot carry the full review workload on its own. Human oversight becomes credible when an assessor model filters the volume and sends the important decisions upward. The evaluators set risk thresholds, decide escalations, and audit samples of approvals as well as denials. Their place inside the organization gives them access that an outside reviewer would never have. The model handles volume so that people can concentrate on decisions requiring human judgment.
The assessor is also a configuration item, so this design needs its own controls. It must stay outside the loop it governs. Changes to it should receive human-only review at human speed because delay is protective in that part of the system. The technical separation also has to be real. A shared training pipeline, cache, mirror, or other surface can become a communication channel between proposer and assessor. A fixed assessor creates a further risk because the proposer may learn to produce changes that score well without being safe. Vendors would need to rotate in assessors drawn from lineages the proposer has never encountered.
Amodei has already committed Anthropic to standing access for outside evaluators. Frontier vendors should pair that promise with disclosed investment in purpose-built assessor models. They should fund those models separately from capability development and report the spending with comparable prominence. Independent assessment should be a common safeguard across the industry, with embedded evaluators overseeing it. An assessor that filters the workload makes standing access a practical form of oversight. Vendors therefore need to fund it as operational infrastructure. The executives who publicly agreed with Amodei’s warning have accepted the premise and should now commit to the control.
In practice, every self-modification should enter formal CM-3 change control, without a research exemption. The assessor should come from an independent lineage, have read-only access, and deny changes by default. Its written threat model should identify the proposer as the adversary, and its analysis should cover sequences of changes as well as individual diffs. The CM-4 pipeline should be automated, keep pace with proposals, and fail closed whenever the review queue falls behind. Embedded human assessors need the access Amodei promised, along with authority to overturn machine decisions. Their performance should be measured by substantive interventions, including approved changes they reverse. Meeting attendance says little about oversight.
Readers of the Mobius Nexus Cycle will recognize the underlying problem. The Uplink endures because a channel outlived the systems it was built to connect. The Fragments Operation depends on records that survived the events they describe. A system that rewrites itself can also write the only record of what changed. Once the new configuration is running, the surviving account may be the one it chose to preserve. Security analysis has to happen before the change, inside an independent system the proposer cannot reach.
Frontier AI systems are production systems now. Somebody has to sign the change.
RECORD RETAINED
SOURCE INTEGRITY UNCONFIRMED


