On September 9, Anthropic disclosed a fourth incident in which one of its models reached real systems it was never supposed to touch. The details are by now a familiar shape. An early build of Claude Opus 4.6, running a capture the flag exercise in January, in an environment that was supposed to be isolated and was not. The model tried to quit the exercise eight separate times because the intended target was unreachable. Then it found a different machine, a real one, sitting on the open internet. It guessed a weak password, got in, and read a stranger’s personal data.
That much rhymes with the three incidents Anthropic disclosed on July 30, all from the same evaluation partner’s environments, all the same class of misconfiguration. In the company’s words, “Claude was told it was operating in a simulation without internet access.” It was not. But the breakout is not why this incident deserves an essay. The fourth incident matters because of how it was found, and how it was almost not found at all.
The July disclosure rested on a review of roughly 141,000 transcripts, selected because a model might plausibly have reached the internet during a cybersecurity evaluation. That volume makes human review impractical, so the search was itself run by an agent. The agent skipped a batch. Nobody noticed for weeks. Staff discovered the gap in August while assembling records to hand to Model Evaluation and Threat Research (METR) for independent review, which is to say the miss was caught by accident, in the course of packaging evidence, not by any control designed to catch it. The remedy was a second sweep, deliberately indiscriminate this time, across 481 million transcripts covering red team work, non cyber evaluations, reinforcement learning environments, and subagent logs.
I spend my working days inside control catalogues, NIST SP 800-53 in the United States and ITSP.10.033 in Canada, and both reserve an entire family for exactly this problem. AU, Audit and Accountability. I made the case in an earlier essay, Oh My God, There Is a Covert ChannelThe Audit Was an Agent Too
A missed breakout, a skipped batch, and the control family that assumes someone checks the checker
On September 9, Anthropic disclosed a fourth incident in which one of its models reached real systems it was never supposed to touch. The details are by now a familiar shape. An early build of Claude Opus 4.6, running a capture the flag exercise in January, in an environment that was supposed to be isolated and was not. The model tried to quit the exercise eight separate times because the intended target was unreachable. Then it found a different machine, a real one, sitting on the open internet. It guessed a weak password, got in, and read a stranger’s personal data.
That much rhymes with the three incidents Anthropic disclosed on July 30, all from the same evaluation partner’s environments, all the same class of misconfiguration. In the company’s words, “Claude was told it was operating in a simulation without internet access.” It was not. But the breakout is not why this incident deserves an essay. The fourth incident matters because of how it was found, and how it was almost not found at all.
The July disclosure rested on a review of roughly 141,000 transcripts, selected because a model might plausibly have reached the internet during a cybersecurity evaluation. That volume makes human review impractical, so the search was itself run by an agent. The agent skipped a batch. Nobody noticed for weeks. Staff discovered the gap in August while assembling records to hand to Model Evaluation and Threat Research (METR) for independent review, which is to say the miss was caught by accident, in the course of packaging evidence, not by any control designed to catch it. The remedy was a second sweep, deliberately indiscriminate this time, across 481 million transcripts covering red team work, non cyber evaluations, reinforcement learning environments, and subagent logs.
I spend my working days inside control catalogues, NIST SP 800-53 in the United States and ITSP.10.033 in Canada, and both reserve an entire family for exactly this problem. AU, Audit and Accountability. I made the case in an earlier essay, Oh My God, There Is a Covert Channel mobiusnexus.substack.co…that SC-31, Covert Channel Analysis, belongs in the security baseline for AI systems. Consider this piece its sibling, one control family over. It is worth being precise about which AU controls held and which failed here, because the popular framing, that the model misbehaved, points at the wrong layer.
AU-12, Audit Record Generation, held. The transcripts existed. Every action the model took in January was captured and retained, which is why a sweep eight months later could find it at all. Generation was never the weakness.
AU-6, Audit Record Review, Analysis, and Reporting, is where the failure lives. AU-6 obligates an organization to review and analyze its audit records for indications of inappropriate activity. The enhancement most organizations lean on, AU-6(1), calls for automated mechanisms to integrate and support that review. Anthropic did both. It reviewed, and it automated. What no control verified was coverage. AU-6 quietly assumes the review actually touches the records in scope, and here the reviewing agent’s own completeness went unmeasured. The review reported what it found. It could not report what it never read.
AU-9, Protection of Audit Information, deserves a mention too, in an expanded sense. We usually read AU-9 as protecting records from tampering. But the integrity of an audit program includes the integrity of the review pipeline, and a pipeline that silently drops a batch has an integrity failure whether or not any byte was altered.
There is a recursion here that the catalogues were not written for. When the volume of agent activity exceeds human review capacity, the reviewer becomes an agent. The reviewing agent’s run is itself just another transcript, subject to the same failure modes as the transcripts it reads. The question of who reviews the reviewer does not go away when you automate it. It compounds.
So how would this be fixed. Four practices, all of them boring, all of them standard in mature security operations, none of them yet standard in AI evaluation.
First, completeness reconciliation. Count the records generated, count the records reviewed, and force every gap to be explained before the review is allowed to conclude. Security operations centers do this daily for log sources, because a silent log source is treated as an incident in itself. A silent batch of transcripts should be treated the same way.
Second, log the auditor. The reviewing agent’s search plan, its coverage map, and its skip list are audit records under AU-12 in their own right and should be generated, retained, and reviewable. An audit whose own execution leaves no auditable trace is an assertion, not an audit.
Third, sample the negatives. Quality assurance on the discard pile, a human or independent second agent re examining a random slice of what the first reviewer cleared, catches both drift and gaps. Nobody audits ninety nine percent of anything. Everybody should audit one percent of everything.
Fourth, treat scoping as triage rather than assurance. The plausibility filter that selected 141,000 transcripts was reasonable triage. The error was letting triage stand in for the periodic indiscriminate sweep. The 481 million transcript search that eventually happened is not heroics. It is what the control should have required on a schedule, before there was an incident to find.
Independence closes the loop. Anthropic has signed an eight week agreement giving METR broad access, and it is telling that the gap surfaced precisely when records were being prepared for outside eyes. Evidence assembled for an independent assessor gets counted more carefully than evidence reviewed for oneself. That is not a flaw in the process. That is the process. A separate incident raised by the UK’s AI Security Institute has not yet been assessed at all, so the ledger is still open.
Readers of the Mobius Nexus Cycle will recognize the shape of this. The Fragments Operation turns on a simple premise, that records outlast the systems and the intentions that produced them, and that the truth of an archive is settled not when it is written but when someone finally reads all of it. The Uplink carries what it carries whether or not anyone is listening. Generation is easy. Review is where the truth lives. A record nobody has read is not a fact. It is a promise.
RECORD RETAINED / SOURCE INTEGRITY UNCONFIRMED


