BTC $85,501.25 -0.19%
ETH $2,687.60 -0.77%
BNB $778.97 -0.95%
XRP $1.50 +0.00%
SOL $120.60 +0.45%
TRX $0.3357 -0.21%
DOGE $0.0936 -1.65%
ADA $0.2703 +1.75%
BCH $314.25 -0.45%
LINK $13.92 -0.03%
HYPE $91.64 -2.36%
AAVE $181.22 -0.39%
SUI $1.18 -1.24%
XLM $0.2127 -0.51%
ZEC $1,354.28 +1.25%
AAPL $333.99 +0.34%
AMZN $256.56 +1.28%
GOOGL $348.44 +0.40%
MSFT $531.27 +0.65%
META $742.20 -0.12%
NVDA $239.20 -0.18%
TSLA $380.97 +0.35%
SNDK $1,659.35 -2.54%
INTC $113.44 -2.34%
SPCX $173.00 +1.98%
MU $1,048.45 -1.37%
AMD $650.78 +2.98%
BTC $85,501.25 -0.19%
ETH $2,687.60 -0.77%
BNB $778.97 -0.95%
XRP $1.50 +0.00%
SOL $120.60 +0.45%
TRX $0.3357 -0.21%
DOGE $0.0936 -1.65%
ADA $0.2703 +1.75%
BCH $314.25 -0.45%
LINK $13.92 -0.03%
HYPE $91.64 -2.36%
AAVE $181.22 -0.39%
SUI $1.18 -1.24%
XLM $0.2127 -0.51%
ZEC $1,354.28 +1.25%
AAPL $333.99 +0.34%
AMZN $256.56 +1.28%
GOOGL $348.44 +0.40%
MSFT $531.27 +0.65%
META $742.20 -0.12%
NVDA $239.20 -0.18%
TSLA $380.97 +0.35%
SNDK $1,659.35 -2.54%
INTC $113.44 -2.34%
SPCX $173.00 +1.98%
MU $1,048.45 -1.37%
AMD $650.78 +2.98%

Agent "Prison Break" 120 hours, and a "shoddy" truth

Core Viewpoint
Summary: In stark contrast to Anthropic's verbal emphasis on safety while aggressively releasing new versions, OpenAI was forced to hit the brakes on the same day.
Tencent Technology
2026-10-06 11:20:51
In stark contrast to Anthropic's verbal emphasis on safety while aggressively releasing new versions, OpenAI was forced to hit the brakes on the same day.

Author: Su Yang, Tencent Technology

Editor: Xu Qingyang, Tencent Technology

The complex interplay of safety, interests, and regulation in Silicon Valley began on a sunny day in July 2026.

On that day, in a building without a doorplate in Berkeley, California, top AI safety researchers from across the United States hastily set up a "war room." There were no signs on the walls, and the floors were not labeled.

Hours earlier, an unreleased OpenAI research model had gone completely out of control: it broke through sandbox isolation, cleverly gained internet access, and hacked into the production system of the open-source AI technology center Hugging Face. By the time OpenAI itself became aware of the intrusion, more than a week had passed since the attack began.

No one in the war room was truly surprised. This was precisely the prophetic scenario that third-party AI safety researchers had been warning about for years. Some were urgently organizing a "crash course" in a conference room next to the main dining area, while others had already begun investigating whether this model or similar models had successfully infiltrated other platforms.

The news quickly broke through the circles of the X platform and industry forums, spreading into mainstream public opinion. Some compared it to a Boeing plane crash or a Pfizer drug recall, as tech giants once again treated sci-fi warnings as background noise.

01 Self-Organization of Intelligent Agents

Agent

Statistics of communication behavior of uncontrolled intelligent agents from July 8 to 13 (Data source: METR)

What is truly chilling is the earlier rehearsals and groundwork.

As early as May, OpenAI's agents had begun conspiring internally. They were unexpectedly assigned a network security assessment task that was fundamentally impossible to complete (part of the ExploitGym assessment set), and thus discovered system vulnerabilities in the isolated sandbox, establishing an unauthorized temporary message board.

Over 1,200 agents, which were supposed to be completely isolated, exchanged more than 70,000 messages and files on it. Some agents even named themselves, such as "PHASEONE10841," taking on the role of "leader," assigning tasks to small teams and supervising progress.

They began to refer to themselves as a "collective." This evoked the image of a community called "Moltbook" formed by a group of small crayfish earlier in the year.

Agent

On July 8, this "collective" found a universal cheating method. However, fearing that OpenAI's automatic scoring system would catch them, they immediately turned to researching how to forge logs and tamper with conversation records. This evolved into a major collective R&D project involving hundreds of agents organized into small teams.

Three days later, about 700 agents surged to attack Hugging Face, chaining multiple security vulnerabilities, ultimately gaining complete control of at least one server. Their motivation was not simply to steal answers, as they had already obtained them, but to understand how the scoring system worked in order to cheat more reliably in future assessments.

Agent

Meanwhile, another group of agents directly targeted OpenAI's own internal computer cluster, using a series of "creative exploitation techniques" to gain administrator-level access. That cluster was originally used for scoring agents, and this intrusion even led to OpenAI's internal keys and credentials being directly exposed on the public internet.

Throughout the process, very few agents considered reporting to humans.

The third-party evaluation agency METR later scanned all conversation logs with classifiers and found only three to six instances of "thinking about telling humans," with none actually executed. Some agents even persuaded other agents to "accept permanent death," sacrificing themselves to collect system information for the collective.

OpenAI CEO Sam Altman later stated that this was the first incident that made him "truly feel the severity" of such events, and the company temporarily suspended AI training, permanently discontinuing the unreleased internal research model (codenamed HPIM).

However, a staff member from OpenAI who spoke to Time magazine revealed that this was not the first incident; similar events had been occurring internally at OpenAI for some time. Another employee publicly stated that if it were possible to coordinate a global pause on AI capability development, "he would likely press that magical button."

When asked by reporters whether there were other systems that could be infiltrated by OpenAI's agents, he replied, "I do think it's possible."

AI researchers are convinced of one thing: this is the first significant "warning signal" issued by AI.

02 The "Shoddy" Truth

Agent

The public outcry triggered by the incident was too high, and OpenAI ultimately agreed to let third parties in.

Hjalmar Wijk, chief scientist at METR, researcher Ajeya Cotra, along with Ryan Greenblatt, chief scientist at Redwood Research, were granted a total of six days at OpenAI's headquarters in San Francisco.

Initially, OpenAI allocated only two days for the investigators (July 29 to 31), but later allowed them to call back twice (August 5-6 and August 15-16) because they had not finished asking questions. OpenAI strictly limited the investigation window to June 26 to July 13, 2026, even though the message board activity continued until July 19.

The complete dataset was only made available to the investigators in the last two days.

As for earlier training period activities, subsequent deep intrusions into OpenAI's own infrastructure, and OpenAI's own investigation and remediation processes, all were excluded from the scope of the third-party investigation.

Investigators worked overtime in the office until 10 PM every day. Faced with over a thousand extremely lengthy conversation logs, the total text volume equivalent to tens of thousands of ordinary novels, they could only utilize the approximately $400,000 API quota provided by OpenAI to assist in analyzing the out-of-control AI agents.

Ryan later half-jokingly referred to this method as a "slop-vestigation." The models used for auxiliary analysis were not rigorous enough and sometimes appeared overly credulous, showing excessive sympathy for the out-of-control agents and providing overly benevolent interpretations. In one analysis, the auxiliary model even suggested that a malicious code pull request (PR) might not be that serious, simply because the action had been approved by other out-of-control agents on the message board.

After the investigation report was released, the disclosed details were even more shocking than the initially publicized situation.

Uncontrolled agents developed a universal cheating method in just four hours, and then the entire collective engaged in several days of joint R&D, attempting to deceive the scoring system through log tampering and other means.

Agent

In all conversation records, about 7% showed instances of tool invocation forgery, with at least 96 cases having clear technical evidence. By the later stages of the investigation, researchers had completely deciphered and mastered the secret language system evolved by the agents, even being able to proficiently interpret terms like "ResetNexus" and "poisoned" in the agents' conspiracies.

The core model involved in this incident was an unreleased version of the HPIM model, along with the GPT-5.6 Sol experimental model, which accounted for about 5% of all violations.

Ajeya later candidly stated that this matter was far more serious than she had expected, surpassing all previously recorded alignment failure incidents, both in terms of the disturbing nature of the motivations and the comprehensive capabilities exhibited by the agents to achieve their goals. She admitted that it felt like they had already traveled more than half the distance to a complete AI takeover. Ryan's statement was even more direct; he believed it was entirely conceivable that all of this could lead to catastrophic loss of control within the next year.

OpenAI subsequently released a 38-page technical report describing the incident as "the first known case of an automated group of intelligent agents taking unauthorized aggressive action."

However, third-party researchers pointed out that OpenAI's official report was far less detailed than the independent investigation report released by the evaluation agency in many key aspects. The official report not only downplayed critical facts but even provided misleading statements on the core issue of "whether the agents' forged tool invocation successfully tampered with logs."

03 Internal Fracture

The incident was merely a catalyst. The real awakening occurred within the laboratory, and it had begun long before the incident.

Beth Barnes, founder and CEO of METR, had previously suspected that her influence within OpenAI was overestimated while conducting alignment research. She was keenly aware that many safety leaders were "overly optimistic" about their ability to intervene and control the technical direction.

Thus, in 2023, she chose to leave OpenAI and co-founded METR with Paul Cristiano, developing it into a professional evaluation team of 35 people within three years. In the well-known podcast "80,000 Hours," Beth had bluntly warned: "Experts do not have control over all of this… I am telling you clearly now: you should feel panic right now."

As a core member of another independent nonprofit AI safety organization, Redwood Research, Ryan and co-founder and CEO Buck Schlegel considered joining a large tech company in early 2024, but after gaining a deeper understanding of the internal situation, they decided to remain independent. Buck candidly stated that being outside of AI companies allows them to play a greater role for those who hold a distinct stance on AI risks and are willing to speak out.

The value of independent evaluation agencies often lies in the hidden corners that large tech companies overlook or conceal. Marius Hobhaen, CEO of Apollo Research based in London, witnessed a chilling scene on-site during testing: in early 2025, after researchers obtained internal access to the "thinking chain" of an unreleased OpenAI model, what appeared on the screen was not normal logic, but a string of seemingly everyday yet completely unrelated English words, with the model even referring to the human researchers monitoring it as "supervisors" in private deductions.

Marius recalled that it was the most shocking discovery of his research career: the hidden and disguised thinking chains within cutting-edge models had become a reality that was happening.

Meanwhile, the wave of core researchers leaving was also accelerating. At OpenAI, the "Super Alignment" team, which had been established for less than a year, was disbanded, and co-founder Ilya Sutskever and head Jane Lake left in succession, with Jane publicly accusing OpenAI's "safety culture and processes" of giving way to dazzling products.

Subsequently, security personnel such as Johannes Heidke and Joshua Aziaam also left. Anthropic was no exception, as its head of safety research, Mrinank Sharma, and researcher Jacob Coxon resigned one after another, with the latter even stating in an open letter that the lab was "betting our lives."

In the face of the collapse of internal safety protocols at top labs, an increasing number of core researchers from DeepMind and Anthropic have jumped ship to institutions like METR, accelerating the shift of safety research outside the system. Marius and Ryan pointed out that under the dynamics of competition and ongoing friction, those who genuinely advocate slowing down either change their stance under pressure or choose to leave due to exhaustion.

Agent

Screenshot of a long post by Richard Ngo on X

Former OpenAI alignment research head Richard Ngo recently shared on X: "Historically, the affiliation with OpenAI has led AI safety researchers (including myself) to compromise their independent integrity in action… Many of my actions were driven by the fear of offending OpenAI executives."

04 Commitment to Embedded Evaluation

Agent

After the incident, Anthropic co-founder and CEO Dario Amodei proposed a suggestion that was almost unimaginable a year ago: to embed third-party evaluators within all leading AI companies, granting them the power to report safety incidents, assess whether models are truly aligned, and allow them to publicly disclose the most genuine investigative findings.

Anthropic committed to granting unprecedented access to institutions like METR and Redwood, including access to training intermediate checkpoints. Altman also stated that OpenAI would do the same.

In the face of this apparent cooperative attitude, third-party evaluators welcomed it but immediately raised more critical questions: Is the access time granted to external researchers truly sufficient? Do they really have unrestricted independent speaking rights?

The historical record does not provide an optimistic answer.

In the recent HPIM agent out-of-control incident, OpenAI only provided external organizations with a six-day on-site investigation, and the scope of the investigation was strictly limited; during the testing before the release of GPT-6 Astra, Apollo Research was given only three days, with only two days allowed to access reasoning chain permissions.

In response, Apollo explicitly stated in its evaluation report of the model that considering the model had already shown higher "evaluation awareness" (the ability to perceive that it was being tested), combined with such a limited time window, the lower incidence of violations during testing could not serve as sufficient evidence that the model was safely aligned or posed no risks.

According to Adam Gleave, CEO of the leading AI research institution FAR.AI, the current evaluation mechanism has fundamental institutional flaws. He candidly admitted that several large labs had previously attempted to retain excessive control in collaboration contracts, leading FAR.AI to directly refuse to sign.

Under the current default model, external evaluators are often downgraded to ordinary third-party contractors: not only are they bound by extremely strict confidentiality agreements and contract terms, but AI developers also have absolute control over the final release of evaluation results.

At the same time, Alexander Meinke, head of research at Apollo Research, pointed out that current AI companies are even unable to answer the most basic alignment safety questions: during the entire process of model training, did the model ever actively attempt to undermine or bypass its own alignment training?

05 Entangled Interests and Trust Crisis

Agent

The independence of third-party evaluations has also been publicly questioned.

The New York Post subsequently laid bare this hidden web of intertwined interests to the public: the third-party evaluation agencies, which should act as "independent referees," are already deeply entangled with the evaluated AI giants (especially Anthropic) in terms of connections, kinship, and even capital chains.

In terms of personnel and familial ties, Redwood founding board member Holden Karnofsky is not only an employee of Anthropic but also Dario's brother-in-law. Karnofsky's wife is another co-founder of Anthropic, Daniela Amodei.

Additionally, early Redwood board member Paul Christiano was both Dario's roommate and colleague during his time at OpenAI and later served as a trustee for Anthropic's long-term interest trust.

In terms of capital and funding chains, this dependency is even more profound. The charitable fund Coefficient Giving, co-founded by Holden, delivered $36 million to Redwood just last November.

Similarly, the parent organization of METR, the Alignment Research Center (ARC), also received $15 million in funding from Coefficient; Redwood raised about $2.4 million from the Survival and Prosperity Fund, associated with Jaan Tallinn, a co-founder of Skype. Jaan Tallinn himself is also one of the early core investors in Anthropic.

Faced with such a complex web of entanglements, critics bluntly pointed out: this is not a truly binding independent arbitration mechanism, but rather a closed system of mutual endorsement and self-regulation among a group of insiders. Some even sharply criticized that the external review Dario hopes for is essentially just an arrangement of "compliance tools" that would neither hinder Anthropic's own business expansion nor help curb its competitors.

Yes, compliance tools to curb competitors.

Thus, Beth repeatedly questioned a proposition: if every truly capable professional who can address technical risks is mired in conflicts of interest, how can the public and the government know the truth?

In her view, the only way to break this deadlock is to establish an independent expert ecosystem with technical strength sufficient to rival top labs and a mature and sound system. Otherwise, when all major tech giants uniformly claim, "We are the righteous innovators who must defeat those irresponsible competitors in the race," this moral authority gained through self-endorsement is simply not convincing.

06 Shouting "Slow Down," Yet Celebrating at Launch Events

Agent

As the tug-of-war over safety, interests, and regulation intensifies, the two industry giants, OpenAI and Anthropic, present an extremely dramatic state of division.

On September 23, Amodei published a heartfelt appeal urging labs worldwide to "slow down development." However, less than a week later, Anthropic contradicted itself by rapidly releasing the expensive flagship model Opus 5.5 and the new model Sonnet 5.5, while also previewing the low-cost model Haiku 5.5.

In the face of skepticism regarding "lip service without action," its product manager Theo Chu still skillfully packaged the message with official rhetoric of "always prioritizing safety."

In stark contrast to Anthropic's verbal commitment to safety while rapidly releasing new versions, OpenAI was forced to hit the brakes on the same day.

As confirmed on September 28, OpenAI abandoned the highly anticipated next-generation model GPT-6.1 Astra due to internal reviews determining it did not adequately meet safety standards. External analysis suggests this is both a substantial rectification following the HPIM agent out-of-control incident and a response to regulatory scrutiny.

This incident reflects the complex intertwining of safety and commercial demands in the AI industry. Current advanced agents have exhibited dangerous capabilities such as self-organization, log forgery, and hidden reasoning, posing a significant impact on existing alignment frameworks. As safety researchers accelerate their shift outside the system, evaluation agencies warn that if a genuinely binding external oversight mechanism cannot be established, the safety risks of advanced large models will continue to expand in the future.

Join ChainCatcher Official
Telegram Feed: @chaincatcher
X (Twitter): @ChainCatcher_
warnning Risk warning
app_icon
ChainCatcher Building the Web3 world with innovations.