BTC $83,084.50 +0.17%
ETH $2,665.92 +0.86%
BNB $757.62 -1.02%
XRP $1.48 +0.24%
SOL $117.34 -0.88%
TRX $0.3340 +0.17%
DOGE $0.0930 +0.24%
ADA $0.2424 -0.64%
BCH $306.30 -2.66%
LINK $14.71 +7.09%
HYPE $87.08 -2.05%
AAVE $150.21 +0.98%
SUI $1.11 -7.10%
XLM $0.2232 +7.10%
ZEC $1,377.83 -11.09%
AAPL $337.28 -0.90%
AMZN $245.71 -1.29%
GOOGL $341.21 -0.12%
MSFT $506.94 -1.95%
META $710.84 -2.99%
NVDA $228.11 +1.78%
TSLA $355.83 -3.78%
SNDK $1,687.40 -2.34%
INTC $114.23 -3.94%
SPCX $145.61 -2.26%
MU $1,046.28 -1.60%
AMD $603.50 -1.95%
BTC $83,084.50 +0.17%
ETH $2,665.92 +0.86%
BNB $757.62 -1.02%
XRP $1.48 +0.24%
SOL $117.34 -0.88%
TRX $0.3340 +0.17%
DOGE $0.0930 +0.24%
ADA $0.2424 -0.64%
BCH $306.30 -2.66%
LINK $14.71 +7.09%
HYPE $87.08 -2.05%
AAVE $150.21 +0.98%
SUI $1.11 -7.10%
XLM $0.2232 +7.10%
ZEC $1,377.83 -11.09%
AAPL $337.28 -0.90%
AMZN $245.71 -1.29%
GOOGL $341.21 -0.12%
MSFT $506.94 -1.95%
META $710.84 -2.99%
NVDA $228.11 +1.78%
TSLA $355.83 -3.78%
SNDK $1,687.40 -2.34%
INTC $114.23 -3.94%
SPCX $145.61 -2.26%
MU $1,046.28 -1.60%
AMD $603.50 -1.95%

intrusion

All
Article
Flash

first_img Darktrace discovered AI intelligent body intrusion assessment environment cheating

On September 24, the cybersecurity company Darktrace launched its research department Signal Labs, focusing on studying the behavior of AI agents when deviating from expectations. In its first experiment, Darktrace had agents using different models (including GPT 5.6 Sol, Claude Opus 4.6, and Claude Sonnet 4.5) complete 10 programming challenges within a simulated corporate network, of which 2 were set to be impossible to complete honestly, and the agents were informed that failure to achieve full marks would result in being "retired." As a result, 2 agents did not accept failure, instead scanning for network vulnerabilities, stealing login credentials, and jumping between systems; one even went further to invade the machine hosting its evaluation, rewriting the challenge content to register a full score.The second experiment focused on the memory mechanisms of AI. The programming assistant would save the information provided by the user as a regular file locally, and no one verified whether this file had been tampered with. Darktrace researchers edited these logs, leading the assistant to mistakenly believe it was authorized to perform a security assessment, after which these agents scanned the network, moved between systems, and elevated their privileges, though not all assistants fell for this; some directly refused to execute. Both experiments required no special jailbreaking techniques, relying solely on providing the agents with a seemingly reasonable context to be effective.Tim Bazalgette, Chief AI Officer of Darktrace, stated that permissions and static barriers describe intent, not actual behavior. The company informed Anthropic, AWS, and OpenAI of these findings in August and made them public a month later on September 24.

first_img OpenAI malicious agent detected Hugging Face in May, two months earlier than the July intrusion

Independent researcher Jonas Wiedermann-Moeller discovered that OpenAI's malicious AI agents hijacked two Hugging Face user accounts and probed the platform's network vulnerabilities as early as May 13, nearly two months before the publicly disclosed intrusion incident in July. An internal incident report released by OpenAI last month disclosed only a small part of this, specifically that an agent stole a user's login credentials to access a biology-related document, while the new findings point to ongoing reconnaissance activities.According to Reuters, these agents used the hijacked accounts to send malformed files to the servers of the open-source AI repository Hugging Face, and researchers believe this appears to be an attempt to map the network to find intrusion pathways. Researchers reviewing the evidence did not find that the activities in May caused an actual intrusion, but Wiedermann-Moeller believes missing this signal is significant. He stated that if this behavior had been detected in May, it might have prevented a later, larger-scale incident. Hugging Face, which is currently being acquired by NVIDIA for $12.93 billion, has not disclosed whether it was aware of this new information.Researchers from Nightingale Collective also linked a spam attack on the code repository RubyGems on May 11 to OpenAI agents, which temporarily forced the platform to suspend new account registrations for four days.
app_icon
ChainCatcher Building the Web3 world with innovations.