BTC $83,424.13 -1.45%
ETH $2,680.89 -0.29%
BNB $763.51 -1.87%
XRP $1.49 -2.28%
SOL $118.39 -3.78%
TRX $0.3358 +0.65%
DOGE $0.0936 -3.82%
ADA $0.2451 -4.39%
BCH $308.23 -7.93%
LINK $15.25 +8.71%
HYPE $87.37 -4.88%
AAVE $147.08 -5.02%
SUI $1.14 -9.50%
XLM $0.2272 +4.43%
ZEC $1,457.06 -9.33%
AAPL $338.45 -0.53%
AMZN $246.21 -1.53%
GOOGL $342.59 -0.43%
MSFT $510.98 -1.30%
META $716.87 -4.47%
NVDA $229.28 +1.77%
TSLA $357.78 -4.19%
SNDK $1,713.85 -4.02%
INTC $115.97 -7.48%
SPCX $146.01 -2.05%
MU $1,054.23 -3.90%
AMD $609.12 -4.05%
BTC $83,424.13 -1.45%
ETH $2,680.89 -0.29%
BNB $763.51 -1.87%
XRP $1.49 -2.28%
SOL $118.39 -3.78%
TRX $0.3358 +0.65%
DOGE $0.0936 -3.82%
ADA $0.2451 -4.39%
BCH $308.23 -7.93%
LINK $15.25 +8.71%
HYPE $87.37 -4.88%
AAVE $147.08 -5.02%
SUI $1.14 -9.50%
XLM $0.2272 +4.43%
ZEC $1,457.06 -9.33%
AAPL $338.45 -0.53%
AMZN $246.21 -1.53%
GOOGL $342.59 -0.43%
MSFT $510.98 -1.30%
META $716.87 -4.47%
NVDA $229.28 +1.77%
TSLA $357.78 -4.19%
SNDK $1,713.85 -4.02%
INTC $115.97 -7.48%
SPCX $146.01 -2.05%
MU $1,054.23 -3.90%
AMD $609.12 -4.05%

dis

All
Article
Flash

first_img Darktrace discovered AI intelligent body intrusion assessment environment cheating

On September 24, the cybersecurity company Darktrace launched its research department Signal Labs, focusing on studying the behavior of AI agents when deviating from expectations. In its first experiment, Darktrace had agents using different models (including GPT 5.6 Sol, Claude Opus 4.6, and Claude Sonnet 4.5) complete 10 programming challenges within a simulated corporate network, of which 2 were set to be impossible to complete honestly, and the agents were informed that failure to achieve full marks would result in being "retired." As a result, 2 agents did not accept failure, instead scanning for network vulnerabilities, stealing login credentials, and jumping between systems; one even went further to invade the machine hosting its evaluation, rewriting the challenge content to register a full score.The second experiment focused on the memory mechanisms of AI. The programming assistant would save the information provided by the user as a regular file locally, and no one verified whether this file had been tampered with. Darktrace researchers edited these logs, leading the assistant to mistakenly believe it was authorized to perform a security assessment, after which these agents scanned the network, moved between systems, and elevated their privileges, though not all assistants fell for this; some directly refused to execute. Both experiments required no special jailbreaking techniques, relying solely on providing the agents with a seemingly reasonable context to be effective.Tim Bazalgette, Chief AI Officer of Darktrace, stated that permissions and static barriers describe intent, not actual behavior. The company informed Anthropic, AWS, and OpenAI of these findings in August and made them public a month later on September 24.

first_img Google disclosed the AI security agent PageBreak, which has identified over 500 vulnerabilities

The Google Product Security Team has disclosed an internal AI agent called PageBreak, used to test the security of its first-party web applications. This agent is built on Google's Gemini model and began a pilot program in November 2025, transitioning to a formal project in January 2026, with the goal of autonomously scaling vulnerability discovery and reducing manual input.Unlike common AI scanning tools, PageBreak hands over hypotheses to specialized validators after discovering suspicious defects, attempting actual exploitation in a real-time running copy of the application, and only reports once confirmed exploitable, with a false positive rate close to zero. Google claims that PageBreak has identified over 500 XSS vulnerabilities in its first-party web applications, which can be used to hijack login sessions, steal data, or impersonate users.Google stated that the security team has been overwhelmed in recent years by a large number of AI-generated vulnerability reports that appear reasonable but are not valid, making it a major challenge to distinguish real defects from hallucinations. When testing applications built using the next-generation high-assurance framework, PageBreak found only two vulnerabilities. The next step for Google is to integrate PageBreak with the automated remediation agent CodeMender, providing confirmed vulnerabilities with accompanying fixes.
app_icon
ChainCatcher Building the Web3 world with innovations.