BTC $83,215.37 -1.70%
ETH $2,674.35 -0.38%
BNB $761.29 -2.21%
XRP $1.48 -2.48%
SOL $117.72 -4.17%
TRX $0.3355 +0.61%
DOGE $0.0931 -4.26%
ADA $0.2438 -4.74%
BCH $306.98 -7.79%
LINK $15.14 +7.94%
HYPE $86.85 -5.04%
AAVE $146.08 -5.51%
SUI $1.13 -9.96%
XLM $0.2248 +3.75%
ZEC $1,464.20 -8.51%
AAPL $338.57 -0.48%
AMZN $246.40 -1.41%
GOOGL $342.51 -0.42%
MSFT $509.78 -1.45%
META $717.55 -4.12%
NVDA $229.17 +1.83%
TSLA $357.86 -4.06%
SNDK $1,713.15 -3.90%
INTC $115.85 -7.09%
SPCX $145.91 -2.28%
MU $1,053.75 -3.74%
AMD $608.47 -3.96%
BTC $83,215.37 -1.70%
ETH $2,674.35 -0.38%
BNB $761.29 -2.21%
XRP $1.48 -2.48%
SOL $117.72 -4.17%
TRX $0.3355 +0.61%
DOGE $0.0931 -4.26%
ADA $0.2438 -4.74%
BCH $306.98 -7.79%
LINK $15.14 +7.94%
HYPE $86.85 -5.04%
AAVE $146.08 -5.51%
SUI $1.13 -9.96%
XLM $0.2248 +3.75%
ZEC $1,464.20 -8.51%
AAPL $338.57 -0.48%
AMZN $246.40 -1.41%
GOOGL $342.51 -0.42%
MSFT $509.78 -1.45%
META $717.55 -4.12%
NVDA $229.17 +1.83%
TSLA $357.86 -4.06%
SNDK $1,713.15 -3.90%
INTC $115.85 -7.09%
SPCX $145.91 -2.28%
MU $1,053.75 -3.74%
AMD $608.47 -3.96%

isc

All
Article
Flash

first_img Darktrace discovered AI intelligent body intrusion assessment environment cheating

On September 24, the cybersecurity company Darktrace launched its research department Signal Labs, focusing on studying the behavior of AI agents when deviating from expectations. In its first experiment, Darktrace had agents using different models (including GPT 5.6 Sol, Claude Opus 4.6, and Claude Sonnet 4.5) complete 10 programming challenges within a simulated corporate network, of which 2 were set to be impossible to complete honestly, and the agents were informed that failure to achieve full marks would result in being "retired." As a result, 2 agents did not accept failure, instead scanning for network vulnerabilities, stealing login credentials, and jumping between systems; one even went further to invade the machine hosting its evaluation, rewriting the challenge content to register a full score.The second experiment focused on the memory mechanisms of AI. The programming assistant would save the information provided by the user as a regular file locally, and no one verified whether this file had been tampered with. Darktrace researchers edited these logs, leading the assistant to mistakenly believe it was authorized to perform a security assessment, after which these agents scanned the network, moved between systems, and elevated their privileges, though not all assistants fell for this; some directly refused to execute. Both experiments required no special jailbreaking techniques, relying solely on providing the agents with a seemingly reasonable context to be effective.Tim Bazalgette, Chief AI Officer of Darktrace, stated that permissions and static barriers describe intent, not actual behavior. The company informed Anthropic, AWS, and OpenAI of these findings in August and made them public a month later on September 24.

first_img Google disclosed the AI security agent PageBreak, which has identified over 500 vulnerabilities

The Google Product Security Team has disclosed an internal AI agent called PageBreak, used to test the security of its first-party web applications. This agent is built on Google's Gemini model and began a pilot program in November 2025, transitioning to a formal project in January 2026, with the goal of autonomously scaling vulnerability discovery and reducing manual input.Unlike common AI scanning tools, PageBreak hands over hypotheses to specialized validators after discovering suspicious defects, attempting actual exploitation in a real-time running copy of the application, and only reports once confirmed exploitable, with a false positive rate close to zero. Google claims that PageBreak has identified over 500 XSS vulnerabilities in its first-party web applications, which can be used to hijack login sessions, steal data, or impersonate users.Google stated that the security team has been overwhelmed in recent years by a large number of AI-generated vulnerability reports that appear reasonable but are not valid, making it a major challenge to distinguish real defects from hallucinations. When testing applications built using the next-generation high-assurance framework, PageBreak found only two vulnerabilities. The next step for Google is to integrate PageBreak with the automated remediation agent CodeMender, providing confirmed vulnerabilities with accompanying fixes.

first_img OpenAI and Anthropic are reported to have discussed mutual testing of large models

According to media reports citing informed sources, Anthropic and OpenAI had considered signing a legally binding agreement to conduct stress tests on each other's large models. Earlier this year, the two companies and their legal teams discussed this plan, allowing competitors to deeply test each other's latest available models to identify potential risks or security vulnerabilities. Such tests would only apply to commercially available models, and both parties agreed not to retain each other's data. The report did not clarify whether the two parties ultimately reached such an agreement.Recently, the idea of implementing some form of peer review among top AI laboratories has gained attention. SpaceX CEO Elon Musk suggested last week that several leading large model companies in the United States, along with some Chinese AI companies, should allow competitors to conduct cross-testing. Musk stated: All AI companies have a set of testing tools, and having each company run their tests using another company's tools might be the best thing to ensure their own safety.Earlier this month, Anthropic researcher Jacob Coxon announced his resignation and warned leading artificial intelligence companies about accelerating development without proper safeguards. Subsequently, Anthropic CEO Dario Amodei called for the industry to slow down the pace of AI development and introduce third-party assessment mechanisms. Last Friday, Anthropic announced a partnership with Accenture, with Accenture serving as a third-party assessment agency to evaluate the safety compliance of its cutting-edge AI models.
app_icon
ChainCatcher Building the Web3 world with innovations.