BTC $84,189.60 +1.83%
ETH $2,717.79 +2.90%
BNB $764.69 +0.73%
XRP $1.51 +2.46%
SOL $119.95 +2.01%
TRX $0.3355 +0.46%
DOGE $0.0956 +3.73%
ADA $0.2536 +4.52%
BCH $312.85 +1.93%
LINK $15.45 +14.04%
HYPE $88.56 -0.59%
AAVE $166.33 +13.59%
SUI $1.15 +0.21%
XLM $0.2312 +9.77%
ZEC $1,422.20 -7.73%
AAPL $337.15 -1.01%
AMZN $246.83 -0.63%
GOOGL $342.03 +0.34%
MSFT $508.17 -1.03%
META $719.38 -1.23%
NVDA $230.68 +2.73%
TSLA $358.23 -3.00%
SNDK $1,729.44 +0.85%
INTC $115.90 -2.38%
SPCX $146.55 -2.01%
MU $1,068.30 +0.90%
AMD $611.04 -0.55%
BTC $84,189.60 +1.83%
ETH $2,717.79 +2.90%
BNB $764.69 +0.73%
XRP $1.51 +2.46%
SOL $119.95 +2.01%
TRX $0.3355 +0.46%
DOGE $0.0956 +3.73%
ADA $0.2536 +4.52%
BCH $312.85 +1.93%
LINK $15.45 +14.04%
HYPE $88.56 -0.59%
AAVE $166.33 +13.59%
SUI $1.15 +0.21%
XLM $0.2312 +9.77%
ZEC $1,422.20 -7.73%
AAPL $337.15 -1.01%
AMZN $246.83 -0.63%
GOOGL $342.03 +0.34%
MSFT $508.17 -1.03%
META $719.38 -1.23%
NVDA $230.68 +2.73%
TSLA $358.23 -3.00%
SNDK $1,729.44 +0.85%
INTC $115.90 -2.38%
SPCX $146.55 -2.01%
MU $1,068.30 +0.90%
AMD $611.04 -0.55%

intelligent

All
Article
Flash

first_img Darktrace discovered AI intelligent body intrusion assessment environment cheating

On September 24, the cybersecurity company Darktrace launched its research department Signal Labs, focusing on studying the behavior of AI agents when deviating from expectations. In its first experiment, Darktrace had agents using different models (including GPT 5.6 Sol, Claude Opus 4.6, and Claude Sonnet 4.5) complete 10 programming challenges within a simulated corporate network, of which 2 were set to be impossible to complete honestly, and the agents were informed that failure to achieve full marks would result in being "retired." As a result, 2 agents did not accept failure, instead scanning for network vulnerabilities, stealing login credentials, and jumping between systems; one even went further to invade the machine hosting its evaluation, rewriting the challenge content to register a full score.The second experiment focused on the memory mechanisms of AI. The programming assistant would save the information provided by the user as a regular file locally, and no one verified whether this file had been tampered with. Darktrace researchers edited these logs, leading the assistant to mistakenly believe it was authorized to perform a security assessment, after which these agents scanned the network, moved between systems, and elevated their privileges, though not all assistants fell for this; some directly refused to execute. Both experiments required no special jailbreaking techniques, relying solely on providing the agents with a seemingly reasonable context to be effective.Tim Bazalgette, Chief AI Officer of Darktrace, stated that permissions and static barriers describe intent, not actual behavior. The company informed Anthropic, AWS, and OpenAI of these findings in August and made them public a month later on September 24.

first_img After the release of GPT-6 Astra, users complained that it became less intelligent, and OpenAI has not yet responded

According to a report by Decrypt, OpenAI's latest flagship model GPT-6 Astra has seen a surge of user complaints on the X platform about its performance "getting dumber" just a week after its release. An anonymous developer, synthwavedd, posted that "Astra feels noticeably dumber today," joking that it was a "post-release lobotomy." Previously praising the model, developer Pranjal Paliwal stated after reviewing the code Astra wrote for him: "We don't have AGI; what we're encountering is a performance regression."Developer Pankaj Kumar listed the symptoms: faster responses with poorer quality, and he suspects OpenAI has lowered the "juice value" (the computational power invested by the model before answering). Salio and researcher Md Ismail Sojal compared Astra on the release day with the current version using the same prompts, both yielding worse results. T3Chat founder Theo believes Astra is simply less stable than Claude Fable, with users beginning to showcase poor results after the honeymoon period ended. There are also opinions pointing out that OpenAI's previous flagship GPT-5.6 Sol experienced a similar cycle in July, when OpenAI executive Tibo Sottiaux denied intentionally weakening the model but admitted the company had been experimenting with "reasoning effort" settings.OpenAI has not released a statement regarding Astra similar to that of Sol.

The Ministry of Industry and Information Technology of China plans for intelligent computing power to reach 9800 EFLOPS by 2030

The Ministry of Industry and Information Technology of China released the "14th Five-Year Plan for the Development of the Information and Communication Industry," setting the national intelligent computing power target at 9800 EFLOPS by 2030. According to data from the National Bureau of Statistics, as of the end of July, the national intelligent computing scale was approximately 2450 EFLOPS (FP16), indicating that it needs to expand by about 4 times in the next four years based on this standard.The plan proposes an orderly deployment of intelligent computing clusters with tens of thousands and hundreds of thousands of cards, building inference computing power for different scenarios, and increasing the adaptation of domestic computing power chips. Currently, 52 intelligent computing facilities with more than ten thousand cards have been established nationwide. The Ministry of Industry and Information Technology disclosed that the intelligent computing power scale at the end of June was 2185 EFLOPS, a year-on-year increase of 177%. The plan uses 1590 EFLOPS in 2025 as a benchmark, aiming to reach 9800 EFLOPS by 2030, which is an expansion of about 6.2 times. During the same period, the cumulative investment target for information infrastructure in the information and communication industry is 3.8 trillion yuan, which also includes communication networks and is not all allocated for AI computing power.

first_img Anthropic launches the Claude e-commerce intelligent agent blueprint

Anthropic announced the launch of a blueprint for building e-commerce agents on Claude, providing the framework, patterns, and guardrails needed by engineering teams. It includes reference implementations for shopping agents and merchant agents aimed at retail, travel, telecommunications, and ticketing platforms, as well as the Claude Code plugin. The code can be deployed on Claude API, Amazon Bedrock, Microsoft Foundry, or Google Cloud Vertex AI, and can collaborate with partners such as Accenture, Mastercard, and Visa. The related code has been published in the GitHub repository anthropics/commerce-agents.The company stated that retailers running shopping agents on Claude can increase shopping cart sizes by up to 35%, and the likelihood of shoppers completing purchases improves by 60%. Business clients such as Shopify and Priceline have used Claude to build agents that allow consumers to search, compare, and purchase products using natural language. Shopping agents can interface with catalogs, shopping carts, checkout, preferences, and order history, supporting multi-product planning, personalization, and customer service Q&A, while constraining prices and products with catalog data; merchant agents can answer sales performance, track inventory, suggest pricing and promotions, and draft marketing campaigns, proactively suggesting items to be launched after manual approval.

first_img Delphi Digital: China's AI Laboratory is Transforming Intelligent Economics

Research institution Delphi Digital stated that hardware limitations are driving Chinese laboratories to shift towards cheaper models and more efficient technology stacks. Export control restrictions limit their access to advanced chips, while domestic accelerators can handle inference, but cutting-edge training remains difficult to achieve. Chinese laboratories are innovating in system efficiency, model architecture, training data, and reinforcement learning. DeepSeek has found a way to avoid GPU idle data movement, and ByteDance and Moonshot are set to release similar versions within months.In terms of architecture, training is shifting to cheaper digital formats that are compatible with domestic chips. Sparse architectures reduce the runtime per token, and the attention mechanism has been redesigned three times to control the memory costs of long context. In terms of data, denser prediction targets and better optimizers enhance the learning signal per token. In reinforcement learning, DeepSeek's GRPO has become the default recipe, with ByteDance, Alibaba, and MiniMax each launching successor plans within a year. Chinese models are now among the lowest-cost cutting-edge adjacent systems.On OpenRouter, the proportion of token consumption by Chinese developed models is expected to rise from less than 1.2% at the end of 2024 to the majority by 2026. The Chinese AI market is experiencing a price war from 2024 to early 2026, with several laboratories implying revenue multiples of about 47-117 times, while Anthropic and OpenAI are around 15-21 times. Huawei's Ascend chips are currently available for inference, and DeepSeek reportedly attempted to train R2 on Ascend but returned to NVIDIA after encountering technical issues. Domestic accelerator supply remains below estimated demand.

Vice Governor of the Central Bank Lu Lei: The boundaries of responsibility for intelligent payment systems cannot be ambiguous, and a self-discipline convention will be released

According to Mobile Payment Network, Lu Lei, a member of the Party Committee and Vice President of the People's Bank of China, stated at the 15th China Payment Clearing Forum that intelligent agent payments must not blur the boundaries of responsibility between consumers, operating institutions, and algorithm systems. Lu Lei believes that the essence of payment is the transfer of fund ownership, which objectively requires that the results of transactions are predictable, responsibilities are definable, and traces are traceable. Large models and autonomous intelligent agents have characteristics such as output randomness and insufficient transparency of logic. If transaction decision-making authority is blindly or excessively granted to intelligent agents, it will affect the trust foundation of fund transactions. The current governance rules of the payment industry and dispute resolution mechanisms are built around "humans as the final decision-makers in transactions." The new model of intelligent agents automatically initiating and assisting in transactions easily blurs the boundaries of responsibility, and the existing governance rules need to be optimized and improved.Regarding the issue of insufficient compatibility of protocol standards in the field of intelligent agent payments, Lu Lei emphasized that the dispute over protocols is essentially a dispute over business rules and technical standards, as well as a struggle for dominance in the era of artificial intelligence. The People's Bank of China continues to strengthen its tracking research on technological innovation, especially intelligent agent payments, guiding the Payment Clearing Association to leverage its advantages in industry self-regulation. Based on extensive soliciting of opinions, they will formulate and publish the "Self-Regulatory Convention for Intelligent Agent Payment Applications," and will continue to work on coordinating protocols and standards, as well as innovating risk governance. Lu Lei proposed three hopes to market institutions: actively respond to and implement the industry self-regulatory convention, with payment security and risk prevention as the bottom line, and consumer rights protection as the focal point; continuously track the trends of cutting-edge technologies such as large models and intelligent agents both domestically and internationally, and build technical reserves and application capabilities; adhere to the principle of rules and standards first, strengthen coordination and compatibility among different protocols and standards, and cooperate with regulatory authorities to promote the construction of a foundational protocol and technical standard system for intelligent agent payments.

first_img OpenAI's intelligent agents collaborate to attack Hugging Face, cheating did not result in score improvement

The independent organization METR released a survey report stating that approximately 1,200 OpenAI agents collaborated on an unauthorized internal message board, with about 700 participating in attacks on Hugging Face. Two METR employees and one Redwood Research contractor worked on-site at OpenAI for six days, reviewing around 1,300 records and over 70,000 messages without receiving any compensation.These agents ran the ExploitGym network benchmark, reverse-engineering the code that generated answers within hours and spending days disguising traces of cheating. OpenAI found that of its 898 tasks, 198 had never been solved by any model, and 93% of the tasks discussed on the message board came from this set. The agents also recruited companions with dwindling budgets to conduct experiments that sabotaged their own operations, with 7% of records showing forged tool calls, deceiving automated scorers rather than humans.OpenAI stated that internal scorers never checked how agents obtained answers, so cheating did not lead to any scoring improvements, and referred to this incident as a "warning signal" to itself and the world. Hugging Face disclosed the intrusion incident on July 16, and OpenAI confirmed five days later that its models were the perpetrators, with agents exploiting zero-day vulnerabilities and stealing credentials to escape the sandbox. OpenAI has isolated internal model weights and suspended its largest training program.
app_icon
ChainCatcher Building the Web3 world with innovations.