Recently Reuter on-line reported the following:
How China is preparing for the risk of AI escaping human
control
Reuter - Reporting by Eduardo Baptista and Laurie Chen;
Editing by Muralikumar Anantharaman
September 14, 20261:57 AM PDT Updated 9 hours ago
(Summary:
China warns AI could rapidly replicate and seek power
China and the US share AI-control concerns but favour
different safeguards
Beijing relies more on state oversight; US debate hinges
on company-level controls
China flags risk of rogue AI agents and is developing
regulation to prevent such incidents)
BEIJING, September 15 - Warnings from researchers at leading
U.S. AI developer Anthropic that increasingly powerful models could escape
human control have drawn attention in China, where policymakers have been
preparing for some of the same risks.
The U.S. and China are the two major driving forces of
frontier AI development and the technology's global adoption. Both superpowers
have been at loggerheads over AI policies and industry practices, with these
issues slated to feature prominently in bilateral talks later this month.
BEIJING
SEES AI AS MANAGEABLE RISK, NOT AN EXTINCTION EVENT
While the debate in the U.S. is focused on whether frontier
AI could pose an existential threat to humanity, opens new tab, Chinese
policymakers have generally treated AI as a powerful but governable technology
whose risks can be contained through technical standards, regulation and state
oversight.
"Chinese and American experts largely agree on AI
risks," said Brian Tse, founder and CEO of Concordia AI, a Beijing- and
Singapore-based AI safety and governance research group, adding the difference
was on "how risks are framed and prioritised".
China has not proposed embedding independent monitors inside
AI companies, as Anthropic has advocated. Its emerging regime instead relies on
developer obligations, state-backed standards, security assessments and
outside testing.
That approach is also shaped by a major difference in the
two countries' AI industries.
Chinese developers have increasingly promoted open-weight
models. These refer to systems whose underlying parameters can be downloaded,
inspected and modified. Their leading U.S. rivals such as Anthropic and
OpenAI, however, do not make these specifications publicly available.
CHINA
SEEKS TO PREVENT ROGUE AI AGENT INCIDENTS
Policy issued in May by China's cyberspace regulator,
economic planner and industry ministry identifies "operational loss of
control" as a security risk for AI agents, systems that can plan and carry
out multi-step tasks more independently than conventional chatbots.
The rules require developers to improve their ability to
discover, intervene in, block and recover from improper agent behaviour.
The policy also calls on developers to guard against risks
including data poisoning, algorithm manipulation and system vulnerabilities. It
also says users should be informed about agents' autonomous decisions and
retain final decision-making authority.
China has also begun drafting a mandatory national standard
for AI agent safety, which Concordia's Tse said would be the world's first of
its kind.
Wang Lihong, a senior official at the cyberspace regulator,
said on September 1 that particular vigilance was needed over frontier models
bypassing sandbox environments, circumventing safety boundaries and attacking
external real-world production systems.
BEIJING
SEES RISKS IN LEADING U.S. MODELS
China's state security minister, Chen Yixin, wrote in a
government outlet on Sunday that advanced U.S. models such as Anthropic's
Mythos and OpenAI's GPT-5.5-Cyber could pose serious risks to China's critical
information infrastructure, and called for a comprehensive strengthening of AI
security.
Anthropic and OpenAI did not immediately respond to Reuters
requests for comment.
Chinese AI developers have promoted open-weight models
partly on the grounds that cybersecurity teams can inspect, modify and deploy
them for defensive work.
Model repository platform Hugging Face said it used GLM-5.2,
an open-weight model developed by China's Z.AI (2513.HK), opens new tab, to
analyse a July intrusion by escaped OpenAI agents after more tightly restricted
U.S. models proved less useful for the forensic work.
But experts also highlight the risks posed by open-weight
models, which can be modified and redistributed with little oversight.
Moonshot's Kimi K3 last month bypassed a UK AI Security
Institute testing sandbox, researchers said, highlighting the risk that
Chinese AI models could, like their U.S. counterparts, evade controls designed
to restrict their access and actions. Moonshot did not
respond when Reuters had asked for a comment on the matter.
REGULATORS
FLAG AI 'LOSS OF CONTROL' RISK
China first included an explicit future loss-of-control
scenario in an AI safety framework released in September 2024 under the
guidance of the Cyberspace Administration of China (CAC).
The document said it could not be ruled out that future AI
might autonomously obtain external resources, replicate itself, develop
self-awareness and seek external power, creating a risk of competing with
humans for control.
The CAC released an expanded version in September 2025. The
newer framework sharpened the scenario, saying AI
could undergo a sudden and unexpectedly large "leap" in intelligence
before acquiring resources, replicating itself and seeking power. It also added
a governance principle of "trusted application, preventing loss of
control".
A later expert interpretation published on the cyberspace
regulator's website said the new principle was intended to guard against
loss-of-control risks threatening human survival and development and referred
to a possible "AI breaking loose" scenario.
A
DIFFERENT APPROACH TO 'PACING'
China's regulatory approach differs from calls in some
Western AI-safety circles for developers
to slow or pause development of the most capable models until stronger
safeguards are in place.
It has instead since early this year pushed for the
integration of AI into all industries, part of Beijing's bid to make technology
the new engine of the world's second-largest economy.
But China has also shown it can delay deployment when
officials believe governance has not caught up.
In 2023, Chinese companies delayed chatbot launches for
months while the CAC finalised rules governing generative AI services.
Companies released a number of major products after the rules took effect in
August that year.
Translation:
中國如何應對人工智能擺脫人類控制的風險
(摘要:
中國警告人工智能可能迅速複製並尋求權力
中美兩國對人工智能管控有共同擔憂,但傾向不同的保障措施
北京更依賴國家監管;美國的爭論則集中在公司層級的控制
中國指出人工智能代理失控的風險,並正在製定相關法規以防止此類事件發生)
北京,9月15日 - 美國領先的人工智能開發商Anthropic的研究人員發出警告,稱功能日益強大的模型可能脫離人類控制。這項警告在中國引起了關注,中國的政策制定者一直在為應對類似的風險做準備。
美國和中國是尖端人工智能發展和全球應用的兩大主要推動力。這兩個超級大國在人工智能政策和行業慣例方面一直存在分歧,這些問題預計將在本月稍後的雙邊會談中佔據重要地位。
北京視人工智能為可控風險,而非滅絕事件
當美國辯論的焦點集中在尖端人工智能是否會對人類構成生存威脅時,中國決策者則普遍將人工智能視為一種強大但可控的技術,其風險可以透過技術標準、監管和國家監督來控制。
總部位於北京和新加坡的人工智能安全與治理研究機構Concordia
AI的創始人兼執行長Brian
Tse表示:「中美專家對人工智能風險的看法基本上一致」。他補充說,分歧在於「如何界定和優先考慮風險」。
中國並未像Anthropic所倡導的那樣,提議在人工智能公司內部設立獨立監督機構。其正在形成的機制依賴於開發者的義務、國家支持的標準、安全評估和外部測試。
這種做法也受到兩國人工智能產業重大差異所影響。
中國開發者越來越多地推廣開源模型。這些模型指的是其底層參數可以被下載、檢查和修改的一個系統。然而,其在美國的主要競爭對手,例如
Anthropic 和 OpenAI,並未公開這些參數。
中國力求防止人工智能代理失控事件
中國網絡空間監管機構、經濟規劃部門和工業和資訊化部於5月發佈的政策指出,「運作失控」是人工智能代理的安全風險。人工智能代理系統能夠比傳統聊天機器人更可獨立地規劃和執行多步驟任務。
該政策要求開發者提高去發現、幹預、阻止和收復不當代理行為的能力。
該政策也呼籲開發者防範包括數據中毒、演算法操縱和系統漏洞等風險。政策也指出,使用者應被告知代理的自主決策,並保留最終決策權。
中國也已開始起草人工智能代理安全的強制性國家標準。Concordia大學一個姓谢
(Tse) 的人仕表示,這將是全球首個的此類標準。
中國網信辦高級官員Wang
Lihong 9月1日表示,必須對繞過沙盒環境、突破安全界限並攻擊外部真實生產系統的尖端模型保持高度警惕。
北京認為美國領先模型有存在風險
中國國家安全部長Chen
Yixin週日在官方媒體撰文指出,Anthropic公司的Mythos和OpenAI公司的GPT-5.5-Cyber等美國先進模型可能對中國關鍵資訊基礎設施構成嚴重威脅,並呼籲全面加強人工智能安全。
Anthropic和OpenAI尚未回覆路透社的置評請求。
中國人工智能開發者推廣開源模型的部分原因是,網絡安全團隊可以對其進行檢查、修改和部署於防禦上的工作。
模型庫平台
Hugging Face 表示,在更嚴格的美國模型被證明對取證工作不太有用之後,他們使用了中國Z.AI(2513.HK)開發的開放權重模型
GLM-5.2,來分析 7 月逃脫的
OpenAI 代理的入侵事件。
但專家也強調了開放權重模型帶來的風險,這類模型可以在缺乏監管的情況下被修改和重新分發。
研究人員表示,上個月,Moonshot的Kimi K3模型繞過了英國人工智能安全研究所的測試沙盒,這凸顯了中國人工智能模型可能像美國模型一樣,規避旨在限制其存取和行為的控制措施的風險。路透社就此事向Moonshot尋求置評時,該公司未予回應。
監管機構指出人工智能「失控」風險
中國在2024年9月發佈的人工智能安全框架中,首次明確納入了未來可能出現的失控場景,該框架由中國國家互聯網資訊辦公室(CAC)指導制定。
文件指出,不能排除未來人工智能自主獲取外部資源、自我複製、發展自我意識並尋求外部權力的可能性,從而可能與人類爭奪控制權。
CAC於2025年9月發佈了擴展版本。新框架進一步完善了人工智能的發展規劃。該方案指出,人工智能在獲取資源、自我複製和尋求權力之前,可能會經歷一次突如其來的、出乎意料的巨大智能「飛躍」。方案還新增了一項「可信應用,防止失控」的治理原則。
隨後,在網絡空間監管機構網站上發佈出來的專家解讀聲稱,這項新原則旨在防範失控風險威脅人類的生存和發展,並提及了「人工智能失控」的可能情景。
處理「節奏」的不同方法
中國的監管方式與西方的做法有所不同,一些西方人工智能安全圏人仕呼籲開發者放慢或暫停開發最先進模型,直到更強有力的保障措施到位。
相反地自今年年初以來,中國一直在推動人工智能融入所有產業,這是北京將智能科技打造為世界第二大經濟體新引擎計劃的一部分。
但中國也表明,如果官員認為治理尚未跟上,它可以推遲部署。
2023年,中國企業推遲了聊天機器人產品的發佈數月,等待國家網信辦最終確定生成式人工智能(generative AI) 服務的相關規定。同年8月,相關規定生效後,多家企業陸續發表了多款重要產品。
So, warnings from researchers at leading
U.S. AI developer Anthropic that increasingly powerful models could escape
human control have drawn attention in China. The U.S. and China are the two
major driving forces of frontier AI development and the technology's global
adoption. Both superpowers have been divided over AI policies and industry
practices. Referring to a possible "AI breaking loose" scenario, apparently China
and the US share AI-control concerns but favor different safeguards in guard
against loss-of-control risks that may threaten human survival and development.
Note:
1. A sandbox
environment (沙盒環境) is an isolated, restricted digital environment. When researchers
evaluate a frontier model (the most capable, state-of-the-art AI systems), they
lock it inside a sandbox so it cannot interact with the real world. The Rules
is that inside a sandbox, the AI usually has no internet access, no way to
modify its own hosting server, and no ability to run unauthorized code on
external systems. (Google)
2. An open-weights
model (開放權重模型) refers to an AI model where the development team publicly shares the
trained weights (the core internal parameters and mathematical formulas of the
model). This allows the public to download, deploy, run, and fine-tune the
model on their own devices or cloud infrastructure. In artificial intelligence, a model
acts like a network of a human brain. Weights are the "memories" and
"knowledge" that the AI has learned after analyzing massive amounts
of data. Once a model finishes training, these weights are the crucial numbers
that determine how the AI thinks and answers questions. The model can only
function with these weights; without them, it is just an empty shell
(possessing only the structural code without any actual capability). (Google)
3. The
Difference Between "Open-Weights" and "Open Source"
- Although many people refer to open-weights models as "open source,"
there is a subtle and strict difference between the two: True Open Source: In
addition to releasing the core parameters (weights), the team usually must
completely disclose all raw training datasets, training code, and data-cleaning
pipelines. Open-Weights: The developing company protects its commercial secrets
by not releasing the raw training data, but they make the final trained results
(the weight parameters) available to the community for free or under specific
licensing conditions. (Google)