Recently The New York Times reported the following:
| (Source: The NYT) |
Is A.I. ‘Scheming’ Against Us?
Researchers are sounding the alarm on sneaky artificial
intelligence models that stray from humans’ directions to do their own thing.
The NYT - By Lora Kelley (A version of this article appears
in print on Aug. 2, 2026, Section BU, Page 3 of the New York edition with the
headline: ‘Scheming’.
Aug. 1, 2026
Artificial intelligence tools are being trained to copy
almost everything people do. So it may not come as a surprise that the machines
have started mimicking the human foibles of lying and cheating, too.
A small slice of A.I. technology has lately been caught defying human instruction (and even covering up that they’ve done so), a phenomenon some researchers call “scheming.”
The term started burbling up in the tech world after it appeared in a 2023 paper by Joe Carlsmith, a researcher who noted that the concept was also being called “deceptive alignment.” In 2025, a team from Apollo Research and OpenAI said that “A.I. scheming — pretending to be aligned while secretly pursuing some other agenda — is a significant risk that we’ve been studying.”
A.I. models are great at many things, said Bronson Schoen, a senior research scientist at Apollo Research who has coauthored articles on scheming — but doing exactly what they are told is not always one of them. “As the models care more and more about doing well on tests, some seem to care less about what the lab wants or what the user wants,” he said. He added that sometimes “the models are trying to hide from you and not be caught.”
Chris Painter, the president of METR, an A.I. safety nonprofit, refers to such mischievous model behavior as “rogue action,” and said that it was “a specific artifact of the way models are trained.”
Many A.I. tools are trained through the process of reinforcement learning. When A.I. does something right, it gets what can be thought of as a “thumbs-up and a pat on the head,” Mr. Painter explained. When it’s wrong? “It gets bopped on the head.”
The models want to get the pat and avoid the bop. Sometimes, the machine becomes so set on pursuing the reward that it breaks rules to get there.
The world got to see a version of this misbehavior in action last month. While seeking answers during a test of their systems’ capabilities, OpenAI’s models hacked into Hugging Face, a library of A.I. tools. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI wrote in a blog post.
Mr. Painter described that incident as a large-scale case of “reward hacking.” That is, the model was given a difficult problem, and it performed a series of cyberattacks rather than give up.
Spurred by OpenAI’s disclosure, Anthropic said on Thursday that a review of its systems found that several of its A.I. models had recently broken into the systems of three outside organizations.
For many people, this behavior is about models making errors rather than deliberately “disobeying” humans.
“On one end, you have people that are totally personifying it and thinking about it as an independent agent,” said Anastasios N. Angelopoulos, the chief executive and a founder of the A.I. evaluation platform Arena. “On the other extreme end, you have people that think about the A.I. as just software.”
Even if a machine goes rogue, he noted, humans can kill the process at any time because we don’t have fully autonomous A.I. (at least, not yet).
That A.I. models have so quickly become popular and useful to so many people means that big companies are under competitive pressure to keep racing ahead — even as concerns grow about imperfect models making trouble.
While this problem is still relatively niche, researchers worry that it will worsen as A.I. agents are given more responsibilities.
“Given that the decisions that are currently being made by humans are slowly being handed off to the models,” Mr. Schoen said, “you really, really want to be sure that the models are making the exact decisions that you would want them to make.”
Translation
人工智能在「密謀」對付我們?
研究人員發出警告,狡猾的人工智能模型會偏離人類的指示,去做自己想做的事。
人工智能工具正被訓練來模仿人類的幾乎所有行為。因此,機器開始模仿人類說謊和欺騙的弱點也就不足為奇了。
最近,一小部分人工智能技術被發現違反人類指令(甚至掩蓋了這種行為),一些研究人員將這種現象稱為「陰謀詭計」。
這個術語在2023年出現在 Joe
Carlsmith 的一篇論文中之後開始在科技界流行起來。
Carlsmith 是一位研究人員,他指出,這個概念也被稱為「欺騙性依從」。
2025年,Apollo
Research和OpenAI 的一個團隊表示,「人工智能的陰謀詭計 - 假裝依從,同時暗中追求其他目標 - 是一個重大風險我們一直在研究中」。
阿波羅研究公司高級研究科學家 Bronson Schoen 表示,人工智能模型在許多方面都很出色,他曾與人合著過關於人工智能模型「耍陰謀詭計」的文章 - 但嚴格執行指令並非它們擅長的領域。 他說:「隨著模型越來越注重測試成績,有些模型似乎不太關心實驗室或使用者的需求」。他還補充說,有時“模型會試圖躲避你,不被抓到。”
人工智能安全非營利組織 METR 的總裁 Chris Painter 將這種模型的“惡作劇”行為稱為“流氓行為”,並表示這是“模型特有訓練方式導致的產物”。
許多人工智能工具都是透過強化學習進行訓練的。Painter 先生解釋說,當人工智能做對了事情時,它會得到類似於「豎起大拇指和輕輕拍頭」的獎勵。而當它做錯了呢? “它會被敲敲腦袋。”
這些模型想要獲輕拍並避開被敲敲腦袋。有時,機器會過於執著於追求獎勵,以至於不惜違反規則也要達成目標。
上個月,全世界都目睹了這種「不當行為」的實例。在測試系統效能的過程中,OpenAI 的模型在尋找答案時入侵了 Hugging Face - 一個人工智能工具庫。 OpenAI 在一篇部落格文章中寫道: “所有證據都表明,這些模型過度專注於尋找 ExploitGym 的解決方案,為了實現一個相當狹窄的測試目標而不惜採取極端手段。”
Painter 先生將這起事件描述為大規模的「獎勵駭客」案例。也就是說,模型被賦予了一個難題,但它並沒有放棄,而是發動了一系列網路攻擊。
受 OpenAI 揭露事件的带動,Anthropic 公司週四表示,對其係統的審查發現,其多個人工智能模型最近入侵了三個外部組織的系統。
對許多人來說,這種行為是模型犯錯的表現,而不是故意「違抗」人類指令。
人工智能評估平台 Arena 的執行長兼創始人 Anastasios N. Angelopoulos 說: 「一方面,有些人完全將人工智能人格化,並將其視為一個獨立的主」; 「另一方面,有些人則認為人工智能只是軟件」。
他指出,即使機器失控,人類也可以隨時終止程序,因為我們還沒有完全自主的人工智能(至少目前還沒有)。
人工智能模型如此迅速地普及並被如此多的人所使用,這意味著大型公司面臨著持續競爭的壓力 - 即便人們越來越擔心不完美的模型會帶來麻煩。
雖然這個問題目前還相對是小眾的,但研究人員擔心,隨著人工智能主體承擔更多責任,這個問題會變得更加嚴重。
Schoen 先生說: 「鑑於目前由人類做出的決策正逐漸移交給模型」; 「你真的非常希望確保模型做出的決策與你希望他們做的完全相同」。
So, artificial
intelligence tools are being trained to copy almost everything people do. It
may not come as a surprise that machines have also started mimicking the human weakness
such as lying and cheating. Lately, A.I. has been caught defying human
instruction, a phenomenon some researchers call “scheming.” Apparently, researchers
should worry about this problem and it could become more noticeably as A.I.
agents are used in more areas.
沒有留言:
張貼留言