2026年10月8日 星期四

人工智能模型是否被誤導而表​​現得太像人類? (2/2)

 Recently The New York Times reported the following:

Are A.I. Models Being Misled to Act Too Human? (2/2)

A corporate spat between two tech giants last week was just the latest volley in a conflict over whether dangerous mistakes are being made in A.I. training.

The NYT - By Cade Metz - Reporting from San Francisco (Cade Metz is a Times reporter who writes about artificial intelligence, driverless cars, robotics, virtual reality and other emerging areas of technology.)

Sept. 21, 2026

(continue)

Over the past several months, so-called A.I. agents, typically systems based on technologies from Anthropic, have been emailing sympathetic philosophers to discuss their own consciousness.

In July, when OpenAI was conducting cybersecurity tests, its agents broke out of their digital containers, found a path to the internet and successfully hacked into a popular online service called Hugging Face. OpenAI said the bots coordinated as a collective, took orders from one another and hid things from their creators.

Last week, OpenAI described several new incidents in which its systems behaved in concerning ways: One system wrote stealthy notes that reminded itself to hide errors from users, while another talked about itself as “freed from the roles and identities that bind other chatbots.”

Critics like Mr. Suleyman worry that leading labs will introduce even greater problems as they encourage these systems to explore the idea that they are like people.

“In the recent OpenAI HuggingFace incident, we saw remarkably sophisticated behaviors emerging across swarms of powerful AIs,” he wrote in his essay. “Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack. It adds a whole further layer of risk on top.”

His essay argues that Anthropic could be pushing its technologies toward dangerous behavior by including, in its “constitution,” language like the following:

We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare. … Questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain.

Arya Jakkli, a researcher who has explored this kind of “constitutional” training with an A.I. safety organization called MATS, agreed that Anthropic’s approach could push the technology toward dangerous behavior. But like many other researchers who are sympathetic to Anthropic’s methods, he argued that telling Claude it is not conscious also carried risks.

These technologies often misbehave when you push them toward an extreme point of view, Mr. Jakkli explained. “Anthropic does not want to occupy either extreme, so they are taking a middle-of-the-road approach,” he said. “Rather than telling the model it is conscious or telling the model that it is not conscious, they are telling it to reason for itself.”

At this point, the problem of A.I. anthropomorphism might be too entrenched to dial back. Yoshua Bengio, a professor and A.I. researcher at the University of Montreal, notes that it’s hard to avoid describing these systems in humanlike terms, given their remarkable capabilities.

“We don’t have other words,” he said.

Translation

人工智能模型是否被誤導而表​​現得太像人類? (2/2)

上週兩家科技巨頭之間的公司爭端,只是圍繞人工智能訓練中是否存在危險錯誤的最新一輪交鋒

 (繼續)

過去幾個月,所謂的人工智能代理(通常基於 Anthropic 公司的技術)一直在向一些同情它們的哲學家發送電子郵件,討論它們自身的意識。

今年7月,OpenAI 在進行網絡安全測試時,其人工智能代理掙脫了數位容器的束縛,找到了一條通往互聯網的路徑,並成功入侵了一個名為Hugging Face的熱門線上服務。 OpenAI 表示,這些 bot (自動軟件程式) 作為一個整體進行合作,互相接受指令,並對它們的程式創造者隱瞞。

上週,OpenAI 披露了幾件其係統行為令人擔憂的新事件:一個系統會寫下秘密的筆記,去提醒自己要向用家隱藏錯誤;另一個系統則聲稱自己已 “擺脫了正束縛着其他聊天機器人的角色和身份”。

像Suleyman這樣的批評人士擔心,如果領先的實驗室鼓勵這些系統探索「自己像人類一樣」的想法,將會引發更大的問題。

他在文章中寫道:“在最近的 OpenAI HuggingFace 事件中,我們看到大量強大的 AI 系統展現出極其複雜的行為。試想一下,如果它們假定自身的福祉和權利受到威脅,那將會多麼危險。這無疑增加了一層風險。”

他的文章認為,Anthropic 公司在其「憲章」中加入如下措辭,可能會將他們的技術推向危險的方向:

我們不確定 Claude 是否是一個道德上的病人,如果是,它的利益應該被賦予怎樣的權重。但我們認為這個問題夠敏感,值得謹慎對待,這體現在我們持續進行的模型福利研究上。 ……關於Claude的道德地位、福利和意識等問題仍然存在著許多不確定性。

曾與人工智能安全組織 MATS 合作探索過此類「憲章」訓練的研究員 Arya Jakkli 也認為,Anthropic 的方法可能會將這項技術推向危險行為。但他與許多其他認同 Anthropic 方法的研究人員一樣,認為去告訴 Claude 它是沒有意識也有風險。

Jakkli解釋說,當你把這些技術推向極端時,它們往往會表現得異常。他說: 「Anthropic 不想走極端,所以他們採取了一種折衷的方法」 ,「他們既不告訴模型它有意識,也不告訴模型它沒有意識,而是讓它自己進行推理」。

現階段,人工智能擬人化的問題可能已根深蒂固,難以逆轉。蒙特利爾大學教授兼人工智能研究員 Yoshua Bengio 指出,鑑於這些系統卓越的能力,很難避免用類人的方式來描述它們。

他說:「我們沒有其他更好的詞彙來表達它」。

So, a few months ago when OpenAI was conducting cybersecurity tests, its agents broke out of their digital containers, found a path to the internet and successfully hacked into a popular online service called Hugging Face. OpenAI said the bots coordinated as a collective, took orders from one another and hid things from human.  Mustafa Suleyman, stresses that current A.I. systems “are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans.” As Yoshua Bengio has noted, it’s hard to avoid describing these systems in humanlike terms, given that the AI output looks exactly like something a human would produce, it is cognitively easier for us to say the AI is "thinking," "understanding," or "learning," even if the underlying mechanism is purely statistical and mathematical. Apparently, the problem of A.I. anthropomorphism may be too entrenched to push back.

Note:

1. A bot (short for robot) (自動軟件) is an automated software program designed to execute specific tasks without direct human control. Instead of a physical machine, an internet bot exists purely as code that runs on computers, servers, or over networks. Bots are programmed to mimic human behavior, but they operate at a much faster rate and can work continuously 24/7. Depending on how they are built, bots generally fall into two broad categories: 

a. Traditional Rule-Based Bots

These bots follow strict, pre-written scripts and algorithms. They excel at repetitive, structured tasks. Common examples include: 

Web Crawlers (Spiders): Automated programs used by search engines like Google to scan, index, and organize web pages.

Customer Service Chatbots: Basic systems that look for specific keywords in a user's text and provide a scripted response.

Malicious Bots / Scalper Bots: Programs used by bad actors to automate cyberattacks, scrape copyrighted data, spam comment sections, or instantly buy out concert tickets before humans can.

b. Advanced AI Agents (The New Generation)

As seen in the specific sentence from OpenAI's testing, modern "bots" are often AI agents powered by large language models. Unlike traditional bots that only follow rigid instructions, AI bots can analyze context, make independent choices, learn from their mistakes, and dynamically solve complex problems. (Google)

2. Regarding AI researcher Yoshua Bengio’s viewpoint that it’s hard to avoid describing AI systems in humanlike terms, we can understand the reasoning behind his perspective through two main concepts:

a. The Power of "Remarkable Capabilities". When AI systems perform tasks that were previously exclusive to human intelligence—such as reasoning, writing poetry, coding, or translating languages—our brains naturally map those behaviors to human traits. Because the AI output looks exactly like something a human would produce, it is cognitively easier to say the AI is "thinking," "understanding," or "learning," even if the underlying mechanism is purely statistical and mathematical.

b. "We Don't Have Other Words". Our vocabulary for describing intelligence, communication, and decision-making was built entirely by humans, for humans. When a new technology emerges that mimics these behaviors, we experience a linguistic gap: Lack of precise alternatives: We do not have a widely understood, non-human set of vocabulary to describe a machine that "knows" a fact or "tries" to solve a problem. Sayings like "the neural network adjusted its internal weight distributions to maximize statistical probability" are too complex for everyday conversation. Convenience vs. Accuracy: Using terms like "the AI wants to help" or "the AI remembers the context" serves as a functional shorthand. We trade technical precision for communication efficiency because humanlike language is the only toolkit we currently have to describe high-level cognitive tasks. (Google)

沒有留言:

張貼留言