2026年10月10日 星期六

中國真的在竊取美國公司的人工智能技術嗎?

Recently The New York Times reported the following:

Is China Really Stealing A.I. From American Companies?

As Xi Jinping visits President Trump in Washington, the leaders are expected to discuss claims that China is surreptitiously copying American A.I. technologies, a process known as distillation.

The NYT - By Cade MetzReporting from San Francisco  (Cade Metz is a Times reporter who writes about artificial intelligence, driverless cars, robotics, virtual reality and other emerging areas of technology.)

Sept. 25, 2026Updated 11:02 a.m. ET

For nearly two years, companies like Anthropic, OpenAI and Google have accused Chinese companies of unfairly copying their artificial intelligence technologies. As the Chinese leader Xi Jinping visits President Trump this week in Washington, these claims have become a flashpoint in their discussions.

Amid growing concern over the dangers of A.I., the leaders are expected to discuss the rapid rise of the technology, with many U.S. officials, tech executives and everyday people calling for a slowdown. American companies like Anthropic and OpenAI are leading the A.I. race, but Chinese start-ups like Z.ai and Moonshot are not far behind.

One of the more contentious issues on the table is Silicon Valley’s repeated claim that Chinese companies are copying the leading American technologies using a technique called “distillation.”

Many experts say the claims from companies like Anthropic and OpenAI are overblown. “The narrative that all of the capabilities of the Chinese technologies coming from an Anthropic model is not as true as people say it is,” said Charles O’Neill, a head of model training at Baseten, a U.S. company that sells access to Chinese technologies.

Here is a guide to a niche technical process that has now become a geopolitical issue.

What exactly is distillation?

First developed in the early 2010s, distillation was initially a way for A.I. researchers to build more efficient technologies. Using this technique, they could collect data from an established system and use that data to build a new system able to run on less expensive hardware.

“Think of one model as the teacher and the other as a student,” Geoffrey Hinton, a former Google researcher who helped develop the technique, recently told The Times.

Companies still use distillation in this way. But more recently, the technique has also become a way for companies to collect data from technologies built by others.

Basically, an outside company looking to distill a competitor’s system will set up many accounts on the competitor’s service, say Anthropic’s Claude or OpenAI’s GPT. The distilling company will then assign those systems to particular tasks, such as computer coding or mathematics.

The distiller can then use the information generated by these established A.I. systems to help train a new system of their own.

Is distillation illegal?

Some legal scholars say that distillation could violate the Defend Trade Secrets Act, a law that allows companies to sue over the theft of trade secrets. But U.S. courts have not decided this issue.

Copyright law does not necessarily apply, because distillation is an effort to copy the general behavior of the system rather than to copy its text word for word.

Critics of Anthropic and OpenAI point out that the two companies have engaged in similar behavior.

Anthropic is facing multiple lawsuits that accuse the San Francisco start-up of illegally using copyrighted internet data to train its systems. Last year, as part of a landmark legal settlement, the company agreed to pay $1.5 billion to a group of authors and publishers after a judge ruled that it had illegally downloaded and stored millions of copyrighted books. The settlement was the largest payout in the history of U.S. copyright cases.

OpenAI and other companies face similar suits, including one brought by The New York Times in 2023. That suit contends that OpenAI and its partner Microsoft used millions of articles published by The Times to train chatbots, which now compete with the news outlet as a source of information. OpenAI and Microsoft deny the claims.

Are Chinese companies really copying American systems?

Chinese firms have most likely distilled proprietary American systems. But it is not completely clear what they have done.

After the Chinese start-up DeepSeek shocked American researchers and investors last year by releasing a particularly powerful and efficient system, OpenAI accused DeepSeek of distilling its technologies.

As Chinese systems continued to improve over the past nine months, Anthropic made similar claims. In June, Anthropic sent a letter to Senators Tim Scott and Elizabeth Warren, accusing the Chinese tech giant Alibaba of distilling its technologies.

The Chinese companies have not responded publicly to the accusations.

Does distillation alone explain how China’s models remain competitive?

No. When companies distill a proprietary system, they only have access to part of it: the words and characters and other data generated by the system. They do not have access to the system’s underlying computer code, as they would with an open source technology.

Using outputs from a proprietary system like Claude Fable from Anthropic , a Chinese company can accelerate the development of its own technology. But those outputs are only part of the puzzle.

“You are not getting access to everything you would use when you are distilling your model in-house,” said Rehaan Ahmad, a co-founder of the Silicon Valley start-up alphaXiv, a company that tracks the latest A.I. research.

Before using those outputs, a Chinese company must first build a system that is already pretty powerful in its own right. Through distillation, it can then hone this model in important ways, but this requires additional work, money, computing power and talent.

Can Anthropic and OpenAI stop companies from distilling?

They can curb the practice — and they have tried.

If Anthropic or OpenAI suspect that someone is distilling their systems, they may shut down the offender’s account. But others will most likely pop up. And if Anthropic and OpenAI shut down too many accounts, they may end up barring legitimate users. “

It is basically impossible to stop distillation,” Lino Le Van, another alphaXiv researcher, said.

Translation

中國真的在竊取美國公司的人工智能技術嗎?

習近平主席訪問華盛頓會見特朗普總統之際,預計兩國領導人將討論有關中國暗中抄襲美國人工智能技術的指控,這一過程被稱為「蒸餾技術」。

近兩年來,Anthropic、OpenAI 和 Google 等公司一直指責中國公司不公平地抄襲他們的人工智能技術。本週,中國領導人習近平訪問華盛頓會見特朗普總統,這些指控已成為雙方會談的焦點。

鑑於人們對人工智能潛在危險的擔憂日益加劇,預計兩國領導人將討論這項技術的快速發展,許多美國官員、科技公司高層和一般民眾都呼籲放緩人工智能的發展速度。美國公司例如Anthropic 和 OpenAI 在人工智能競賽中處於領先地位,但 Z.ai 和 Moonshot 這樣的中國初創公司緊隨其後。

目前最具爭議的問題之一是矽谷一再聲稱中國公司正在利用一種名為「蒸餾」的技術來抄襲美國領先技術。

許多專家認為 Anthropic 和 OpenAI 等公司的說法言過其實。美國 Baseten 模型訓練公司的主管Charles O’Neill 說道: 「認為中國技術的全部功能都源自 Anthropic 模型的說法並不像人們所說的那樣屬實」,該公司銷售中國技術的使用權。

以下介紹這項如今已演變為地緣政治議題的小眾技術的流程。

蒸餾究竟是什麼?

蒸餾技術最初於2010年代初期開發,最初是人工智能研究人員建構更有效率技術的一種方式。利用這種技術,他們可以從現有系統中收集數據,並利用這些數據建立一個能夠在成本更低的硬體上運行的新系統。

曾參與開發這項技術的前 Google 研究員 Geoffrey Hinton 最近告訴 The Times 稱:「你可以把一個模型想像成老師,另一個模型想像成學生」。

企業仍然這樣使用這種蒸餾技術。但最近,這項技術也成為企業從他人所建構的技術中收集資料的一種方式。

基本上,外來公司如果想要蒸餾爭對手的系統,就會在競爭對手的服務平台上建立多個帳戶,例如 Anthropic 的 Claude 或 OpenAI 的 GPT。然後,這家想去做蒸餾的公司會將這些帳戶系統分配特定的任務,例如電腦程式設計或數學運算。

之後,蒸餾者就可以利用這些成熟的人工智能系統所產生的資訊來訓練自己的新系統。

蒸餾是違法嗎?

一些法學學者認為,這種蒸餾可能違反《保護商業機密法》,該法允許公司就商業機密被竊取提起訴訟。但美國法院尚未就此問題作出裁決。

版權法未必適用,因為這種蒸餾旨在複製系統的整體行為,而非逐字逐句地複製其文本。

Anthropic 和 OpenAI 的批評者指出,這兩家公司都有這類似的行為。

Anthropic 面臨多起訴訟,指控這家位於舊金山的初創公司非法使用受版權保護的網路資料來訓練其係統。去年,作為一項具有里程碑意義的法律和解協議的一部分,該公司同意向一群作家和出版商支付 15 億美元,一名法官裁定該公司非法下載並儲存了數百萬本受版權保護的書籍。和解協議是美國版權案件史上金額最高的賠償。

OpenAI 和其他公司也面臨類似的訴訟,包括《紐約時報》在2023年提起的訴訟。該訴訟稱,OpenAI及其合作夥伴 Microsoft 利用《紐約時報》發佈的數百萬篇文章來訓練聊天機器人,這些機器人如今與該新聞媒體爭奪資訊來源的地位。 OpenAI 和 Microsoft 否認了這些指控。

中國公司真的在抄襲美國系統嗎?

中國公司很可能對美國專有系統進行了蒸餾。但他們具體做了什麼尚不完全清楚。

去年,中國新創公司 DeepSeek 發佈了一套強大且高效的系統,震驚了美國研究人員和投資者。之後,OpenAI指責DeepSeek蒸餾了其技術。

在過去九個月裡,隨著中國系統的不斷改進,Anthropic 也提出了類似的指控。今年6月,Anthropic 致函參議員 Tim Scott 和 Elizabeth Warren,指責中國科技巨頭阿里巴巴蒸餾了其技術。

中國公司尚未公開回應這些指控。

僅僅蒸餾就能解釋中國模型為何保持競爭力嗎?

不能。當公司蒸餾專有系統時,他們只能存取其中的一部分:系統產生的文字、字元和其他資料。他們無法像使用開源技術那樣存取系統的底層電腦程式碼。

中國公司去利用類似 Anthropic 公司的 Claude Fable 這樣的專有系統的輸出是可以加速自身技術的開發。但這些輸出結果只是問題的一部分。

追蹤最新的人工智能研究的矽谷初創公司 alphaXiv 的聯合創始人 Rehaan Ahmad 表示:「在內部蒸餾模型時你無法獲得所需的所有信息」。

在使用這些輸出之前,中國公司必須先建立一個本身已經相當強大的系統。透過蒸餾,便能以重要的方式磨練這個模型,但這需要額外的工作、資金、計算能力和人才。

Anthropic 和 OpenAI 能否阻止公司進行蒸餾?

他們可以遏制這種行為 - 而且他們也嘗試過。

如果 Anthropic 或 OpenAI 懷疑有人在蒸餾他們的系統,他們可能會關閉違規者的帳戶。但很可能還會出現其他帳號。如果 Anthropic 和 OpenAI 關閉過多的帳戶,最終可能會誤封合法用戶。

另一位 alphaXiv 研究員 Lino Le Van 說:「基本上不可能停止蒸餾」。

            So, first developed in the early 2010s, distillation was initially a way for A.I. researchers to build more efficient technologies. Basically, an outside company looking to distill a competitor’s system will set up many accounts on the competitor’s service. The distilling company will then assign those systems to particular tasks. The reality is that if Anthropic or OpenAI suspect that someone is distilling their systems, they may shut down the offender’s account. But others will most likely pop up. Apparently, it is basically impossible to stop distillation.

沒有留言:

張貼留言