Recently The New York Times reported the following:
The Best Genetic Risk Tools Don’t Work Equally for
Everyone (2/2)
Genetic prediction models are poised to revolutionize
medicine. But because they have been trained on European DNA, they threaten to
widen health care disparities.
The NYT - By Emily Baumgaertner Nunn - Emily Baumgaertner
Nunn is a national health reporter for The Times, focusing on public health
issues that primarily affect vulnerable communities.
July 20, 2026, 5:01 a.m. ET
(continue)
Major demographic events in history, like mass migrations,
explain why genetic findings from people in Africa or of African descent are
vital to study. Communities there have accumulated great genetic diversity over
hundreds of thousands of years. By contrast, present-day Europeans descend from
a small group of Africans who expanded out of Africa about 50,000 years ago. As
a result, they have much less diversity.
“If you’re going to pick a population to study for the sake of everybody’s good, the Europeans are the worst you could choose,” said Jay Kaufman, an epidemiologist at McGill University who has written about race and genomics.
The most obvious solution is to recruit more people of color, something the All of Us program says it has prioritized. The group works with local partner organizations like churches and community health centers to host discussion sessions and enroll rural communities through a mobile clinic. It also returns individualized health findings and genetic risks to the participants to establish a sense of reciprocity.
Still, recruitment can be an upstream endeavor, particularly in the United States, which has a history of mistreatment of people of color, from the Tuskegee Syphilis Study to discriminatory sickle-cell screening to the story of Henrietta Lacks, a Black patient whose cancer cells were used for research without her knowledge or consent.
“Why would you contribute if you’ve been wronged in that way in the past?” Dr. Martin said. “There is some earned mistrust that needs to be addressed — to have that level of trust to be willing to say, ‘Take my data, monitor it for decades, do whatever you want with it.’”
In the meantime, researchers designing polygenic risk scores have found technical ways to diversify data. Modeling tools can now pull weighted information from a variety of sources simultaneously, rather than drawing only from the largest European cohorts and then forcing their algorithms onto other groups. Mount Sinai Health System in New York and U.C.L.A. in Los Angeles, for example, have both established more diverse biobanks than the UK Biobank, with each of them containing tens of thousands of genomes.
Scientists also draw from rapidly growing biobanks in China, Japan, South Korea and Taiwan, which bolster the accuracy of polygenic risk scores for East Asian groups. Localized initiatives are underway in Peru, Mexico, Qatar and elsewhere. And a continentwide program called H3Africa is collecting genetic data in African populations and supporting efforts by scientists there to study how genes and environments cause diseases.
Academic networks have also tapped sources beyond biobanks, integrating data from large, disease-specific cohort projects, including the Atherosclerosis Risk in Communities Study, or ARIC, which intentionally enrolled thousands of Black participants from North Carolina and Mississippi.
For breast cancer, a project called Confluence aggregates data from over 300 individual studies across 62 countries, providing information from 400,000 breast cancer cases and more than 1.5 million controls. The influx of data from minority groups is being used to sharpen breast cancer polygenic risk scores for everyone.
Private genomics companies and start-ups looking to develop and sell polygenic risk scores do not typically have access to those vast networks. Until recently, they relied almost exclusively on data from the UK Biobank. But the All of Us program recently expanded its policies so that companies can use its more diverse data.
Will all of these efforts, taken together, be enough? It depends on whom you ask. Some experts argue that even an ethnically representative data pool is still imperfect, since variations in other factors — socioeconomic status, health access, age or even sex — can still cause a score’s accuracy to decay. Others believe that the statistical issue was overblown in the first place.
Dr. Kenny, who has been helping to test existing polygenic risk scores for 11 common conditions in racially and ethnically diverse patients, argues that the bias is too complex to resolve quickly.
“I don’t think these things are perfectly portable yet — and maybe never will be perfectly portable,” she said. “But there are things that are starting to narrow that gap.”
Translation
最佳基因風險評估工具並非對每個人都同樣有效(2/2)
基因預測模型有望徹底改變醫學。但由於這些模型是基於歐洲人的DNA訓練的,它們有可能加劇醫療保健方面的差異
(繼續)
歷史上重大的人口發展里程,例如大規模遷徙,解釋了為什麼研究非洲人,或非洲裔人群的基因發現至關重要。這些人群在數十萬年的時間裡累積了豐富的基因多樣性。相較之下,現代歐洲人的祖先是大約5萬年前從非洲遷徙出來的一小群非洲人。因此,他們的基因多樣性要低得多。
曾撰寫過關於種族和基因組學相關文章的麥吉爾大學流行病學家Jay Kaufman說道:「如果你要選擇一個群體進行研究,以造福所有人,那麼歐洲人絕對是最糟糕的選擇」。
最顯而易見的解決方案是招募更多有色人種,這也是「我們所有人」(All of Us)計劃表示優先考慮的事項。該計劃與當地夥伴機構如教會和社區健康中心合作,舉辦討論會,並透過流動診所招募農村社區居民參與。他們也會向參與者提供個人化的健康檢查結果和遺傳風險評估,以建立一種互惠互利的意識。
然而,招募工作可能是一項帶挑戰和困難工作,尤其是在美國。美國歷史上曾多次虐待有色人種,從Tuskegee梅毒實驗到歧視性的鐮狀細胞貧血症篩檢,再到Henrietta Lacks的故事 - 一位黑人患者的癌細胞在未經她知情或同意的情況下被用於研究。
Martin博士問: 「如果你過去曾經遭受過這樣的不公,你又怎麼會願意做出貢獻呢?」; 「目前存在一些根深蒂固的不信任需要加以解決 - 要建立起足夠的信任,才能讓人願意地說: ‘拿取我的數據,監測它幾十年,你想怎麼用都行’」。
同時,設計多基因風險評分的研究人員已經找到了一些技術手段來使數據多樣化。建模工具現在可以同時從各種來源提取加權訊息,而不是僅僅從最大的歐洲群體中提取訊息,然後將其演算法強加給其他群體。例如,紐約的西奈山醫療系統和洛杉磯的加州大學洛杉磯分校都建立了比英國生物銀行更多樣化的生物樣本庫,每個樣本庫都包含數以萬計的基因組。
科學家也利用中國、日本、韓國和台灣地區快速成長的生物樣本庫,這些樣本庫提高了東亞人口多基因風險評分的準確性。秘魯、墨西哥、卡達和其他地區也正在進行類似的本地化計劃。此外,一項名為H3Africa的全非洲大陸計劃正在收集非洲人群的遺傳數據,並支持當地科學家研究基因和環境如何導致疾病。
學術網絡也開始利用生物樣本庫以外的數據來源,去整合大型特異疾病群體研究的數據,例如動脈粥狀硬化風險社群研究(ARIC)。 ARIC 特意招募了數千名來自北卡羅來納州和密西西比州的黑人參與。
在乳癌領域,一個名為 Confluence 的計劃匯聚了來自 62 個國家/地區 300 多項獨立研究的數據,提供了 40 萬個乳癌病例和超過 150 萬個對照的資訊。來自少數族裔群體的大量數據正被用於改善適用於所有人的乳癌多基因風險評分。
希望開發和銷售多基因風險評分的私人基因組學公司和初創公司通常無法存取這些龐大的數據網路。直到最近,他們幾乎完全依賴英國生物樣本庫的數據。但「我們所有人」(All of Us)計劃最近擴大了其政策範圍,允許其他公司使用其擁有的更多樣化的數據。
所有這些努力加起來是否足夠?這取決於你問誰。一些專家認為,即使是具有種族代表性的資料池也並非完美無缺,因為其他因,素例如社會經濟地位、獲得醫療服務能力、年齡甚至性別的差異仍然會導致評分準確性下降。另一些專家則認為,統計學上的問題一開始就被誇大了。
Kenny博士一直在幫助測試針對11種常見疾病的現有多基因風險評分,研究對象涵蓋不同種族和民族的患者。她認為,偏差過於複雜,難以迅速解決。
她說:“ 目前我不認為這些方法能完全通用 - 或許永遠也無法完全通用”; “但有些方法正在縮小那差距。”