A recent disclosure from U.S. intelligence and cybersecurity agencies alleges that six Chinese artificial intelligence companies executed a systematic, large-scale campaign to extract core knowledge from leading American AI models. The operation, which reportedly extracted billions of tokens of training data since late 2024, has been framed not merely as corporate espionage, but as a stark revelation of the AI industry's foundational security crisis.
According to a report by BleepingComputer detailing the U.S. government assessment, the companies utilized techniques known as distillation attacks. This method exploits a core function of large language models—their ability to generate text. By sending vast volumes of automated queries to public model interfaces, adversaries can methodically reconstruct the underlying patterns, reasoning capabilities, and encoded knowledge that define a frontier system's value. The harvested outputs are then used to train competing models at a fraction of the original cost and effort.
The implications extend far beyond the immediate geopolitical friction. The incident underscores a critical vulnerability in the very open-access model that propelled the AI revolution. While terms of service and export controls exist, they have proven inadequate against industrial-scale extraction, revealing a profound governance vacuum. The attack surface is inherent to the business model: making powerful AI accessible via APIs also makes it exploitable.
Existing legal and technical safeguards are being exposed as obsolete. Rate limiting and automated filtering struggle to distinguish between legitimate high-volume research and methodical theft, with false positives posing their own service risks. The scale of this alleged operation—coordinated and state-backed—suggests that the competitive landscape itself may be unsustainable if model capabilities can be cloned through API calls. This shifts the core economic concern from intellectual property theft to the viability of investing billions in developing frontier models.
The event marks a significant expansion of the U.S.-China technology competition, moving from hardware restrictions on chips and manufacturing tools to the data layer itself. Traditional export regimes are ill-suited to govern knowledge encoded in neural networks, forcing policymakers to confront the need for entirely new international norms and technical standards for model provenance and output certification.
In response, the AI sector is expected to accelerate investment in protection technologies, including cryptographic watermarking of outputs and advanced anomaly detection. However, the efficacy of such measures against determined, well-resourced actors remains uncertain. For organizations deploying AI, this incident is a clear signal to adopt a security-first posture, treating model supply chain integrity with the rigor traditionally applied to software security.
The allegations, as presented in the U.S. report, point to a systemic inflection point. As AI models become foundational infrastructure, the era of treating their access and replication as an open, low-friction endeavor appears to be ending. The focus now urgently shifts to defining the security and governance frameworks necessary to protect the intellectual property and strategic advantages embedded within them.
美國情報與網絡安全機構近日披露,六家中國人工智能公司執行了一項系統性的大規模行動,從領先的美國AI模型中提取核心知識。據悉,該行動自2024年底起已提取了數十億個標記(tokens)的訓練數據,不僅被定性為企業間諜行為,更揭示了AI產業面臨的基礎安全危機。
根據BleepingComputer詳述美國政府評估的報告,相關公司採用了名為「蒸餾攻擊」的技術。此方法利用大型語言模型的核心功能——文本生成能力。對手透過向公開模型界面發送大量自動化查詢,可系統性地重建定義前沿系統價值的底層模式、推理能力與編碼知識。這些獲取的輸出隨後被用於以遠低於原始成本和精力的訓練競爭模型。
此事件的影響遠超即時的地緣政治摩擦。它凸顯了推動AI革命的開放獲取模式存在關鍵漏洞。儘管存在服務條款和出口管制,但這些措施已被證明無法應對工業規模的提取行動,暴露出嚴重的治理真空。攻擊面本身源自商業模式:透過API提供強大的AI能力,同時也使其可被利用。
現行的法律和技術保障措施正被揭露為過時。速率限制和自動化過濾機制難以區分合法的高流量研究與系統性竊取,而誤報本身也構成服務風險。此次被指控行動的規模——有組織且受國家支持——表明,如果模型能力可以通過API調用被複製,競爭格局本身可能難以為繼。這將核心經濟關注點從知識產權被竊,轉向投入數十億美元開發前沿模型的可行性。
此事件標誌著中美技術競爭的重大擴展,從硬件層面的芯片和製造工具限制,延伸至數據層面本身。傳統出口管制體制並不適合管理編碼在神經網絡中的知識,迫使政策制定者面對建立全新國際規範和技術標準的需求,以保障模型來源和輸出認證。
作為回應,預計AI行業將加速投資保護技術,包括輸出內容的加密數字水印和先進異常檢測。然而,這些措施對抗有決心且資源充足的行為者的有效性仍不確定。對於部署AI的組織而言,此事件明確發出信號,應採取安全優先的姿態,以傳統應用於軟件安全的嚴謹態度對待模型供應鏈完整性。
美國報告中提出的指控,指向一個系統性的轉折點。隨著AI模型成為基礎設施,將其訪問和複製視為開放、低摩擦事業的時代似乎正在終結。現今的緊迫焦點轉向定義必要的安全與治理框架,以保護嵌入其中的知識產權和戰略優勢。
