The Codeberg forge has enacted a pair of new policies targeting large language models, drawing what may be the sharpest institutional line yet between open-source communities and the AI industry. The community-hosted platform has committed to two measures: promising not to use any hosted projects to train LLMs, and banning the hosting of LLM-generated software.
The announcement, covered by LWN.net, shifts the enforcement mechanism from individual project maintainers—who have historically set their own rules about AI contributions—up to the infrastructure layer. This represents a notable precedent: a major hosting platform explicitly declaring that its repositories are not available as a data source for commercial AI development.
Protecting a Commons
At the heart of the new policies is a philosophical argument about what constitutes legitimate software libre. As Codeberg explained in its accompanying blog post, "sharing the result of a prompt and calling it 'libre software' does not make the world a better place." The organization frames the move as defense of the FLOSS commons—a shared resource built on human collaboration, mutual obligation, and the ethical commitments embedded in free and open-source licensing.
The training data prohibition confronts the prevailing industry practice of scraping public code repositories to build commercial AI systems. Developers who contribute to FLOSS projects typically do so under licenses designed to govern human-to-human collaboration. Codeberg's policy reframes the systematic harvesting of that code for model training as a matter of data stewardship and community consent, not merely a licensing question.
Analysis: The Enforcement Challenge
The more contentious measure is the outright ban on hosting LLM-generated code. While the intent is clear—preserving the integrity of human-authored software—the practical implications are formidable. As AI coding assistants become standard tools in development workflows, distinguishing between code written with AI assistance and code produced primarily by an LLM is a fundamentally different problem than detecting plagiarism or license violations.
How will Codeberg assess whether a project crosses this threshold? What counts as "primarily" LLM-generated when developers routinely use AI tools for debugging, refactoring, and boilerplate generation? These questions have no clean answers, and the difficulty of fair, consistent enforcement at scale is arguably the single greatest challenge to the policy's viability.
Analysis: A Community Divided
The controversy crystallizes a deeper ideological rift within the free software community. On one side are developers who view AI tools as natural extensions of the existing toolchain—compilers, linters, and IDEs have always automated aspects of software creation. On the other are contributors who see uncritical adoption of LLM-generated code as incompatible with the collaborative, human-centered ethics that have sustained the FLOSS ecosystem for decades.
Codeberg's action positions the platform squarely in the latter camp, potentially attracting developers seeking an AI-free environment while simultaneously risking alienation of those who see such restrictions as reactionary. The broader ecosystem's response will be telling: if other major hosting foundations adopt similar policies, a patchwork of platform-specific AI governance may emerge across the open-source landscape.
For technology organizations that rely on open-source components, the developments at Codeberg underscore an evolving reality. Software provenance—the chain of custody and rights associated with code—is no longer a peripheral concern. It is becoming a material factor in supply chain risk, licensing compliance, and the broader governance of AI adoption within development teams. The question is no longer whether platforms will set policies around AI and code, but what shape those policies will take and how organizations will need to adapt.
Codeberg 代碼倉庫平台已制定一項針對大型語言模型的新政策,劃下了開源社群與 AI 產業之間迄今最明確的制度界線。該社群託管平台承諾採取兩項措施:不使用任何託管的項目來訓練 LLM,並禁止託管 LLM 生成的軟件。
此項由 LWN.net 報導的公告,將執行機制以往由個別項目維護者自行設定 AI 貢獻規則,提升至基礎設施層級。這代表一項顯著的先例:一個主要託管平台明確聲明其代碼倉庫不作為商業 AI 開發的數據來源。
捍衛公共領域
新政策的核心,是關於何謂正當自由軟件(libre)的哲學論辯。正如 Codeberg 在其附隨的博客文章中解釋:「分享一個提示詞的結果並將其稱為『自由軟件』,並不會讓世界變得更美好。」該組織將此舉定位為對 FLOSS 公共領域的捍衛——這是一個建立在人類協作、相互義務以及嵌入於自由及開源軟件許可證中倫理承諾之上的共享資源。
訓練數據禁令直面業界普遍採用的抓取公共代碼倉庫以建構商業 AI 系統的做法。貢獻於 FLOSS 項目的開發者,通常是在旨在規範人與人之間協作的許可證下進行貢獻。Codeberg 的政策將系統性收割這些代碼用於模型訓練,重新詮釋為數據管理與社群同意的問題,而非僅僅是許可證問題。
分析:執法挑戰
更具爭議的措施是直接禁止託管 LLM 生成的代碼。雖然意圖明確——維護人類編寫軟件的完整性——但其實際影響相當嚴峻。隨著 AI 編碼助手成為開發工作流程中的標準工具,區分是借助 AI 編寫的代碼,還是由 LLM 主要生成的代碼,本質上不同於偵測剽竊或許可證違規。
Codeberg 將如何評估一個項目是否跨越此門檻?當開發者常規性地使用 AI 工具進行除錯、重構和樣板代碼生成時,什麼情況可被視為「主要」由 LLM 生成?這些問題沒有清晰的答案,而大規模下公平、一致的執法難度,可說是該政策可行性的最大單一挑戰。
分析:分裂的社群
這場爭議凸顯了自由軟件社群內部更深層的意識形態分歧。一方是開發者,他們視 AI 工具為現有工具鏈的自然延伸——編譯器、檢查工具和整合開發環境(IDE)一直自動化著軟件創建的某些方面。另一方則是貢獻者,他們認為不加批判地採用 LLM 生成的代碼,與維繫 FLOSS 生態系統數十年、以協作為中心且以人為本的倫理觀相悖。
Codeberg 的行動使該平台明確站在後者陣營,可能吸引尋求無 AI 環境的開發者,同時也有風險疏遠那些視此類限制為反動的人士。更廣泛的生態系統回應將具有啟發性:如果其他主要託管基金會也採取類似政策,整個開源領域可能出現一系列針對特定平台的 AI 治理拼湊方案。
對於依賴開源組件的技術組織而言,Codeberg 的發展凸顯了一個正在演變的現實。軟件來源溯源——代碼的監管鏈與相關權利——不再是一個邊緣問題。它正成為供應鏈風險、許可證合規性以及開發團隊中更廣泛的 AI 採用治理中的實質因素。問題已不再是平台是否會圍繞 AI 和代碼制定政策,而是這些政策將採取何種形式,以及組織將如何適應。
