Google's agentic code-review tool has completed its first full year inside the Linux kernel workflow. One headline number needs careful reading: the roughly 500 CVEs count Sashiko being cited on — not vulnerabilities it discovered.
Google's Sashiko has now logged a first full year of results inside the Linux kernel's review workflow, and the figures released so far amount to the most concrete production-scale evidence yet that agentic AI can operate meaningfully in one of software's most demanding collaborative environments. The tool has participated in roughly 191,000 patch reviews and has been cited in connection with roughly 500 kernel CVEs.
Sashiko was built by Google engineers to carry out agentic review of Linux kernel changes, drawing on Google's own AI models; the figures were shared by those engineers and first reported by Phoronix on 11 May. They are self-reported and have not been independently audited. Agentic review describes tooling that can autonomously analyse a proposed patch, flag issues, and contribute commentary to the review thread. Sashiko sits downstream of a long line of automated kernel tooling; the step-change is the move from pattern-matching linters to an agent that engages with the surrounding review context. As built, the tool functions as a way to scale the first-pass scrutiny applied to the volume of patches flowing into the kernel tree, where human reviewer bandwidth has long been a known bottleneck.
Volume versus influence: reading the metrics carefully
The headline figures deserve precise reading, because the two numbers measure different things and should not be conflated.
The 191,000 figure describes review volume — how many patch reviews Sashiko has participated in. It is a measure of throughput and adoption within the kernel community's workflow.
The roughly 500 CVE figure describes something else entirely: it counts kernel CVEs on which Sashiko has been cited. In practice, this means a reference to Sashiko — a mention in a patch description, a review-thread comment, or a security advisory — appearing in connection with one of those CVEs. That is a citation, a trace of the tool's presence in the ecosystem, not a claim that Sashiko independently discovered 500 vulnerabilities. Google has not stated, and the reported figures do not establish, what proportion of those cited CVEs Sashiko originated versus merely commented on. Treating the number as a "detection count" would materially misrepresent what it shows: evidence of influence and integration, not a measured vulnerability-discovery rate.
Equally absent from the public record is any published false-positive rate — how often Sashiko's flags turn out to be spurious — or a baseline comparison against human-only review outcomes at similar scale. Those gaps matter when weighing the tool's practical value.
A scale multiplier, not a substitute
Our own reading casts Sashiko as a scale multiplier for triage rather than a replacement for human reviewers, with the community's evident willingness to engage as the observable basis for that reading. The distinction is not incidental. The kernel's review culture is famously adversarial and rigorous; maintainers scrutinise patches closely, and contributions earn acceptance on technical merit alone. In our reading, that a Google-built tool has been adopted at this volume in a community with a well-documented allergy to top-down impositions may be the more telling datapoint than the raw counts.
It is also worth noting who is doing the reporting. Google has a clear commercial stake in demonstrating AI-review credibility: Linux underpins Android, and by extension a significant portion of Google's mobile-device ecosystem. A healthier, better-reviewed kernel is a direct corporate asset. That does not invalidate the results, but it does mean independent verification — and metrics from non-Google-affiliated deployments — will be the more meaningful benchmark as Sashiko scales into security-critical subsystems.
What Hong Kong teams can take from this
For DevOps and development teams in Hong Kong evaluating AI-assisted code review, Sashiko's trajectory offers practical signals rather than a turnkey solution. The kernel case suggests agentic review delivers its clearest value in high-volume, well-structured review environments — repositories with strong conventions, clear style rules, and testable change units. Teams auditing their own codebases should first assess whether their internal review culture and code hygiene are strong enough for an AI reviewer to add signal rather than noise; weak review cultures may not see comparable results. Beyond Sashiko itself, teams evaluating other AI code-review offerings that integrate with Git-based workflows should demand concrete metrics — flag precision, false-positive rates, and measured impact on review latency, not just throughput counts.
Second, data governance deserves explicit attention from any team sending code to vendor-hosted AI services. Internal repositories frequently contain credentials, proprietary algorithms, customer data references, and trade secrets embedded in comments or configuration files. Before routing source code through a third-party AI review service, teams should establish clear policies on which repositories are in scope, where code is processed and retained, whether outputs are used for model training, and how access is logged; sensitive codebases may warrant self-hosted tooling instead, with vendor terms reviewed as rigorously as any data-processing agreement and sign-off involving both engineering and legal stakeholders.
Whether Sashiko's results hold as it extends into more security-critical parts of the kernel — and whether comparable metrics emerge from deployments outside Google's own engineers — will be the benchmarks worth revisiting. Until independent data appears, the honest summary is this: a notable milestone in production-scale AI review, reported by the party with the most to gain, and a data point rather than a verdict.
Google 的 agentic 代碼審查工具已在 Linux kernel 工作流程中完成首個全年運作。當中一個關鍵數字需要小心解讀:約 500 個 CVE 的計數,計算的是 Sashiko 被引用提及的次數 — 並非它發現的漏洞數量。
Google 的 Sashiko 已在 Linux kernel 的審查工作流程中完成首個全年運作,迄今公布的數據,構成了迄今為止最具體的生產規模證據,證明 agentic AI 可以在軟件界要求最嚴苛的協作環境之一中有實質作為。該工具參與了約 191,000 個 patch 審查,並在約 500 個 kernel CVE 相關紀錄中被引用提及。
Sashiko 由 Google 工程師開發,運用 Google 自身的 AI 模型,對 Linux kernel 變更進行 agentic 審查;上述數據由這些工程師提供,並由 Phoronix 於 5 月 11 日首次報道。數據屬自行申報,未經獨立審計。所謂 agentic 審查,是指工具能夠自主分析建議中的 patch、標記問題,並在審查討論串中提交意見。Sashiko 位於一連串自動化 kernel 工具的下游;其突破性在於從模式比對式的 linting 工具,躍升為能夠理解並參與周遭審查脈絡的 agent。就目前設計而言,該工具的作用是擴大對流入 kernel tree 的大量 patch 所作初步審查的規模 — 人類審查者的處理能力一直是公認的瓶頸。
數量與影響力:正確解讀數據
上述關鍵數字需要精確解讀,因為兩個數字衡量的是不同事物,不應混為一談。
191,000 這個數字描述的是審查數量 — Sashiko 參與了多少次 patch 審查,反映其在 kernel 社羣工作流程中的吞吐量與採用程度。
約 500 個 CVE 的數字描述的則完全是另一回事:它計算的是 Sashiko 被引用提及的 kernel CVE 數量。實際上,這意味著在其中某個 CVE 相關脈絡中,出現了對 Sashiko 的引用 — 可能是 patch 說明中的提及、審查討論串中的留言,或安全公告中的引用。這是引用,是該工具在生態系統中曾經存在的痕跡,而非聲稱 Sashiko 獨立發現了 500 個漏洞。Google 未有說明,現有報道數據亦未能確立,這些被引用的 CVE 中有多少比例由 Sashiko 原創提出、又有多少僅是評論性質。將此數字視為「檢出數量」,會嚴重誤導其所反映的內容:這是有關影響力與整合程度的證據,而非經量化的漏洞發現率。
同樣未見於公開紀錄的,還有任何已發表的誤報率 — Sashiko 標記的問題中有多少最終被證實屬虛假 — 以及在同等規模下與純人工審查結果的基準比較。在評估該工具的實際價值時,這些缺口不容忽視。
是規模倍增器,而非替代品
我們的解讀認為,Sashiko 是分流審查的規模倍增器,而非人類審查者的替代品,而社羣明顯願意與之互動,正是此一解讀的可觀察依據。這項區別並非無關宏旨。kernel 的審查文化向來以對抗性強、一絲不苟著稱;maintainer 對 patch 仔細審視,貢獻能否被接納純粹取決於技術實力。在我們看來,一個由 Google 開發的工具,在一個對自上而下強加措施出了名抗拒、且有充分紀錄的社羣中,能以如此規模被採用,這一點可能比原始計數更能說明問題。
另一點值得注意的是由誰負責發布數據。Google 在證明 AI 審查可信度方面有明確的商業利益:Linux 是 Android 的根基,繼而支撐 Google 流動生態系統的相當大部分。一個更健康、審查更完善的 kernel,對企業而言是直接資產。這無損結果的有效性,但確實意味著 — 隨著 Sashiko 擴展至安全關鍵的 subsystem — 獨立驗證以及來自非 Google 關聯部署的數據,才會是更有意義的基準。
香港團隊可從中借鏡
對於香港正在評估 AI 輔助代碼審查的 DevOps 及開發團隊而言,Sashiko 的發展軌跡提供的是實際參考訊號,而非現成方案。Kernel 案例顯示,agentic 審查在數量龐大、結構良好的審查環境中能發揮最清晰的價值 — 即具備嚴格規範、明確風格規則、以及可測試變更單元的 repository。審核自身代碼庫的團隊,應先評估其內部審查文化與代碼衛生是否足夠強健,令 AI 審查者能提供有用訊息而非雜訊;審查文化薄弱的團隊,未必能取得可比的成效。除 Sashiko 本身外,評估其他整合 Git 工作流程的 AI 代碼審查方案時,團隊應要求具體數據指標 — 例如標記精準度、誤報率,以及對審查延遲的實測影響,而不只是吞吐量數字。
其次,任何將代碼傳送至供應商託管 AI 服務的團隊,都應明確重視數據治理。內部 repository 常常包含嵌於註解或配置文件中的憑證、專有演算法、客戶數據參考及商業機密。在將源代碼送往第三方 AI 審查服務之前,團隊應制定清晰政策,界定哪些 repository 在處理範圍之內、代碼在何處處理及留存、輸出是否用於模型訓練,以及存取紀錄如何保存;涉及敏感代碼庫者,或應考慮採用自託管工具,供應商條款須像審視任何數據處理協議一樣嚴謹,並須工程與法律相關持份者共同簽署核准。
Sashiko 的成果能否在其延伸至 kernel 更多安全關鍵部分後依然成立 — 以及 Google 自身工程師以外的部署會否產生可比的數據 — 將是值得持續關注的基準。在獨立數據出現之前,最誠實的總結是:這是生產規模 AI 審查的一個重要里程碑,由獲益最大的一方自行發布,是一個數據點,而非最終定論。
