Google engineers are pushing to extend compiler-driven performance optimizations — currently applied only to the main Linux kernel image — to loadable kernel modules, according to Phoronix. The effort also aims to clear technical limitations that today prevent the AutoFDO and Propeller optimization techniques from being applied to modules at all.
Profile-guided and feedback-directed compiler optimizations have, in recent years, become an established lever inside large software projects: the compiler is fed data about how code actually behaves at runtime, then recompiles the same source with better-informed decisions about inlining, layout, and branch prediction. Inside the Linux world, that lever has been pulled almost exclusively on vmlinux — the monolithic kernel image. Everything that ships as a loadable .ko module has been left out.
Google now wants to change that, and to do so on a footing that stands a better chance of surviving upstream review than earlier attempts. Phoronix notes that Linus Torvalds has previously pushed back on bringing PGO into the mainline kernel. Rather than relitigating that debate head-on, the current push is reportedly scoped as out-of-tree functionality first, with module-level support as an explicit goal — a narrower, more incremental approach than a direct mainline-merge proposal.
Three techniques, briefly. Profile-guided optimization (PGO) compiles the code once, runs it under real workloads to record which paths are hot and which branches are predictable, then recompiles with that knowledge. AutoFDO, closely associated with Google's kernel toolchain work, achieves a similar effect without instrumented builds, instead harvesting profile data from production samples — a practical advantage for codebases as large and as varied as the kernel's. Propeller, also closely associated with Google's compiler-optimization efforts, takes a link-time, layout-reordering approach built on profile feedback to reduce instruction-cache pressure. All three are, in essence, ways of making the compiler's guesses match reality.
Why modules matter. The kernel image is only part of the story. Networking stacks, storage drivers, filesystems, container runtimes, and security modules frequently ship as loadable modules — and modules are exactly where hyperscalers and mobile-vendor kernels concentrate much of their differentiated code. If profile-guided techniques can only touch vmlinux, a large fraction of the code actually executing on production hosts never benefits. Extending the pipeline to modules closes that gap, and — crucially for anyone running a stock distribution kernel rather than a vendor-built one — it is the prerequisite for these gains ever showing up outside the small set of organizations that already build their own kernels with the toolchains and profiling infrastructure to support PGO today.
The practical payoff is significant. The performance benefits of profile-guided kernel builds have so far been concentrated in the hands of operators who can afford the build complexity. A module-capable, out-of-tree-first workflow is a plausible on-ramp toward making those gains more broadly attainable — including for teams that consume kernels from a distribution rather than compiling them in-house. Whether a concrete patch series exists, which engineers are behind it, whether any target kernel release or merge window is in play, and what — if any — quantified performance figures accompany the proposal are details addressed in Phoronix's full coverage.
For operators, the open question is real. Whether module-level, out-of-tree-first scoping is enough to overcome the reproducibility and profile-stability objections that have historically blocked PGO in the mainline kernel is not yet settled. For infrastructure teams, the stakes are concrete. Kubernetes fleets running dense, CPU-bound workloads, telco environments pushing NFV functions to their latency limits, and low-latency trading infrastructure where every microsecond of kernel-path jitter is measured — all of these sit in the population that benefits most from a kernel tuned by real workload profiles rather than generic heuristics. Whether that benefit arrives for stock builds, or remains confined to vendor kernels, is the thing to watch as this proposal matures.
Source: Phoronix, Google Working To Leverage More Compiler Optimizations Within The Linux Kernel.
據 Phoronix 報導,Google 工程師正推動將 compiler 驅動的效能優化擴展至可載入核心模組(loadable kernel modules),目前這些優化只應用於主 Linux 核心映像。該項工作亦旨在清除技術限制,令 AutoFDO 與 Propeller 優化技術目前根本無法應用於模組的情況得以改變。
近年來,profile-guided 與 feedback-directed 編譯器優化已成為大型軟件專案中一項成熟的手段:編譯器會接收程式碼在執行期實際行為的數據,再以更充分的資訊重新編譯同一份原始碼,就 inlining(函式內聯)、佈局(layout)及 branch prediction(分支預測)作出更好的決定。在 Linux 生態圈中,這一手段幾乎只用於 vmlinux —— 即單體核心映像(monolithic kernel image)。所有以可載入 .ko 模組形式發佈的程式碼則一直被排除在外。
Google 現在希望改變這一點,而且是以較早期嘗試更有機會通過 upstream review(上游審核)的方式進行。Phoronix 指出,Linus Torvalds 此前曾反對將 PGO 引入 mainline 核心。據報導,目前的推動並非正面重提該爭論,而是先定位為 out-of-tree(樹外)功能,並將模組層級支援列為明確目標 —— 這是比直接提案合入 mainline 更為收窄、更為漸進的路徑。
三項技術簡介。 Profile-guided optimization(PGO)先編譯一次程式碼,在真實工作負載下執行以記錄哪些路徑屬熱路徑(hot path)、哪些分支具有可預測性,然後運用這些資料重新編譯。AutoFDO與 Google 的核心工具鏈工作密切相關,它毋須在編譯時加入插樁(instrumented builds)即可達到類似效果,而是從 production samples 中擷取 profile data —— 對於核心這樣規模龐大且高度多樣的程式碼庫而言,這是一項實際優勢。Propeller同樣與 Google 的編譯器優化工作密切相關,它採用 link-time(連結期)的佈局重排(layout reordering)方式,以 profile feedback 為基礎來降低 instruction cache 的壓力。三者本質上都是令編譯器的推斷與現實吻合的方法。
為何模組如此重要。 核心映像只是故事的一部分。網絡協定堆疊(networking stacks)、儲存驅動程式、檔案系統、容器 runtime 以及安全模組,經常以可載入模組形式發佈 —— 而模組正是 hyperscalers(超大規模雲服務商)及手機廠商核心大量集中其差異化程式碼之處。若 profile-guided 技術只能觸及 vmlinux,生產主機上實際執行的相當一部分程式碼便永遠無法受惠。將 pipeline 擴展至模組,可補上這道缺口;而 —— 對任何使用發行版所提供的核心而非供應商自建核心的用戶來說尤其關鍵 —— 這是相關效益能否走出目前那批已具備支援 PGO 所需工具鏈及 profiling 基礎設施、能自行編譯核心的小圈子的先決條件。
其實際回報相當可觀。Profile-guided 核心建置的效能效益,至今仍集中在有能力負擔複雜建置流程的營運者手中。一套支援模組、以 out-of-tree 為先的工作流,是令這些效益變得更廣泛可達的可行起步 —— 包括那些從發行版獲取核心、而非自行編譯核心的團隊。至於是否已有一套具體的 patch series、背後是哪些工程師、是否有既定的核心版本或 merge window,以及提案是否附有任何量化的效能數據,均在 Phoronix 的完整報道中有進一步交代。
對營運者而言,這個疑問確實存在。 以模組層級、out-of-tree 為先的定位,是否足以克服歷來阻礙 PGO 進入 mainline 核心的可重複性(reproducibility)及 profile 穩定性(profile stability)反對意見,至今仍未有定論。對基建團隊而言,利害關係是具體的。密集型 CPU-bound 工作負載的 Kubernetes fleet、將 NFV 功能推向延遲極限的電信環境、以及每微秒核心路徑抖動(jitter)都要量度的低延遲交易基建 —— 這些都屬於最能受惠於經真實工作負載 profile 調校核心、而非依賴通用啟發式規則的群體。相關效益最終是否會惠及標準建置,抑或仍只侷限於供應商核心,正是這項提案走向成熟時需要密切關注的問題。
資料來源:Phoronix,Google Working To Leverage More Compiler Optimizations Within The Linux Kernel。
