If you've ever monitored a Linux virtual machine and noticed the st% column in top spiking into double digits, you know what steal time feels like: your vCPU is ready to run, but the hypervisor has yanked it away to serve another tenant or handle housekeeping. Until now, the kernel's CPU frequency governors—like powersave, performance, and schedutil—have been largely blind to that signal. A new "steal governor" queued for the Linux 7.4 merge window, expected later in October, aims to change that by making steal time a first-class input into frequency scaling decisions.
As highlighted by Phoronix, the steal governor is purpose-built for virtualized environments rather than adapted from bare-metal assumptions. Its core idea is straightforward: when the hypervisor reports frequent vCPU preemption—such as in a "noisy neighbor" scenario—there's little point in ramping up clock speed, since the core won't get to finish its work anyway. Conversely, when a vCPU enjoys long uninterrupted runs, the governor can confidently request higher frequencies to maximize performance during guaranteed execution windows.
The implementation rests on two integrated techniques. First, steal-driven vCPU backoff lets the governor detect periods of heavy host preemption and dial back frequency requests, avoiding wasted energy and thermal pressure on cores that the VM can barely use. Second, preferred CPU awareness allows the scheduler to favor cores that historically deliver more stable, less contested run time—essentially steering workloads toward quieter corners of the host when possible. Together, these mechanisms address the persistent performance variability that plagues cloud instances, promising more consistent latency and throughput for workloads that depend on predictable timing.
For cloud and DevOps practitioners, the practical implications are meaningful. Containerized microservices running on shared clusters often suffer from tail-latency spikes caused by unpredictable vCPU scheduling. Database workloads inside VMs can see query throughput dip when a neighboring tenant bursts. The steal governor offers the kernel a data-driven way to adapt its frequency policy to the realities of multi-tenant infrastructure, potentially smoothing out jitter without requiring changes to application code or hypervisor configuration.
The feature also marks a broader strategic maturation for Linux. The kernel now runs the vast majority of its global compute footprint on top of hypervisors—public cloud instances, private VMware or KVM clusters, and managed Kubernetes platforms alike. A CPU governor designed around virtualization signals that the kernel community is treating that deployment model as the default rather than an edge case. It acknowledges that bare-metal optimizations, while still valuable, no longer cover the dominant use case.
Open questions remain around real-world validation. Specific benchmark data comparing the steal governor against previous defaults under varied load profiles will be needed before cloud providers consider adopting it as a new baseline image setting. Enterprise distributions—the RHELs, Ubuntu LTS releases, and SUSE versions that anchor production environments—will also need to decide how quickly to carry the feature back from upstream. Given that the code is entering the 7.4 merge window now, a realistic timeline for availability in production-grade images likely stretches into 2027.
For teams running workloads in Hong Kong's rapidly growing cloud market—whether on hyperscale providers or local infrastructure-as-a-service platforms—the steal governor represents a welcome shift: the kernel finally learning to schedule smarter in the environment where most of its instances actually live.
如果你曾經監控過一個 Linux 虛擬機,並注意到 top 指令中的 st% 欄位飆升至兩位數,你就會明白「被偷取時間」的感覺:你的 vCPU 已準備就緒,但管理程式(hypervisor)已將其強行調走,以服務或另一個租戶或處理系統管理任務。到目前為止,Linux 核心的 CPU 頻率調節器——如 powersave、performance 和 schedutil——基本上對此信號是視而不見的。一個新的「steal governor」已排入 Linux 7.4 的合併窗口,預計將於十月稍後推出,旨在改變這種情況,讓「被偷取時間」成為頻率調整決策的一個首要輸入參數。
正如 Phoronix 所強調,此 steal governor 是專為虛擬化環境而設計,並非從實體機(bare-metal)的假設改編而來。其核心理念很直接:當管理程式報告頻繁的 vCPU 搶佔——例如在「吵鬧的鄰居」場景中——提高時脈頻率幾乎沒有意義,因為該核心反正也無法完成其工作。相反,當 vCPU 享有長時間的連續運行時,調節器便能自信地要求更高的頻率,以在有保證的執行窗口期內最大化效能。
其實作仰賴兩項整合的技術。第一,由 steal 驅動的 vCPU 回退(backoff),讓調節器能偵測到主機搶佔嚴重的時段,並降低頻率請求,從而避免虛擬機幾乎無法使用的核心產生能源浪費和熱量壓力。第二,偏好的 CPU 認知,使排程器能傾向選擇那些在歷史上提供更穩定、較少競爭運行時間的核心——本質上是在可能的情況下,將工作負載引導至主機上較為安靜的角落。這些機制共同解決了困擾雲端實例的持續性效能變異問題,為依賴可預測時序的工作負載帶來更一致的延遲與吞吐量。
對於雲端與 DevOps 從業人員而言,其實際影響顯著。在共享叢集上運行的容器化微服務,經常因不可預測的 vCPU 排程而飽受尾部延遲飆升之苦。虛擬機內的資料庫工作負載,在鄰近租戶出現流量高峰時,其查詢吞吐量也可能下降。steal governor 為核心提供了一種數據驅動的方式,使其頻率策略能適應多租戶基礎設施的實際狀況,有望平滑抖動,而無需更改應用程式代碼或管理程式設定。
此特性也標誌著 Linux 更廣泛的戰略成熟。目前,其全球絕大多數的運算負載都運行於管理程式之上——無論是公有雲實例、私有 VMware 或 KVM 叢集,還是託管式 Kubernetes 平台。一個圍繞虛擬化設計的 CPU 調節器,表明核心社群正將此部署模式視為預設選項,而非邊緣案例。它承認了實體機優化雖然仍有價值,但已不足以涵蓋主流使用場景。
關於實際驗證仍有待解決的問題。在雲端供應商考慮將其作為新的基礎映像設定之前,需要在不同負載配置下,比較 steal governor 與先前預設值的具體基準測試數據。企業發行版——如 RHEL、Ubuntu LTS 版本及 SUSE 版本,它們是生產環境的支柱——也將需要決定多快將此特性從上游回移。鑒於該代碼目前正進入 7.4 合併窗口,其在生產級映像中的可用時間表,很可能要延伸至 2027 年。
對於在香港蓬勃發展的雲端市場中運行工作負載的團隊——無論是使用超大規模供應商,還是本地基礎設施即服務(IaaS)平台——steal governor 代表了一個受歡迎的轉變:核心終於學會在其大多數實例實際所處的環境中,更智能地進行排程。
