Malware campaign targets self-hosted inference servers, converting them into both profit centers and scanning infrastructure
A cryptomining campaign built around malware called PoeLLM is compromising exposed AI services and using them for far more than coin generation, according to a report published by BleepingComputer on 7 October. Rather than mining quietly and moving on, the operators convert each infected server into an active scanning and exploitation platform — a design that lifts the threat out of opportunistic theft and into persistent network foothold territory.
How the campaign works
The campaign goes after unsecured self-hosted AI inference servers — deployments commonly running frameworks such as Ollama or vLLM — that are reachable directly from the internet, frequently with no authentication in place. Once an instance is exposed, PoeLLM is deployed to carry out two tasks in parallel.
The first is the most visible: cryptocurrency mining. Infected machines burn GPU and CPU cycles generating coins for the attackers, drawing down the power and capacity that the operator believes is serving legitimate model inference.
The second is what gives the campaign its teeth. PoeLLM-driven infrastructure also scans for other vulnerable services and launches exploitation attempts from the compromised host itself. Victims are effectively conscripted into the attack chain, with scanning traffic originating from infrastructure the organization owns and trusts. That creates real consequences for defenders: outbound scanning from a legitimate server can resemble the traffic a security appliance or load balancer routinely generates, and sustained compute load from a rogue mining process can be indistinguishable from a busy inference node.
This detection ambiguity is the practical sting of the campaign. An administrator chasing a performance complaint may see a GPU running at capacity and conclude the model is simply under load — not that a hijacked process is quietly mining and probing the wider network.
Why AI infrastructure is a target
Security practitioners have warned for months that the pace of AI adoption is running ahead of the security maturity surrounding it. The pattern in this campaign follows a familiar arc: new platforms inherit the exposure problems of the web and database eras — misconfigured interfaces, missing authentication, services bound to public interfaces — while the tooling guarding them remains less mature. In PoeLLM's case, the models themselves are not the objective; the servers hosting them are.
Indicators and what defenders should look for
The researchers behind the analysis have released indicators of compromise tied to the campaign, including the malware artifacts and infrastructure connected to the scanning and mining activity. Behavioral signs worth monitoring include unexpected GPU utilization sustained outside business hours, egress connections to unusual destinations or high-volume port scans originating from AI hosts, and new processes or containers appearing on systems that should only be running inference workloads.
Practical guidance
Organizations running self-hosted inference servers should treat them as internet-facing workloads: place them behind authentication rather than exposing raw API endpoints, segment them from the rest of the network so a compromised node cannot reach internal systems, apply egress filtering to catch scanning and mining-traffic patterns, and alert on GPU and network anomalies rather than assuming heavy compute equals legitimate demand.
The exposure class applies broadly. Any organization running self-hosted AI inference — regardless of size or sector — that binds its inference stack to a public interface without authentication is a candidate for the same treatment.
Source: BleepingComputer.
惡意軟件攻擊針對自行託管的 AI 推理伺服器,把它們同時變成牟利用途的礦機與掃描基建
BleepingComputer 在 10 月 7 日報道了一個以 PoeLLM 惡意軟件為核心的加密貨幣挖掘攻擊,該攻擊正入侵暴露於互聯網的 AI 服務,並利用它們從事遠超單純挖礦的活動。攻擊者不會低調挖掘後便抽身離去,而是把每台受感染的伺服器轉化成一個主動掃描與入侵的平台 —— 這設計把威脅由機會式偷竊,提升至在受害者網絡中長期立足的層次。
攻擊如何運作
攻擊鎖定那些自行託管、缺乏保護的 AI 推理伺服器 —— 這類部署通常運行 Ollama 或 vLLM 等框架 —— 這些伺服器直接暴露於互聯網,而且往往沒有任何身份驗證。一旦某個實例可被接觸,PoeLLM 便被部署去同時執行兩項任務。
第一項任務最為明顯:加密貨幣挖掘。受感染的機器會耗用 GPU 及 CPU 周期為攻擊者生成加密貨幣,搶佔運算電力與算力,而系統管理員往往以為這些資源正用於合法的模型推理。
第二項任務才是攻擊真正可怕的地方。PoeLLM 所驅動的基建同時會掃描其他系統上的漏洞,並從受感染的主機本身發起入侵嘗試。受害者實際上被強制加入攻擊鏈 —— 掃描流量出自該組織自有並信任的基礎設施。這對防守方帶來實質影響:由合法伺服器發出的對外掃描流量,可能與保安設備或負載平衡器例行產生的流量難以分辨;而由惡意挖礦程序造成的持續高運算負載,亦可能與一個繁忙的推理節點無異。
偵測上的模糊性正是這攻擊的實際痛點。管理員在處理系統變慢時,看到 GPU 運行在滿載狀態,很可能會判斷只是模型負載過高,而不是意識到有被劫持的進程正秘密挖礦並對更廣泛的網絡進行探測。
為何 AI 基建成為目標
保安從業員數月來一直警告,AI 採用的速度正在超越周邊保安控制的成熟度。這次攻擊的模式呼應一個熟悉的現象:新平台繼承了網絡與資料庫時代的暴露問題 —— 設定不當的介面、缺乏身份驗證、綁定至公開介面的服務 —— 而配套的保安工具則尚未成熟。就 PoeLLM 而言,模型本身不是攻擊目標;承載它們的伺服器才是。
指標與防守方應留意的跡象
負責這項分析的研究人員已公開與該攻擊相關的入侵指標(IoC),包括與掃描及挖礦活動相關的惡意軟件檔案與攻擊基礎設施。值得監控的行為跡象包括:在辦公時間以外仍然持續的 GPU 高用量、來自 AI 主機的非典型對外連線或高流量端口掃描,以及出現在只應運行推理工作負載的系統上的新進程或容器。
實務建議
自行託管推理伺服器的機構應把它們視為面向互聯網的工作負載:應把它們置於身份驗證之後,而非直接暴露 API 端點;把它們與網絡的其他部分分段,令受感染節點無法接觸內部系統;加入對外流量過濾以攔截掃描及挖礦流量的模式;並針對 GPU 與網絡異常作出預警,而非假設高運算用量等於合法需求。
這次暴露類別具廣泛影響。任何自行託管 AI 推理 —— 無論規模或行業 —— 但把推理系統綁定至公開介面而不設身份驗證的機構,都有可能成為同一套攻擊的目標。
資料來源:BleepingComputer。
