Rob Gibbon has published a walkthrough on Canonical's Ubuntu blog showing how a language model can be deployed inside a realtime transaction fraud detection loop — not by asking it to write anything, but by using it as a classifier that returns fraud probabilities directly.
The distinction is the heart of the post. Conversational AI, Gibbon argues, is well suited to drafting emails or debugging code, but less appropriate as real-time application middleware. When a financial transaction needs inspecting in the middle of a checkout loop, the requirement is a score, not prose — an essay explaining why a credit card transaction looks suspicious is exactly what the system should not produce on the critical path.
Gibbon's solution repurposes the LLM rather than displacing it. The system uses a fine-tuned Qwen3.5-4B model, referred to as system-one-qwen3.5-4b-scorer, deployed to return classification probabilities for a transaction while bypassing autoregressive text generation. The model stays in the inference path, but the output is a bounded classification result rather than open-ended generated text — the mechanism that makes an LLM viable inside a latency-sensitive scoring step.
The post also documents the deployment stack in detail. The model is served on Kubernetes using Charmed Kubeflow on MicroK8s, with KServe handling model serving and vLLM providing the inference runtime. Orchestration is handled through Juju, with Terraform covering infrastructure provisioning. On the model side, the walkthrough covers merging a LoRA adapter into the base model and applying AWQ 4-bit quantization to keep the memory footprint manageable for serving. Gibbon's post links to the associated GitHub repositories for readers who want to reproduce the setup.
For developers evaluating LLMs beyond the chatbot pattern, the post's value lies less in any single component than in the deployment pattern it makes concrete: a quantized, fine-tuned small model deployed as an on-premises inference service, queried for scores rather than essays, running on infrastructure the operator controls through standard Kubernetes tooling.
Rob Gibbon 於 Canonical 的 Ubuntu blog 上發表了一篇操作指南,展示如何將語言模型部署在即時交易詐騙偵測的流程之中——做法並非要求模型撰寫任何內容,而是將其作為分類器使用,直接回傳詐騙機率。
這正是整篇文章的核心所在。Gibbon 認為,對話式 AI 適合撰寫電郵或除錯(debugging)程式碼,但作為即時應用程式的中介層(middleware)則不太恰當。當一宗金融交易需要在結帳流程中途即時檢視時,系統需要的是一個分數,而非散文——一篇解釋某宗信用卡交易為何可疑的文章,正是系統在關鍵路徑上 不應 產出的內容。
Gibbon 的方案是重新利用(repurpose)LLM,而非將其取代。系統採用經過微調(fine-tuning)的 Qwen3.5-4B 模型,命名為 system-one-qwen3.5-4b-scorer,部署後可為交易回傳分類機率,同時略過自回歸(autoregressive)文本生成。模型仍然留在推論路徑(inference path)之內,但輸出的是一個範圍明確的分類結果,而非開放式生成文本——正是這一機制令 LLM 能夠在延遲敏感的評分步驟中實際可行。
該文亦詳細記錄了整套部署技術堆疊(stack)。模型透過 MicroK8s 上的 Charmed Kubeflow 部署於 Kubernetes,由 KServe 負責模型服務(model serving),vLLM 則提供推論運行時(inference runtime)。編排(orchestration)由 Juju 處理,基礎設施供給則由 Terraform 覆蓋。在模型方面,指南涵蓋了如何將 LoRA 適配器(adapter)合併至基礎模型,以及套用 AWQ 4-bit 量化(quantization),以便將服務時的記憶體佔用量控制在合理範圍之內。Gibbon 的文章亦附上了相關 GitHub repository 的連結,供讀者重現整套設置。
對於希望超越聊天機器人模式、探索 LLM 其他應用的開發者而言,這篇文章的價值與其說在於任何單一組件,不如說在於它所具體呈現的部署模式:一個經過量化及微調的小型模型,以本地部署(on-premises)的推論服務形式運行,查詢時取得的是分數而非文章,並執行於操作者可透過標準 Kubernetes 工具自行掌控的基礎設施之上。
