19/09/2026
Kubernetes ပေါ်မှာ AI Inference workloads တွေကို run ရုံတင်မကဘဲ တကယ့် GPU ကုန်ကျစရိတ်နဲ့ Token-level economics ကိုပါ ထိန်းကျောင်းနိုင်ဖို့ ဘာတွေလိုအပ်မလဲ။
ပုံမှန် Microservices တွေမှာဆိုရင် CPU နဲ့ RAM requests/limits တွေကို အခြေခံပြီး Kubernetes cluster ထဲ schedule လုပ်ရတာ အဆင်ပြေပါတယ်။ ဒါပေမဲ့ Large Language Models (LLMs) လို dynamic ဖြစ်တဲ့ AI workloads တွေ Production ထဲ ရောက်လာတဲ့အခါ static metrics တွေနဲ့တင် မလုံလောက်တော့တာကို တွေ့ရပါတယ်။
AI inference ရဲ့ သဘာဝအရ request mix တွေ၊ KV cache saturation နဲ့ prefill/decode phases တွေအပေါ် မူတည်ပြီး compute လိုအပ်ချက်က dynamic ဖြစ်နေတာပါ။ ပုံမှန် Kubernetes scheduling က ဒီလို GPU memory behavior တွေကို အတွင်းကျကျ မသိနိုင်တဲ့အတွက် GPU utilization အပြည့်မရဘဲ Cloud bills တွေ အဆမတန် မြင့်တက်လာတတ်ပါတယ်။
ဒီ Article မှာ Enterprise ကြီးတွေ (ဥပမာ China Merchants Bank လို heterogeneous accelerators ပေါင်း သောင်းနဲ့ချီ run နေတဲ့ architecture) မှာ ဒီပြဿနာကို ဘယ်လို ဖြေရှင်းထားသလဲဆိုတာ ရှင်းပြထားတာ တွေ့ရပါတယ်။ ရိုးရိုး Pod-level scheduling အဆင့်ကနေ specialized cloud-native ecosystem ဘက်ကို ပြောင်းလဲချဉ်းကပ်လာကြတာပါ။
အထူးသဖြင့် Quota queuing အတွက် Kueue၊ GPU sharing နဲ့ slicing အတွက် HAMi၊ event-driven autoscaling အတွက် KEDA နဲ့ dataset caching အတွက် Fluid စတဲ့ specialized tools တွေကို တွဲသုံးပြီး GPU utilization ကို အမြင့်ဆုံးရအောင် ဆွဲတင်လာကြပါတယ်။ ဒါမှလည်း latency မကျစေဘဲ Token တစ်ခုချင်းစီရဲ့ ကုန်ကျစရိတ် (cost per token) ကို တိတိကျကျ တိုင်းတာထိန်းညှိနိုင်မှာ ဖြစ်ပါတယ်။
AI application တွေကို production မှာ scale လုပ်ဖို့ ပြင်ဆင်နေတဲ့ DevOps နဲ့ Platform Engineer တွေအတွက် GPU memory management နဲ့ intelligent queueing architecture အကြောင်း ပြန်လည်ဆန်းစစ်ကြည့်ဖို့ အတော်လေး စဉ်းစားစရာ ကောင်းပါတယ်။
ကိုယ့်ရဲ့ cluster ထဲမှာ AI model တွေ host လုပ်ဖို့ ပြင်ဆင်နေတယ်ဆိုရင်တော့ ဒီ article ထဲက cluster scheduling model တွေနဲ့ FinOps ချဉ်းကပ်ပုံတွေကို ဖတ်ကြည့်သင့်တဲ့ reference တစ်ခုအဖြစ် မျှဝေပေးလိုက်ပါတယ်။
Reference:
https://thenewstack.io/kubernetes-ai-inference-costs/