Vận hành cụm Kubernetes trong môi trường Production đòi hỏi sự khắt khe về tính sẵn sàng cao (High Availability), kiểm soát tài nguyên CPU/RAM chặt chẽ và cấu hình Zero-Downtime Rolling Update.
1. Thiết Lập Resource Requests & Limits
Sai lầm phổ biến nhất của các kỹ sư khi triển khai ứng dụng lên Kubernetes là không khai báo hoặc khai báo sai requests và limits, dẫn đến hiện tượng OOMKilled hoặc CPU Throttling.
apiVersion: apps/v1
kind: Deployment
metadata:
name: showtech-api-service
namespace: production
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25%
maxUnavailable: 0
selector:
matchLabels:
app: showtech-api
template:
metadata:
labels:
app: showtech-api
spec:
containers:
- name: api
image: ghcr.io/showtech/api:v2.4.0
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "1000m"
memory: "1Gi"
readinessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /livez
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
2. Phân Biệt Readiness vs Liveness Probes
- Readiness Probe: Kiểm tra xem container đã sẵn sàng nhận traffic từ Service hay chưa.
- Liveness Probe: Phát hiện lỗi Deadlock để Kubelet tự động khởi động lại container.
Tuyệt đối không kiểm tra kết nối Database trong **Liveness Probe** vì nếu database bị ngắt tạm thời, toàn bộ Pod trong cụm sẽ bị restart đồng loạt, biến sự cố nhỏ thành sự cố sập hoàn toàn hệ thống!
3. Bảo Vệ Ứng Dụng Với PodDisruptionBudget (PDB)
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: showtech-api-pdb
namespace: production
spec:
minAvailable: 2
selector:
matchLabels:
app: showtech-api
4. Tổng Kết
Hãy luôn tự động hóa quy trình kiểm thử hạ tầng bằng CI/CD pipeline (GitLab CI / GitHub Actions) kết hợp công cụ ArgoCD GitOps để đạt chuẩn bảo mật và ổn định cao nhất.
Nhận xét
Đăng nhận xét