在 CentOS 上排查 Kubernetes(k8s)集群故障,一般遵循 “从整体到局部、从节点到 Pod、从网络到存储” 的思路。下面给你一套系统化排查流程 + 常用命令,适合生产环境实战。
kubectl get nodes
kubectl get nodes -o wide
重点看:
STATUS 是否为 ReadyROLES 是否正确(master / worker)如果节点 NotReady:
kubectl describe node <node-name>
systemctl status kubelet
journalctl -u kubelet -f
常见错误:
CentOS 常见:
systemctl status docker
docker info
systemctl status containerd
ctr ns ls
kubectl get pods -A
kubectl get pods -n <namespace>
常见状态:
PendingCrashLoopBackOffImagePullBackOffEvictedTerminatingkubectl describe pod <pod-name> -n <namespace>
重点看:
kubectl logs <pod-name> -n <namespace>
kubectl logs <pod-name> -c <container-name> -n <namespace>
kubectl logs <pod-name> --previous
kubectl exec -it <pod-name> -n <namespace> -- sh
如果进不去,说明:
kubectl get svc -A
kubectl describe svc <svc-name> -n <namespace>
检查:
kubectl get endpoints <svc-name> -n <namespace>
kubectl exec -it <pod> -- nslookup kubernetes.default
DNS 常见问题:
kubectl get pods -n kube-system
常见插件:
查看日志:
kubectl logs -n kube-system <cni-pod>
kubectl get componentstatuses
或:
kubectl get pods -n kube-system
重点组件:
kube-apiserverkube-controller-managerkube-schedulerjournalctl -u kube-apiserver
# 或
kubectl logs -n kube-system kube-apiserver-<node>
systemctl status etcd
etcdctl endpoint health
常见问题:
kubectl top node
kubectl top pod -A
Pod 事件:
kubectl describe pod <pod>
常见:
0/3 nodes are available: 3 Insufficient cpu
kubectl describe node | grep Taint
openssl x509 -in /etc/kubernetes/pki/apiserver.crt -text -noout | grep Not
或:
kubeadm certs check-expiration
证书过期会导致:
df -h
kubelet 默认要求:
/var/lib/docker 或 /var/lib/containerd/var/logfree -h
top
systemctl status firewalld
getenforce
建议:
permissive✅ Pod 起不来
kubectl describe podkubectl logs✅ Node NotReady
✅ 服务访问不通
✅ 集群直接不可用
kubectl get events --sort-by=.metadata.creationTimestamp
kubectl get pods -A -o wide
journalctl -u kubelet -f
kubectl describe node <node>
如果你愿意,可以直接把 具体报错信息或现象(比如某个 Pod 状态、kubelet 日志、NotReady 节点)贴出来,我可以 直接帮你定位原因并给出修复方案。