在 Debian 上排查 Kubernetes(k8s)故障,通常遵循**“自底向上 + 分层定位”的思路:先确认系统/网络,再确认 k8s 组件,最后看工作负载。下面给你一套实用排查清单**。
uname -a
cat /etc/os-release
free -h
df -h
top / htop
k8s 对时间敏感:
timedatectl status
# 如有问题
apt install chrony
systemctl enable --now chrony
sysctl net.bridge.bridge-nf-call-iptables
sysctl net.ipv4.ip_forward
lsmod | grep br_netfilter
# 若无
modprobe br_netfilter
Debian 通常无 SELinux,但注意:
iptables -L -n -v
ufw status
k8s 节点之间需开放:
which docker containerd crictl
crictl info
systemctl status containerd
journalctl -u containerd -f
systemctl status kubelet
journalctl -u kubelet -f
常见错误:
systemctl status kube-apiserver
systemctl status kube-controller-manager
systemctl status kube-scheduler
如是用 kubeadm:
docker ps | grep kube
# 或
crictl ps
kubectl cluster-info
kubectl get nodes -o wide
kubectl get cs
kubectl describe node <node-name>
重点看:
kubectl get pods -A
kubectl describe pod <pod> -n <ns>
kubectl logs <pod> -n <ns>
kubectl logs <pod> -p -n <ns> # 前一个容器
| 状态 | 可能原因 |
|---|---|
| Pending | 调度失败 / 资源不足 |
| CrashLoopBackOff | 应用启动失败 |
| ImagePullBackOff | 镜像拉取失败 |
| Evicted | 节点资源紧张 |
kubectl get svc,ep -n <ns>
kubectl get pods -n kube-system | grep kube-proxy
kubectl exec -it <pod> -- nslookup kubernetes.default
检查 CoreDNS:
kubectl get pods -n kube-system
kubectl logs -n kube-system coredns-xxx
kubeadm certs check-expiration
kubectl config view
cat ~/.kube/config
| 故障现象 | 排查点 |
|---|---|
| Node NotReady | kubelet / 资源 / 网络 |
| Pod 一直 Pending | 调度 / 资源 |
| 集群无法访问 | API Server / 证书 |
| 网络不通 | CNI / kube-proxy |
| 升级后异常 | 版本兼容 / 配置变更 |
如果你愿意,可以告诉我:
我可以直接帮你定位具体原因 + 给修复命令。