排查 Debian 上 Kubernetes(k8s)故障通常要从系统层 → 容器运行时 → k8s 组件 → 应用层逐层定位。下面给你一套实用排查思路 + 常用命令。
# 系统负载、内存、磁盘
top
free -h
df -h
k8s 对时间敏感:
timedatectl status
若不同步:
sudo apt install chrony -y
sudo systemctl enable --now chrony
Debian 常见:
sudo ufw status
iptables -L -n -v
k8s 需要端口(如 6443、10250 等)放行。
sysctl net.bridge.bridge-nf-call-iptables
应为 1,否则:
cat <<EOF > /etc/sysctl.d/k8s.conf
net.bridge.bridge-nf-call-iptables=1
net.ipv4.ip_forward=1
EOF
sysctl --system
systemctl status containerd
journalctl -u containerd -f
systemctl status docker
journalctl -u docker -f
测试拉镜像:
crictl pull nginx
systemctl status kubelet
journalctl -u kubelet -f
常见错误:
systemctl status kube-apiserver
systemctl status kube-controller-manager
systemctl status kube-scheduler
查看日志:
journalctl -u kube-apiserver -f
kubectl get nodes
kubectl describe node <node-name>
常见状态:
NotReadyDiskPressureMemoryPressurekubectl get pods -A
kubectl get pods -n <namespace>
kubectl describe pod <pod-name> -n <ns>
重点看:
kubectl logs <pod-name> -n <ns>
kubectl logs <pod-name> -c <container> -n <ns>
kubectl exec -it <pod> -- sh
kubectl run test --rm -it --image=busybox -- sh
nslookup kubernetes.default
kubectl get pods -n kube-system
示例(Flannel):
kubectl logs -n kube-system kube-flannel-ds-xxxx
kubeadm certs check-expiration
kubectl config view
节点 NotReady → kubelet / 网络 / 资源
Pod 起不来 → 镜像 / 资源 / 配置
API 不通 → apiserver / 证书 / 网络
DNS 失败 → CoreDNS / 网络插件
# 看所有组件日志
journalctl -u kubelet -u containerd -u docker -f
# 重置节点(谨慎)
kubeadm reset
如果你愿意,可以告诉我:
我可以帮你精准定位某一类故障。