温馨提示×

怎样在Debian配Python爬虫

小樊
38
2026-10-05 03:09:02
栏目: 编程语言

下面按从零开始的思路,说明在 Debian 上配置 Python 爬虫环境 的常见做法(适用于 Debian 11 / 12,也基本适用于 Ubuntu)。


一、更新系统

sudo apt update
sudo apt upgrade -y

二、安装 Python 与 pip

Debian 通常自带 Python3,但建议确认并安装 pip:

sudo apt install -y python3 python3-pip python3-venv

查看版本:

python3 --version
pip3 --version

三、创建虚拟环境(强烈推荐)

避免污染系统 Python:

mkdir ~/spider
cd ~/spider
python3 -m venv venv
source venv/bin/activate

激活后提示符前会出现 (venv)。

退出环境:

deactivate

四、安装常用爬虫库

在虚拟环境中安装:

pip install requests beautifulsoup4 lxml

进阶可选:

pip install selenium scrapy playwright

五、示例:简单爬虫

1️⃣ 使用 requests + BeautifulSoup

spider.py

import requests
from bs4 import BeautifulSoup

url = "https://example.com"
resp = requests.get(url)
soup = BeautifulSoup(resp.text, "lxml")

print(soup.title.text)

运行:

python spider.py

六、处理常见依赖问题

1️⃣ SSL / 证书错误

sudo apt install -y ca-certificates

2️⃣ 中文乱码

resp.encoding = resp.apparent_encoding

3️⃣ 需要浏览器(Selenium / Playwright)

sudo apt install -y chromium
playwright install

七、后台运行爬虫

nohup python spider.py > log.txt 2>&1 &

或用 screen:

sudo apt install screen
screen -S spider
python spider.py

八、定时运行(可选)

crontab -e

示例(每天 3 点):

0 3 * * * /home/user/spider/venv/bin/python /home/user/spider/spider.py

九、合规建议(重要)

  • 遵守 robots.txt
  • 控制请求频率
  • 不爬敏感 / 私有数据
  • 设置 User-Agent
headers = {
    "User-Agent": "Mozilla/5.0"
}

如果你愿意,可以告诉我:

  • 爬 哪个网站
  • 用 requests / scrapy / selenium
  • 是否要 定时 / 分布式

我可以直接帮你写一套可用的爬虫配置。

0 踩