在Python爬虫过程中怎么使用代理IP

发布时间：2021-04-29 11:09:11 来源：亿速云阅读：210 作者：小新栏目：编程语言

这篇文章主要介绍了在Python爬虫过程中怎么使用代理IP，具有一定借鉴价值，感兴趣的朋友可以参考下，希望大家阅读完这篇文章之后大有收获，下面让小编带着大家一起了解一下。

python是什么意思

Python是一种跨平台的、具有解释性、编译性、互动性和面向对象的脚本语言，其最初的设计是用于编写自动化脚本，随着版本的不断更新和新功能的添加，常用于用于开发独立的项目和大型项目。

许多网站会在一定时间内检测到某个IP的访问次数(通过流量统计、系统日志等)，如果访问次数多得不像正常人，就会禁止该IP的访问。因此，我们可以设置一些代理服务器，每隔一段时间更换一个代理，即使IP被禁止，仍然可以更换IP继续爬行。

1、设置代理服务器

通过ProxyHandler在request中设置使用代理服务器，代理的使用非常简单，可以在专业网站上购买稳定的ip地址，也可以在网上寻找免费的ip代理。

免费开放代理基本没有成本。我们可以在一些代理网站上收集这些免费代理。如果测试后可以使用，我们可以在爬虫上收集它们。

2、selenium使用代理ip

selenium在使用带有用户名和密码的代理ip时，不能使用无头模式。

def create_proxy_auth_extension(proxy_host, proxy_port,
                                proxy_username, proxy_password,
                                scheme='http', plugin_path=None):
    if plugin_path is None:
        plugin_path = r'./proxy_auth_plugin.zip'
 
    manifest_json = """
        {
            "version": "1.0.0",
            "manifest_version": 2,
            "name": "Chrome Proxy",
            "permissions": [
                "proxy",
                "tabs",
                "unlimitedStorage",
                "storage",
                "<all_urls>",
                "webRequest",
                "webRequestBlocking"
            ],
            "background": {
                "scripts": ["background.js"]
            },
            "minimum_chrome_version":"22.0.0"
        }
        """
 
    background_js = string.Template(
        """
        var config = {
            mode: "fixed_servers",
            rules: {
                singleProxy: {
                    scheme: "${scheme}",
                    host: "${host}",
                    port: parseInt(${port})
                },
                bypassList: ["foobar.com"]
            }
          };
 
        chrome.proxy.settings.set({value: config, scope: "regular"}, function() {});
 
        function callbackFn(details) {
            return {
                authCredentials: {
                    username: "${username}",
                    password: "${password}"
                }
            };
        }
 
        chrome.webRequest.onAuthRequired.addListener(
            callbackFn,
            {urls: ["<all_urls>"]},
            ['blocking']
        );
 
        chrome.webRequest.onBeforeSendHeaders.addListener(function (details) {
        details.requestHeaders.push({name:"connection",value:"close"});
            return {
            requestHeaders: details.requestHeaders
            };
        },
        {urls: ["<all_urls>"]},
            ['blocking']
        );
        """
    ).substitute(
        host=proxy_host,
        port=proxy_port,
        username=proxy_username,
        password=proxy_password,
        scheme=scheme,
    )
 
    with zipfile.ZipFile(plugin_path, 'w') as zp:
        zp.writestr("manifest.json", manifest_json)
        zp.writestr("background.js", background_js)
 
    return plugin_path
chrome_options = webdriver.ChromeOptions()
proxy_auth_plugin_path = create_proxy_auth_extension(
    proxy_host=proxyHost,
    proxy_port=proxyPort,
    proxy_username=proxyUser,
    proxy_password=proxyPass)
chrome_options.add_extension(proxy_auth_plugin_path)
driver = webdriver.Chrome(chrome_options=chrome_options)

感谢你能够认真阅读完这篇文章，希望小编分享的“在Python爬虫过程中怎么使用代理IP”这篇文章对大家有帮助，同时也希望大家多多支持亿速云，关注亿速云行业资讯频道，更多相关知识等着你来学习!

向AI问一下细节

在Python爬虫过程中怎么使用代理IP

python是什么意思

猜你喜欢

最新资讯

相关推荐

相关标签