爬虫爬取贴吧评论

介绍一个小项目，通过爬虫爬取贴吧评论并保存到本地

导入相关的库

import csv
import requests
import re
import time

打开百度贴吧随便选择一篇帖子，检查网页源代码，我们发现评论信息都在源代码中，所以直接解析网页源代码即可

用`requests`发送请求

url = 'https://tieba.baidu.com/p/9285638137?pn=1'
headers = {
'user-agent':
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36'
}
resp = requests.get(url,headers = headers)

我们用requests模拟浏览器向目标url发送了一次请求并且拿到了服务器的响应resp。下一步的工作就是分析resp中的内容，拿到我们想要的数据。

解析数据

html = resp.text
#通过正则表达式获取需要的内容
# 评论内容
comments = re.findall('style="display:;">                    (.*?)</div>',html)
# 评论用户
users = re.findall('class="p_author_name j_user_card" href=".*?" target="_blank">(.*?)</a>',html)
# 评论时间
comment_times = re.findall('楼</span><span class="tail-info">(.*?)</span><div',html)

保存到本地

with open('01.csv', 'a', encoding='utf-8') as f:
    # 将文件加载到csv对象中
    csvwriter = csv.writer(f)
    for u, c, t in zip(users, comments, comment_times):
        # 筛选数据,过滤掉异常数据
        if 'img' in c or 'div' in c or len(u) > 50:
            continue
        csvwriter.writerow((u,t,c))
        print(u, t, c)
    csvwriter.writerow(('评论用户', '评论时间', '评论内容'))

以下是完整代码

import csv
import requests
import re
import time

url = 'https://tieba.baidu.com/p/9285638137?pn=1'
headers = {
'user-agent':
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36'
}
resp = requests.get(url,headers = headers)

html = resp.text
#通过正则表达式获取需要的内容
# 评论内容
comments = re.findall('style="display:;">                    (.*?)</div>',html)

# 评论用户
users = re.findall('class="p_author_name j_user_card" href=".*?" target="_blank">(.*?)</a>',html)
# 评论时间
comment_times = re.findall('楼</span><span class="tail-info">(.*?)</span><div',html)


with open('01.csv', 'a', encoding='utf-8') as f:
    # 将文件加载到csv对象中
    csvwriter = csv.writer(f)
    for u, c, t in zip(users, comments, comment_times):
        # 筛选数据,过滤掉异常数据
        if 'img' in c or 'div' in c or len(u) > 50:
            continue
        csvwriter.writerow((u,t,c))
        print(u, t, c)
    csvwriter.writerow(('评论用户', '评论时间', '评论内容'))

我们发现这样只能拿到第一页的评论内容
检查网页url：https://tieba.baidu.com/p/9285638137?pn=1
翻页之后网页url变成了：https://tieba.baidu.com/p/9285638137?pn=2
区别在于pn=1/2
我们可以修改代码循环遍历所有页面，代码如下

import csv
import requests
import re
import time

def main(page):
    url = f'https://tieba.baidu.com/p/9285638137?pn={page}'
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36'
    }
    resp = requests.get(url,headers=headers)
    html = resp.text
    # 评论内容
    comments = re.findall('style="display:;">                    (.*?)</div>',html)
    # 评论用户
    users = re.findall('class="p_author_name j_user_card" href=".*?" target="_blank">(.*?)</a>',html)
    # 评论时间
    comment_times = re.findall('楼</span><span class="tail-info">(.*?)</span><div',html)
    for u,c,t in zip(users,comments,comment_times):
        # 筛选数据,过滤掉异常数据
        if 'img' in c or 'div' in c or len(u)>50:
            continue
        csvwriter.writerow((u,t,c))
        print(u,t,c)
    print(f'第{page}页爬取完毕')

if __name__ == '__main__':
    with open('01.csv','a',encoding='utf-8')as f:
        # 将文件加载到csv对象中
        csvwriter = csv.writer(f)
        csvwriter.writerow(('评论用户','评论时间','评论内容'))
        for page in range(1,8):  # 爬取前7页的内容
            main(page)
            time.sleep(2)