它是一个用来解析HTML文档和XML文档的库, 它能够将那些复杂的HTML文档, 转换成一个那个树形结构, 而且这个结构也是比较复杂的, 使得我们能够在操作的时候, 去查找、去提取那些所需的内容, 下面我将要非常详细地去讲述它的使用流程, 并且还会结合一些实际的示例来进行说明。
一、安装,以及关于解析器的内容, 这就是第一步的基础安装环节。
pip install beautifulsoup4
pip install lxml # 推荐使用的解析器,也可以使用html.parser(Python内置)
2. 我们来创建对象, 然后解析本地的HTML文件。
from bs4 import BeautifulSoup
# 解析本地HTML文件
with open('example.html', 'r', encoding='utf-8') as f:
soup = BeautifulSoup(f, 'lxml')
(2) 把通过网络抓取得到的HTML代码进行解析, 弄明白里面的内容。
import requests
from bs4 import BeautifulSoup
url = 'https://example.com'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'lxml')
(3) 把那个叫做soup的对象, 直接拿去进行打印操作。
print(soup) # 打印整个HTML文档
print(soup.prettify()) # 格式化输出,更易读
二、基本查找方法中的第一类情况, 是通过查看标签名字来进行查找的。
# 获取第一个a标签
first_a_tag = soup.a
print(first_a_tag)
# 获取第一个p标签
first_p_tag = soup.p
print(first_p_tag)
请注意, 所采取的这样的方式仅仅能够获取文档里面第一个匹配到的标签。
2. 获取该标签的各个属性信息, 然后进行操作。
# 获取所有属性和属性值(返回字典)
a_attrs = soup.a.attrs
print(a_attrs) # 例如:{'href': 'https://example.com', 'class': ['external']}
# 获取特定属性值
href_value = soup.a['href'] # 等同于 soup.a.attrs['href']
print(href_value)
# 安全获取属性(属性不存在时返回None)
non_existent_attr = soup.a.get('nonexistent')
print(non_existent_attr) # 输出: None
3. 获取标签内容
# 假设有以下HTML片段
html_doc = """
直接文本内容
嵌套文本内容内部span后面文本
""" soup = BeautifulSoup(html_doc, 'lxml') # string属性(只获取直系文本) p1_text = soup.find('p').string print(p1_text) # 输出: 直接文本内容 p2_text = soup.find_all('p')[1].string print(p2_text) # 输出: None(因为有嵌套标签) # text/get_text()方法(获取所有文本内容) p2_full_text = soup.find_all('p')[1].text print(p2_full_text) # 输出: 嵌套文本内容内部span后面文本 p2_full_text_alt = soup.find_all('p')[1].get_text() print(p2_full_text_alt) # 同上
三、关于那个高级别的查找途径的第一个手段, 便是使用叫做 find() 的这种方法。
# 查找第一个a标签
first_a = soup.find('a')
# 查找具有特定属性的标签
a_with_title = soup.find('a', title='example')
a_with_class = soup.find('a', class_='external') # 注意class是Python关键字,所以要加下划线
a_with_id = soup.find('div', id='header')
# 使用多个条件查找
specific_a = soup.find('a', {'class': 'external', 'title': 'Example'})
2. () 方法
# 查找所有a标签
all_a_tags = soup.find_all('a')
# 查找多种标签
all_a_and_p = soup.find_all(['a', 'p'])
# 限制返回数量
first_two_a = soup.find_all('a', limit=2)
# 使用属性过滤
external_links = soup.find_all('a', class_='external')
specific_links = soup.find_all('a', {'data-category': 'news'})
# 使用函数过滤
def has_href_but_no_class(tag):
return tag.has_attr('href') and not tag.has_attr('class')
custom_filter_links = soup.find_all(has_href_but_no_class)
3. 该()方法是通过CSS选择器来进行匹配的。
# 通过id选择
element = soup.select('#header') # 返回列表
# 通过class选择
elements = soup.select('.external-link') # 所有class="external-link"的元素
# 通过标签选择
all_p = soup.select('p')
# 层级选择器
# 后代选择器(空格分隔)
descendants = soup.select('div p') # div下的所有p标签(不限层级)
# 子选择器(>分隔)
children = soup.select('div > p') # 直接子级p标签
# 组合选择
complex_selection = soup.select('div.content > p.intro + p.highlight')
四、在实战示例的第1个环节中, 所完成的工作是提取所有的链接。
from bs4 import BeautifulSoup
import requests
url = 'https://example.com'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'lxml')
# 提取所有a标签的href属性
links = [a['href'] for a in soup.find_all('a') if a.has_attr('href')]
print("页面中的所有链接:")
for link in links:
print(link)
从内容当中把那个新闻的标题拿过来, 然后再把摘要也提取出来。
现在呢, 我们来假定一个HTML结构是这样的。
新闻标题1
新闻摘要1...
新闻标题2
新闻摘要2...
提取代码:
news_items = []
for item in soup.select('.news-item'):
title = item.select_one('.title').text
summary = item.select_one('.summary').text
news_items.append({'title': title, 'summary': summary})
print("新闻列表:")
for news in news_items:
print(f"标题: {news['title']}")
print(f"摘要: {news['summary']}\n")
从所给的表格之中, 将里面的那些具体数据给提取出来。
假设有HTML表格:
[PROTECTED_c284d179d6b2fa7d64feafcad2aa7d8c]
提取代码:
table_data = []
table = soup.find('table', {'id': 'data-table'})
# 提取表头
headers = [th.text for th in table.find_all('th')]
# 提取表格内容
for row in table.find_all('tr')[1:]: # 跳过表头行
cells = row.find_all('td')
row_data = {headers[i]: cell.text for i, cell in enumerate(cells)}
table_data.append(row_data)
print("表格数据:")
for row in table_data:
print(row)
五、注意事项与技巧
编码方面的问题。
性能考虑:
对于解析器的选择这一事项。
处理不完整HTML:
# 使用html5lib解析不完整的HTML
soup = BeautifulSoup(broken_html, 'html5lib')
对那份文件的结构进行修改。
# 修改标签内容
tag = soup.find('div')
tag.string = "新内容"
# 添加新标签
new_tag = soup.new_tag('a', href='http://example.com')
new_tag.string = "新链接"
soup.body.append(new_tag)
# 删除标签
tag.decompose() # 完全删除
tag.extract() # 从树中移除但保留
需要对那些注解以及拥有特殊含义的字符串进行处理。
for comment in soup.find_all(text=lambda text: isinstance(text, Comment)):
print(comment)
六、常见问题解答
Q1:find()和()有什么区别?
A1:
关于动态加载的内容, 具体要如何去处理。
A2:
它仅仅能够解析那种静态的HTML代码, 可是呢, 对于那些需要动态加载才能显示出来的内容。
通过使用诸如等这些工具, 来获取那个已经完全渲染完毕之后的页面的相关信息。
通过对网站的API接口进行分析, 然后直接从这个地方把数据给获取过来。
为什么在很多情况下, 无法获得我们所希望得到的那些内容呢。
A3:
其有可能产生的原因, 在于:
这个网页里面使用了动态的方式, 去生成相关的内容。
这些标签是存在隐藏条件的, 比如说, 样式方面可能是设置了隐藏状态这样的代码表现(例如 style=“:none” 这种情况)。
这个选择器不够精准, 意外地匹配到了其他的元素。
解决方案:
你需要去检查网页源代码, 从而确认元素是否存在。
使用选择器的时候, 需要更加精确一些。