☰
BeautifulSoup实战:从入门到网页数据提取
2026/10/1 10:36:26 网站建设 项目流程

它是一个用来解析HTML文档和XML文档的库, 它能够将那些复杂的HTML文档, 转换成一个那个树形结构, 而且这个结构也是比较复杂的, 使得我们能够在操作的时候, 去查找、去提取那些所需的内容, 下面我将要非常详细地去讲述它的使用流程, 并且还会结合一些实际的示例来进行说明。

一、安装,以及关于解析器的内容, 这就是第一步的基础安装环节。

pip install beautifulsoup4

pip install lxml # 推荐使用的解析器,也可以使用html.parser(Python内置)

2. 我们来创建对象, 然后解析本地的HTML文件。

from bs4 import BeautifulSoup

# 解析本地HTML文件

with open('example.html', 'r', encoding='utf-8') as f:

soup = BeautifulSoup(f, 'lxml')

(2) 把通过网络抓取得到的HTML代码进行解析, 弄明白里面的内容。

import requests

from bs4 import BeautifulSoup

url = 'https://example.com'

response = requests.get(url)

soup = BeautifulSoup(response.text, 'lxml')

(3) 把那个叫做soup的对象, 直接拿去进行打印操作。

print(soup) # 打印整个HTML文档

print(soup.prettify()) # 格式化输出,更易读

二、基本查找方法中的第一类情况, 是通过查看标签名字来进行查找的。

# 获取第一个a标签

first_a_tag = soup.a

print(first_a_tag)

# 获取第一个p标签

first_p_tag = soup.p

print(first_p_tag)

请注意, 所采取的这样的方式仅仅能够获取文档里面第一个匹配到的标签。

2. 获取该标签的各个属性信息, 然后进行操作。

# 获取所有属性和属性值(返回字典)

a_attrs = soup.a.attrs

print(a_attrs) # 例如:{'href': 'https://example.com', 'class': ['external']}

# 获取特定属性值

href_value = soup.a['href'] # 等同于 soup.a.attrs['href']

print(href_value)

# 安全获取属性(属性不存在时返回None)

non_existent_attr = soup.a.get('nonexistent')

print(non_existent_attr) # 输出: None

3. 获取标签内容

# 假设有以下HTML片段

html_doc = """

直接文本内容

嵌套文本内容内部span后面文本

""" soup = BeautifulSoup(html_doc, 'lxml') # string属性(只获取直系文本) p1_text = soup.find('p').string print(p1_text) # 输出: 直接文本内容 p2_text = soup.find_all('p')[1].string print(p2_text) # 输出: None(因为有嵌套标签) # text/get_text()方法(获取所有文本内容) p2_full_text = soup.find_all('p')[1].text print(p2_full_text) # 输出: 嵌套文本内容内部span后面文本 p2_full_text_alt = soup.find_all('p')[1].get_text() print(p2_full_text_alt) # 同上

三、关于那个高级别的查找途径的第一个手段, 便是使用叫做 find() 的这种方法。

# 查找第一个a标签

first_a = soup.find('a')

# 查找具有特定属性的标签

a_with_title = soup.find('a', title='example')

a_with_class = soup.find('a', class_='external') # 注意class是Python关键字,所以要加下划线

a_with_id = soup.find('div', id='header')

# 使用多个条件查找

specific_a = soup.find('a', {'class': 'external', 'title': 'Example'})

2. () 方法

# 查找所有a标签

all_a_tags = soup.find_all('a')

# 查找多种标签

all_a_and_p = soup.find_all(['a', 'p'])

# 限制返回数量

first_two_a = soup.find_all('a', limit=2)

# 使用属性过滤

external_links = soup.find_all('a', class_='external')

specific_links = soup.find_all('a', {'data-category': 'news'})

# 使用函数过滤

def has_href_but_no_class(tag):

return tag.has_attr('href') and not tag.has_attr('class')

custom_filter_links = soup.find_all(has_href_but_no_class)

3. 该()方法是通过CSS选择器来进行匹配的。

# 通过id选择

element = soup.select('#header') # 返回列表

# 通过class选择

elements = soup.select('.external-link') # 所有class="external-link"的元素

# 通过标签选择

all_p = soup.select('p')

# 层级选择器

# 后代选择器(空格分隔)

descendants = soup.select('div p') # div下的所有p标签(不限层级)

# 子选择器(>分隔)

children = soup.select('div > p') # 直接子级p标签

# 组合选择

complex_selection = soup.select('div.content > p.intro + p.highlight')

四、在实战示例的第1个环节中, 所完成的工作是提取所有的链接。

from bs4 import BeautifulSoup

import requests

url = 'https://example.com'

response = requests.get(url)

soup = BeautifulSoup(response.text, 'lxml')

# 提取所有a标签的href属性

links = [a['href'] for a in soup.find_all('a') if a.has_attr('href')]

print("页面中的所有链接:")

for link in links:

print(link)

从内容当中把那个新闻的标题拿过来, 然后再把摘要也提取出来。

现在呢, 我们来假定一个HTML结构是这样的。

新闻标题1

新闻摘要1...

新闻标题2

新闻摘要2...

提取代码:

news_items = []

for item in soup.select('.news-item'):

title = item.select_one('.title').text

summary = item.select_one('.summary').text

news_items.append({'title': title, 'summary': summary})

print("新闻列表:")

for news in news_items:

print(f"标题: {news['title']}")

print(f"摘要: {news['summary']}\n")

从所给的表格之中, 将里面的那些具体数据给提取出来。

假设有HTML表格:

[PROTECTED_c284d179d6b2fa7d64feafcad2aa7d8c]

提取代码:

table_data = []

table = soup.find('table', {'id': 'data-table'})

# 提取表头

headers = [th.text for th in table.find_all('th')]

# 提取表格内容

for row in table.find_all('tr')[1:]: # 跳过表头行

cells = row.find_all('td')

row_data = {headers[i]: cell.text for i, cell in enumerate(cells)}

table_data.append(row_data)

print("表格数据:")

for row in table_data:

print(row)

五、注意事项与技巧

编码方面的问题。

性能考虑:

对于解析器的选择这一事项。

处理不完整HTML:

# 使用html5lib解析不完整的HTML

soup = BeautifulSoup(broken_html, 'html5lib')

对那份文件的结构进行修改。

# 修改标签内容

tag = soup.find('div')

tag.string = "新内容"

# 添加新标签

new_tag = soup.new_tag('a', href='http://example.com')

new_tag.string = "新链接"

soup.body.append(new_tag)

# 删除标签

tag.decompose() # 完全删除

tag.extract() # 从树中移除但保留

需要对那些注解以及拥有特殊含义的字符串进行处理。

for comment in soup.find_all(text=lambda text: isinstance(text, Comment)):

print(comment)

六、常见问题解答

Q1:find()和()有什么区别?

A1:

关于动态加载的内容, 具体要如何去处理。

A2:

它仅仅能够解析那种静态的HTML代码, 可是呢, 对于那些需要动态加载才能显示出来的内容。

通过使用诸如等这些工具, 来获取那个已经完全渲染完毕之后的页面的相关信息。

通过对网站的API接口进行分析, 然后直接从这个地方把数据给获取过来。

为什么在很多情况下, 无法获得我们所希望得到的那些内容呢。

A3:

其有可能产生的原因, 在于:

这个网页里面使用了动态的方式, 去生成相关的内容。

这些标签是存在隐藏条件的, 比如说, 样式方面可能是设置了隐藏状态这样的代码表现(例如 style=“:none” 这种情况)。

这个选择器不够精准, 意外地匹配到了其他的元素。

解决方案:

你需要去检查网页源代码, 从而确认元素是否存在。

使用选择器的时候, 需要更加精确一些。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询