☰
LlamaIndex 集成 Valyu 工具:为 AI Agent 接入付费内容深度搜索与网页内容抽取
2026/10/11 7:38:10 网站建设 项目流程
  • 人工智能
  • RAG
  • 大模型

【免费下载链接】llama_index

LlamaIndex is the document processing platform for AI

项目地址:https://gitcode.com/GitHub_Trending/ll/llama_index
点击查看免费下载

Valyu 是面向 LLM 应用的内容检索平台,通过其 deep search API 让 Agent 既能检索互联网公开网页,也能访问经过商业授权的专有内容库。本文围绕 LlamaIndex 仓库中的ValyuToolSpec与ValyuRetriever展开,完整讲解如何安装、配置、调用搜索与内容抽取能力,并深入源码与测试验证每个参数的取值与行为。读完本文,你将能够把 Valyu 作为 Function Tool 接入FunctionAgent/OpenAIAgent,或作为检索器接入 LlamaIndex 的 RAG 检索管线。

Valyu 工具定位与仓库结构

在 LlamaIndex 的集成生态中,Valyu 工具属于llama-index-integrations/tools/下的独立包llama-index-tools-valyu(版本 0.7.0,见 pyproject.toml)。它的职责是让 Agent 通过 Valyu 的深度搜索与内容抽取 API,获得来自「程序化授权的专有内容」和「互联网公开内容」两大类来源的检索结果,并统一转换为 LlamaIndex 的Document数据对象,从而无缝衔接下游的索引、检索与问答流程。

整个模块的文件布局非常清晰:

文件职责
llama_index/tools/valyu/base.py定义ValyuToolSpec,封装search与get_contents两个工具函数
llama_index/tools/valyu/retriever.py定义ValyuRetriever,实现BaseRetriever接口,供检索管线直接使用
llama_index/tools/valyu/init.py导出ValyuToolSpec、ValyuRetriever
README.md官方用法示例
examples/context.py接入OpenAIAgent的完整示例
examples/retriever_example.py使用ValyuRetriever的完整示例
tests/test_tools_valyu.py覆盖初始化、搜索、内容抽取、检索器四类行为的测试
docs/api_reference/api_reference/tools/valyu.mdAPI 参考文档入口(mkdocstrings 自动生成ValyuToolSpec文档)

依赖方面,包要求valyu>=2.0.3作为底层 SDK、llama-index-core>=0.13.0,<0.15提供Document、BaseToolSpec等核心抽象,Python 版本要求>=3.10,<4.0。

安装与凭据准备

安装工具包及其运行时依赖:

pip install llama-index llama-index-core llama-index-tools-valyu

使用前需要获取 Valyu API Key,两种提供方式任选其一:

  1. 在 Valyu 开发者平台申请 API Key 后显式传入构造函数;
  2. 设置环境变量VALYU_API_KEY,这样即使不显式传参,底层 SDK 也能读取。
import os from llama_index.tools.valyu import ValyuToolSpec # 方式一:显式传入 valyu_tool = ValyuToolSpec(api_key=os.environ["VALYU_API_KEY"]) # 方式二:依赖环境变量(构造函数仍要求 api_key 参数非空)

需要说明的是,ValyuToolSpec.__init__对api_key有严格的校验逻辑(base.py):必须是「非空字符串」,否则立即抛出ValueError("api_key must be a non-empty string"),避免带着空凭据发起无意义的网络请求。

核心类 ValyuToolSpec 与工具化机制

ValyuToolSpec继承自llama_index.core.tools.tool_spec.base.BaseToolSpec,通过声明spec_functions = ["search", "get_contents"](base.py)告诉框架哪些方法可以被转换为 Agent 可调用的 Function Tool。

转换动作由BaseToolSpec.to_tool_list()完成(tool_spec/base.py):框架遍历spec_functions,用getattr取到对应方法,读取其 docstring 与函数签名生成ToolMetadata,最终包装成FunctionTool。这也解释了为什么search与get_contents的方法 docstring 写得非常详细——它们会直接作为工具描述暴露给 LLM,成为模型理解工具用途与参数语义的唯一依据。

最简单的接入方式如下(示例来自 README.md):

from llama_index.tools.valyu import ValyuToolSpec from llama_index.core.agent.workflow import FunctionAgent from llama_index.llms.openai import OpenAI import os valyu_tool = ValyuToolSpec( api_key=os.environ["VALYU_API_KEY"], max_price=100, # 默认是 100 美元 ) agent = FunctionAgent( tools=valyu_tool.to_tool_list(), llm=OpenAI(model="gpt-4.1"), ) print( await agent.run( "What are the implications of using different volatility calculation methods (EWMA vs. GARCH) in Value at Risk (VaR) modeling for fixed income portfolios?" ) )

search:深度搜索参数详解

search是 Valyu 工具的核心能力,用于「从专有与公开来源中搜索并检索相关内容」。它的参数分为初始化时固定与调用时灵活指定两类,二者界限在源码注释中明确给出。

初始化时固定的参数

在构造ValyuToolSpec时设置、随后不可在单次search调用中覆盖的参数如下(来自 base.py):

参数类型默认值说明
api_keystr必填Valyu API Key
verboseboolFalse开启后打印搜索响应/抽取响应日志
max_priceOptional[float]100单次搜索操作的最大成本上限(美元),None表示不限制
relevance_thresholdfloat0.5结果相关性得分的最低阈值,取值0.0–1.0
fast_modeOptional[bool]False快速模式:返回更快但更短的结果;传None时交由模型在每次搜索中自行决定
included_sourcesOptional[List[str]]None仅在这些 URL / 域名 / 数据集范围内搜索并返回结果
excluded_sourcesOptional[List[str]]None从搜索结果中排除的 URL / 域名 / 数据集列表
response_lengthOptional[Union[int, str]]None每个结果返回的字符数:预设值"short"(25k)、"medium"(50k)、"large"(100k)、"max"(完整内容),或直接传正整数表示自定义字符数
country_codeOptional[str]None两位字母 ISO 国家码(如"GB"、"US"),用于将搜索结果偏向特定国家
contents_summaryOptional[Union[bool, str, Dict]]None内容抽取的 AI 摘要配置(详见下文)
contents_extract_effortOptional[str]"normal"抽取彻底程度:"normal"快速、"high"更彻底但更慢、"auto"自动判断但最慢
contents_response_lengthOptional[Union[str, int]]"short"内容抽取时每个 URL 返回的字符数(同response_length的预设值语义,默认 25k)

调用时可指定的参数

def search( self, query: str, # 查询语句 search_type: str = "all", # "all" 专有+网页、"proprietary" 仅专有索引、"web" 仅网页 max_num_results: int = 5, # 最大结果数,取值范围 1–20 start_date: Optional[str] = None, # 时间过滤起始日期,YYYY-MM-DD end_date: Optional[str] = None, # 时间过滤结束日期,YYYY-MM-DD fast_mode: Optional[bool] = None, # 本次搜索是否启用快速模式 ) -> List[Document]

调用示例(来自 README.md):

results = valyu_tool.search( query="artificial intelligence trends 2024", included_sources=[ "arxiv.org", "nature.com", ], # 仅搜索学术来源 response_length="medium", # 每个结果 50k 字符 max_num_results=3, relevance_threshold=0.5, )

fast_mode 的优先级逻辑

fast_mode是少数「初始化与调用两个层面都有」的参数,源码(base.py)中的合并规则为:

  1. 若初始化时用户显式设置了fast_mode(非None),则无论模型传入什么都使用用户的值;
  2. 若用户设置为None(允许模型决定),则采用模型本次调用传入的值;
  3. 若两者都未指定,回退到 SDK 默认值False。

返回结果与元数据

search会把 SDK 返回的每个结果封装成 LlamaIndex 的Document,并透传以下元数据字段(base.py):

title、url、source、price(该条结果的价格)、length、data_type、relevance_score。

get_contents:网页内容抽取

get_contents调用 Valyu 的内容抽取 API,从给定 URL 中提取干净、结构化的正文内容。与搜索不同,它的所有抽取参数(summary、extract_effort、response_length)都在初始化时固定,模型调用时只能指定 URL 列表:

def get_contents(self, urls: List[str]) -> List[Document]

单次请求最多支持 10 个 URL。contents_summary参数提供了从「原始内容」到「结构化抽取」的四个档位(base.py):

取值行为
False/None不做任何 AI 处理,返回原始内容
True基础自动摘要
str自定义摘要指令(最多 500 字符)
dictJSON Schema,用于结构化抽取

get_contents返回的Document元数据更丰富(base.py):url、title、source、length、data_type、citation,以及可选字段summary、summary_success、image_url。citation字段为每条抽取内容提供可追溯的引用来源,适合需要输出引用格式的问答场景。

将 Valyu 接入 Agent 的两种方式

方式一:FunctionAgent(Workflow 风格)

即上文 README 示例,使用llama_index.core.agent.workflow.FunctionAgent,to_tool_list()返回的两个工具会被自动注册给模型,模型按需调用search或get_contents。

方式二:OpenAIAgent

仓库自带的 examples/context.py 展示了基于OpenAIAgent的接入,并且同时配置了搜索与内容抽取两组参数:

import os from llama_index.agent.openai import OpenAIAgent from llama_index.tools.valyu import ValyuToolSpec valyu_tool = ValyuToolSpec( api_key=os.environ["VALYU_API_KEY"], max_price=100, # 默认是 100 fast_mode=True, # 快速模式:更快但结果更短 # Contents API 配置 contents_summary=True, # 内容抽取开启 AI 摘要 contents_extract_effort="normal", contents_response_length="medium", ) agent = OpenAIAgent.from_tools( valyu_tool.to_tool_list(), verbose=True, ) # 搜索示例 search_response = agent.chat( "What are the key considerations and empirical evidence for implementing " "statistical arbitrage strategies using cointegrated pairs trading..." ) print(search_response) # URL 内容抽取示例 content_response = agent.chat( "Please extract and summarize the content from these URLs: " "https://arxiv.org/abs/1706.03762 and " "https://en.wikipedia.org/wiki/Transformer_(machine_learning_model)" ) print(content_response)

ValyuRetriever:URL 内容检索器

ValyuRetriever继承llama_index.core.base.base_retriever.BaseRetriever(retriever.py),把「网页内容抽取」包装成标准的检索器接口,可直接用于 LlamaIndex 的检索增强管线(如与RetrieverQueryEngine、后处理器组合使用)。

它的构造参数只包含内容抽取相关配置:api_key、verbose、contents_summary、contents_extract_effort、contents_response_length,以及可选的callback_manager用于追踪操作。

从查询串解析 URL

_retrieve的输入是QueryBundle,其query_str中应包含 URL(用空格或逗号分隔)。_parse_urls_from_query(retriever.py)的实现细节很有价值:

  • 用正则[,\s]+按空格/逗号切分查询串;
  • 仅保留以http://或https://开头的片段(自然语言中的非 URL 词会被自动过滤);
  • 最多截取前 10 个 URL(对齐 API 约束)。

因此用户可以用自然语言下达「请抽取这几个网页」的指令,检索器会自动定位其中的 URL,这在该模块的测试test_valyu_retriever_url_parsing(tests/test_tools_valyu.py)中有完整验证。

检索行为

每个抽取结果被包装为TextNode并放入NodeWithScore,相关性得分统一为1.0(retriever.py),元数据与get_contents保持一致。如果查询串中没有 URL,_retrieve直接返回空列表且不发起任何 API 调用(有专门的测试覆盖)。

examples/retriever_example.py 演示了单 URL、多 URL、以及自然语言夹杂 URL 三种用法:

import os from llama_index.tools.valyu import ValyuRetriever from llama_index.core import QueryBundle valyu_retriever = ValyuRetriever( api_key=os.environ.get("VALYU_API_KEY", "your-api-key-here"), verbose=True, contents_summary=True, # 开启 AI 摘要 contents_extract_effort="normal", contents_response_length="medium", ) # 单 URL 检索 query_bundle = QueryBundle( query_str="https://en.wikipedia.org/wiki/Transformer_(machine_learning_model)" ) nodes = valyu_retriever.retrieve(query_bundle) for node in nodes: print(node.node.metadata.get("title"), len(node.node.text), node.score) # 自然语言夹杂 URL natural_query = QueryBundle( query_str="Please extract content from these research papers: " "https://arxiv.org/abs/1706.03762 and also from " "https://en.wikipedia.org/wiki/Large_language_model" ) nodes = valyu_retriever.retrieve(natural_query)

参数校验规则速查

ValyuToolSpec.__init__内置了一整套参数校验(base.py),ValyuRetriever对内容抽取相关参数也有相同的校验。理解这些规则可以避免运行时异常:

参数校验规则
api_key非空字符串,否则抛ValueError
max_priceNone或非负数值
relevance_threshold必须为0.0–1.0之间的数值
verbose必须为bool
fast_mode必须为bool或None
included_sources/excluded_sourcesNone或字符串列表
response_length/contents_response_length预设值字符串之一(short/medium/large/max),或正整数(自定义字符数)
country_code恰好 2 位字母的 ISO 国家码
contents_summarybool、dict,或不超过 500 字符的字符串
contents_extract_effort必须是normal/high/auto之一

测试覆盖与质量保障

test_tools_valyu.py 通过 mockvalyu.Valyu客户端对三类行为做了系统性验证,可以作为理解工具语义的补充证据:

  • 初始化:验证类继承关系、构造参数是否正确传递到内部状态(test_init、test_init_with_user_controlled_params);
  • 搜索:验证用户配置的默认参数(阈值、价格上限、来源过滤、国家码)会被原样透传给 SDK,且模型无法覆盖这些「初始化级」参数(test_search_uses_user_defaults、test_search_model_can_control_allowed_params);同时验证fast_mode的用户/模型优先级(test_search_fast_mode_user_controls)、verbose 打印、多结果转多Document;
  • 内容抽取:验证默认参数、summary 配置、多 URL、空响应/None 响应的空列表处理(test_get_contents_*系列);
  • 检索器:验证 URL 解析边界(单 URL、逗号/空格分隔、自然语言混排、无 URL、超过 10 个 URL 截断)、空查询不调 API、NodeWithScore的得分与元数据(test_valyu_retriever_*系列)。

使用限制与注意事项

  • get_contents单次最多 10 个 URL,检索器解析查询串时同样截断到 10 个;
  • search的max_num_results取值范围为 1–20;
  • max_price是美元计价的操作成本上限,搜索结果中每个结果的price字段可用于核算单条成本;
  • start_date/end_date必须为YYYY-MM-DD格式;
  • 初始化时固定的搜索参数(max_price、relevance_threshold、included_sources、excluded_sources、response_length、country_code)无法在单次调用中修改,需要不同策略时应构造多个ValyuToolSpec实例;
  • contents_summary自定义指令上限 500 字符,超出会直接抛ValueError;
  • search_type仅支持"all"/"proprietary"/"web"三个取值,"proprietary"只检索 Valyu 自有索引,"web"只做网页搜索。

总体而言,Valyu 工具为 LlamaIndex Agent 补上了「付费专有内容 + 公开网页」双源检索的能力,且凭借BaseToolSpec与BaseRetriever两个标准抽象,可以零成本嵌入已有的 Agent 工作流或 RAG 管线,是面向金融研报、学术文献等高价值场景的实用集成组件。

  • 人工智能
  • RAG
  • 大模型

【免费下载链接】llama_index

LlamaIndex is the document processing platform for AI

项目地址:https://gitcode.com/GitHub_Trending/ll/llama_index
点击查看免费下载

相关推荐

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询