☰
你的第一个 Elastic Agent:从 ES|QL 查询到 Kibana 中 AI 聊天(二)
2026/10/1 7:31:29 网站建设 项目流程

1. 从单条 ES|QL 到 Kibana AI 聊天:链路卡在哪

你已经能在 Kibana 的 Dev Tools 里跑通一条 ES|QL 查询,结果表格也正常返回。接下来想做的事很自然:把这条查询变成 AI 能调用的能力,然后在 Kibana 的聊天面板里用自然语言问它,让它自己去查、去汇总。这一步的难点不在 ES|QL 语法,而在于三件事没打通——查询没有被注册成 Tool、Agent 没有绑定这个 Tool、聊天入口没有指向正确的 Agent ID。

我试过直接拿原始查询去问聊天面板,结果它只会泛泛而谈,因为它根本不知道有这张表、这个字段。Elastic Agent Builder 的设计思路是把「逻辑」「技能」「大脑」拆成三层:ES|QL 查询是逻辑,注册成 Tool 是技能,Agent 是带人格和指令的大脑。三者缺一,聊天链路就是断的。

这篇面向已经跑通单条查询的开发者,交付可复制的 Agent 配置片段、MCP 连接参数,以及在 Kibana 里验证「查询→对话」链路是否生效的具体动作。适合谁:手上有 Elastic Stack 9.2 及以上、已经开启 Agent Builder 相关开关、想把自己的业务查询接进 AI 聊天的人。读完你能得到一个能对话的 Financial Assistant 示例,也能把这套流程套到自己的索引上。

需要说明的是,本文的 AI 能力调用走的是标准 HTTP 接口,如果你希望把模型调用统一到一个可管理的入口,可以用 TaoToken 这类兼容 OpenAI 协议的服务来承接,后面配置章节会给到具体参数。

2. 前置准备:开启 Agent Builder 与 MCP 开关

在写任何配置之前,先把 Kibana 侧的功能开关打开。Elastic Stack 9.2 之后,Agents 界面默认不显示,需要通过 Kibana 的内部设置接口开启。这一步用 Dev Tools 的 Console 执行即可,注意是kbn://协议。

POST kbn://internal/kibana/settings { "changes": { "agentBuilder:enabled": true } }

接着开启 onechat 相关的四个开关,它们分别控制 MCP、A2A、API 和 UI:

POST kbn://internal/kibana/settings { "changes": { "onechat:mcp:enabled": true, "onechat:a2a:enabled": true, "onechat:api:enabled": true, "onechat:ui:enabled": true } }

执行完刷新页面,左侧导航应该能看到 Agents 入口。如果没出现,检查 Kibana 版本是否低于 9.2,低版本没有这套接口。

环境变量方面,建议用一个.env文件集中管理连接信息,避免把 Key 硬编码进 Notebook。下面是我用的结构,你需要替换成自己的地址和 Key:

es_url="https://localhost:9200" kb_url="http://localhost:5601" es_api_key="你的Base64编码APIKey"

读取时用 python-dotenv 加载,打印时只显示 Key 末五位,避免泄露:

import os from dotenv import load_dotenv load_dotenv() es_url = os.getenv("es_url") kb_url = os.getenv("kb_url") es_api_key = os.getenv("es_api_key") print(f"Elasticsearch URL: {es_url}") print(f"Kibana URL: {kb_url}") print(f"ES_API_KEY: ***************************{es_api_key[-5:]}")

请求头统一带上kbn-xsrf,这是 Kibana 对写操作的防护要求,缺了会返回 400:

HEADERS = { "Content-Type": "application/json", "kbn-xsrf": "true", "Authorization": f"ApiKey {es_api_key}", }

先做一次连通性验证,调/api/status确认 Kibana 可达:

import requests response = requests.get(kb_url + "/api/status", headers=HEADERS) response.raise_for_status() status = response.json() print("Cluster:", status.get("name")) print("Version:", status.get("version").get("number"))

看到版本号打印出来,说明前置链路是通的。这一步别跳过,后面所有请求都依赖这个 Headers 和地址。

3. 可复制配置:Tool 定义与 Agent 绑定

这一节是全文的核心,给出可以直接粘贴运行的 JSON 配置。先定义 Tool,再定义 Agent,最后把两者绑定。

Tool 的本质是一段带参数的 ES|QL 查询加上给 LLM 看的描述。描述字段决定了模型什么时候会调用它,所以要把「做什么、返回什么、按什么排序」写清楚。下面这个 Tool 用来找出受负面新闻影响最大的客户持仓:

{ "id": "find_client_exposure_to_negative_news", "type": "esql", "description": "Finds client portfolio exposure to negative news. Scans recent news and reports for negative sentiment, identifies the associated asset, and finds all clients holding that asset. Returns a list sorted by current market value of the position.", "configuration": { "query": "FROM financial_news, financial_reports METADATA _index | WHERE sentiment == \"negative\" | WHERE coalesce(published_date, report_date) >= NOW() - TO_TIMEDURATION(?time_duration) | RENAME primary_symbol AS symbol | LOOKUP JOIN financial_asset_details ON symbol | LOOKUP JOIN financial_holdings ON symbol | LOOKUP JOIN financial_accounts ON account_id | WHERE account_holder_name IS NOT NULL | EVAL position_current_value = quantity * current_price.price | RENAME title AS news_title | KEEP account_holder_name, symbol, asset_name, news_title, sentiment, position_current_value, quantity, current_price.price, published_date, report_date | SORT position_current_value DESC | LIMIT 50", "params": { "time_duration": { "type": "keyword", "description": "The timeframe to search back for negative news. Format is \"X hours\". Default 8760 hours." } } }, "tags": ["retrieval", "risk-analysis"] }

注意params里的time_duration是占位符,查询里用?time_duration引用。这样模型可以自己决定回溯多久,而不是写死。

创建 Tool 的请求:

tool_id = "find_client_exposure_to_negative_news" tool_definition = { ... } # 上面的 JSON response = requests.post( f"{kb_url}/api/agent_builder/tools", headers=HEADERS, json=tool_definition ) response.raise_for_status() print("Tool created:", response.json())

如果返回 400 且提示Tool with id ... already exists,说明之前建过,直接继续即可,不用删。

接下来定义 Agent。Agent 的configuration.instructions是它的行为准则,tools数组里放它可用的工具 ID。除了自定义 Tool,还可以挂上平台内置的检索类工具:

{ "id": "financial_assistant", "name": "Financial Assistant", "description": "An assistant for analyzing and understanding your financial data", "labels": ["Finance"], "avatar_color": "#16C5C0", "avatar_symbol": "FA", "configuration": { "instructions": "You are a specialized Data Intelligence Assistant for financial managers. Respond accurately and concisely. Use available tools to explore indices before answering. DO NOT provide financial advice or predictions. Format all responses in Markdown.", "tools": [ { "tool_ids": [ "platform.core.search", "platform.core.list_indices", "platform.core.get_index_mapping", "platform.core.get_document_by_id", "find_client_exposure_to_negative_news" ] } ] } }

创建 Agent:

agent_id = "financial_assistant" agent_definition = { ... } # 上面的 JSON response = requests.post( f"{kb_url}/api/agent_builder/agents", headers=HEADERS, json=agent_definition ) response.raise_for_status() print("Agent created:", response.json())

如果返回 409,说明 Agent 已存在,同样继续。

这里有个关键点:tool_ids里必须包含你刚创建的 Tool ID,否则 Agent 在对话时不会调用它。很多人卡在这里,以为 Agent 建好了就能用,结果模型只会用内置检索工具,查不到你的业务数据。

如果你希望把模型调用统一管理,可以在环境变量里加一个兼容 OpenAI 协议的 Base URL 和 Key,指向 TaoToken 的 API 入口https://taotoken.net/api,模型 ID 按你实际使用的填。这样 Agent 的推理请求就走这个统一入口,便于后续切换模型或做用量统计。

4. 验证请求:从 ES|QL 到对话链路是否生效

配置写完,必须验证。验证分两步:先确认原始 ES|QL 查询能返回数据,再确认 Agent 对话能触发 Tool 调用。

第一步,直接打/_query端点跑原始查询,确认语法和数据都对:

esql_query = """FROM financial_news, financial_reports METADATA _index | WHERE sentiment == "negative" | WHERE coalesce(published_date, report_date) >= NOW() - TO_TIMEDURATION(?time_duration) | RENAME primary_symbol AS symbol | LOOKUP JOIN financial_asset_details ON symbol | LOOKUP JOIN financial_holdings ON symbol | LOOKUP JOIN financial_accounts ON account_id | WHERE account_holder_name IS NOT NULL | EVAL position_current_value = quantity * current_price.price | KEEP account_holder_name, symbol, asset_name, news_title, sentiment, position_current_value | SORT position_current_value DESC | LIMIT 50""" request_body = {"query": esql_query, "params": [{"time_duration": "2000 hours"}]} response = requests.post(f"{es_url}/_query", headers=HEADERS, json=request_body, verify=False) response.raise_for_status() result = response.json() columns = [col["name"] for col in result["columns"]] rows = result.get("rows", result.get("values", [])) print("Rows:", len(rows))

能打印出行数,说明查询本身没问题。如果这里就报错,先修查询,别往下走。

第二步,向 Agent 发对话请求,看它是否调用了你的 Tool:

user_question = "I'm worried about market sentiment. Can you show me which of our clients are most at risk from bad news?" chat_request_body = { "input": user_question, "agent_id": agent_id } response = requests.post( f"{kb_url}/api/agent_builder/converse", headers=HEADERS, json=chat_request_body ) response.raise_for_status() chat_response = response.json()

响应里有三个关键字段:conversation_id用于继续同一线程,response.message是最终答案,steps是推理过程。检查steps里有没有tool_call类型且tool_id等于你的自定义 Tool:

for i, step in enumerate(chat_response.get("steps", [])): if step.get("type") == "tool_call": print(f"Step {i+1}: {step.get('tool_id')}") print("Params:", step.get("params"))

如果看到find_client_exposure_to_negative_news被调用,且参数里带了time_duration,说明链路完全打通。最终答案里应该包含按持仓市值排序的客户列表。

第三步,把 Agent 通过 MCP 暴露给外部客户端。Kibana 的 MCP 地址是http://localhost:5601/api/agent_builder/mcp。以 Claude Desktop 为例,配置如下:

{ "mcpServers": { "elastic": { "command": "npx", "args": [ "mcp-remote", "http://localhost:5601/api/agent_builder/mcp", "--header", "Authorization:${AUTH_HEADER}" ], "env": { "AUTH_HEADER": "ApiKey 你的Base64编码APIKey" } } } }

这里三件套要齐全:Base URL 是http://localhost:5601/api/agent_builder/mcp,Key 是ApiKey开头的编码串,Model ID 在客户端侧选择。重启客户端后,工具列表里应该能看到你的自定义 Tool,输入同样的问题,返回结果和 Notebook 里一致。

5. 常见报错排查:401、local proxy failed 与 choices 解析

链路跑不通时,报错信息往往指向具体环节。下面是我踩过的几个坑,对照排查能省不少时间。

401 Unauthorized:最常见。检查Authorization头是不是ApiKey开头(注意有个空格),Key 本身是不是 Base64 编码的完整串。如果 Key 里包含特殊字符,确认没有在复制时被截断。另外 Kibana 和 Elasticsearch 的 Key 权限可能不同,写操作需要manage_agent_builder之类的权限。

local proxy failed / connection refused:MCP 客户端连不上 Kibana。先确认kb_url是http://localhost:5601而不是https,本地默认是 HTTP。如果 Kibana 跑在容器里,localhost在客户端侧可能指向不同网络,需要换成实际可达的地址。mcp-remote依赖 Node 环境,确认npx可用。

reading 'choices' of undefined:这个报错通常出现在模型调用返回体不符合预期时。如果你用了兼容 OpenAI 协议的入口,检查 Base URL 是否指向https://taotoken.net/api,路径有没有多写或少写/v1。响应体里没有choices字段,说明请求根本没到模型服务,或者返回的是错误页。打印完整响应体定位:

try: response.raise_for_status() except requests.exceptions.RequestException as e: print("Status:", e.response.status_code) print("Body:", e.response.text)

OAuth / token 过期:如果用了带鉴权的模型入口,Key 过期会返回 401 或 403。重新生成 Key 并更新环境变量,重启 Notebook 内核让新值生效。

Tool 未被调用:Agent 建好了但对话时不调你的 Tool。检查tool_ids数组里是否包含自定义 Tool ID,以及 Tool 的description是否足够清晰。描述太模糊,模型不知道什么时候该用。可以临时把instructions里加一句「涉及客户持仓风险时,必须调用 find_client_exposure_to_negative_news」。

查询返回空:time_duration设得太短,或者sentiment字段值不匹配。先用宽松条件跑一次,确认有数据再收紧。

排查顺序建议:先验 Kibana 连通性,再验 ES|QL 原始查询,再验 Tool 创建,最后验 Agent 对话。哪一步断,问题就在哪。

6. 把链路接到你的业务数据上

跑通示例之后,替换成自己的索引和字段是下一步。核心改动点有三个:ES|QL 查询里的FROM和字段名、Tool 的description、Agent 的instructions。查询改完先用/_query单独验证,确认返回结构符合预期,再更新 Tool 定义。

更新已有 Tool 用 PUT,路径带上 Tool ID:

response = requests.put( f"{kb_url}/api/agent_builder/tools/{tool_id}", headers=HEADERS, json=tool_definition )

Agent 的instructions建议写清楚三件事:角色定位、可用数据范围、输出格式。比如「你是运维助手,只能查询 logs-* 索引,回答用 Markdown 表格」。指令越具体,模型越不容易跑偏。

如果你打算长期在编码或 Agent 场景里用这套链路,可以把模型调用统一到 Coding Plan 这类入口,配合 API Keys 管理多把 Key,接入文档里有完整的参数说明。验证模型是否正常响应时,用模型对话页面快速测一条请求,比在 Notebook 里反复调试快得多。

最后留一个实用技巧:把conversation_id存下来,后续追问时带上它,Agent 能记住上下文,不用每次重复背景。多轮对话的体验会好很多。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询