☰
DeepSeek-Agent-Harness-2026终极指南-第8章第38节-AgentLoop从零实现-流式Agent:边生成边执行的体验升级
2026/10/2 16:56:50 网站建设 项目流程

DeepSeek Agent Harness 2026终极指南 - 第8章第38节 流式Agent:边生成边执行的体验升级

第35-37节的Agent Loop是同步的——每次调模型都要等模型生成完才能继续。用户体验很差,尤其是长回答要等很久。这节做流式Agent:用stream=True逐token渲染回答,同时处理tool_calls的分片重组难题(一个工具调用可能拆成多个chunk)。从此Agent回答像ChatGPT一样实时流出,不再干等。

本文导航

  • 同步 vs 流式:用户体验的天壤之别
  • 流式模式下的chunk结构
  • tool_calls分片重组:最难的坑
  • 流式Agent Loop实现
  • 踩坑实录:那些流式模式里的坑
  • 完整实现:streaming_agent.py
  • 实测:流式回答+工具调用
  • 小结

同步 vs 流式:用户体验的天壤之别

先看同步模式的问题。假设用户问"写一首关于春天的诗",模型要生成200个token:

同步模式:

  1. 发送请求
  2. 等待10秒(模型生成200个token)
  3. 一次性返回200个token
  4. 用户看到完整回答

这10秒里用户盯着空白屏幕,不知道模型在干嘛,以为卡死了。

流式模式:

  1. 发送请求
  2. 0.5秒后开始收到第一个chunk:“春”
  3. 0.6秒后收到第二个chunk:“风”
  4. 0.7秒后收到第三个chunk:“送”
  5. …
  6. 10秒后收到最后一个chunk:“。”
  7. 用户看到文字实时流出

流式模式下用户0.5秒就能看到第一个字,体验完全不同。

流式模式下的chunk结构

用stream=True调模型时,返回的不是一个完整的ChatCompletion对象,而是一个迭代器,每次yield一个ChatCompletionChunk:

stream=client.chat.completions.create(model="deepseek-flash",messages=[{"role":"user","content":"你好"}],stream=True,)forchunkinstream:print(chunk)

每个chunk的结构:

{"id":"chatcmpl-xxx","object":"chat.completion.chunk","created":1234567890,"model":"deepseek-flash","choices":[{"index":0,"delta":{"role":"assistant",// 只在第一个chunk出现"content":"你"// 本次chunk的文本片段},"finish_reason":null// 只在最后一个chunk有值}]}

关键点:

  1. delta.content是本次chunk的文本片段,不是累积的。需要自己拼接。
  2. delta.role只在第一个chunk出现,后续chunk没有。
  3. finish_reason只在最后一个chunk有值(如"stop"),其他chunk是null。

对于纯文本回答,拼接很简单:

content=""forchunkinstream:ifchunk.choices[0].delta.content:content+=chunk.choices[0].delta.contentprint(chunk.choices[0].delta.content,end="",flush=True)

但如果有tool_calls,事情就复杂了。

tool_calls分片重组:最难的坑

当模型决定调用工具时,delta里会有tool_calls字段。但tool_calls的分片比content复杂得多:

// 第1个chunk{"delta":{"tool_calls":[{"index":0,"id":"call_abc123","function":{"name":"get_weather","arguments":""},"type":"function"}]}}// 第2个chunk{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\":"}}]}}// 第3个chunk{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"北京\"}"}}]}}

关键设计:

  1. index字段:标识这是第几个tool_call。如果模型同时调两个工具,会有index=0和index=1两个序列。
  2. id和name只在第一个chunk出现:后续chunk只有arguments片段。
  3. arguments是字符串片段:需要按index分组拼接。

重组算法:

tool_calls_map={}# {index: {"id": ..., "name": ..., "arguments": ""}}forchunkinstream:ifchunk.choices[0].delta.tool_calls:fortc_deltainchunk.choices[0].delta.tool_calls:idx=tc_delta.indexifidxnotintool_calls_map:# 第一个chunk:初始化tool_calls_map[idx]={"id":tc_delta.id,"name":tc_delta.function.name,"arguments":"","type":"function",}# 拼接argumentsiftc_delta.function.arguments:tool_calls_map[idx]["arguments"]+=tc_delta.function.arguments# 转成列表tool_calls=[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]

这个算法的核心是:

  1. 用index作为key,把同一个tool_call的所有chunk分到一组
  2. 第一个chunk提取id和name,后续chunk只拼接arguments
  3. 最后按index排序,转成列表

流式Agent Loop实现

把流式渲染和tool_calls重组整合到Agent Loop里:

defrun_streaming(user_query:str)->str:"""流式Agent Loop"""messages=[{"role":"system","content":"你是一个有用的助手。"},{"role":"user","content":user_query},]tools=registry.to_openai_tools()foriterationinrange(1,MAX_ITERS+1):logger.info(f"Loop 第{iteration}轮 ↻")# 流式调用stream=client.chat.completions.create(model=settings.deepseek_model,messages=messages,tools=tools,stream=True,)# 拼接content和tool_callscontent=""tool_calls_map={}finish_reason=Noneforchunkinstream:delta=chunk.choices[0].delta finish_reason=chunk.choices[0].finish_reason# 拼接contentifdelta.content:content+=delta.contentprint(delta.content,end="",flush=True)# 拼接tool_callsifdelta.tool_calls:fortc_deltaindelta.tool_calls:idx=tc_delta.indexifidxnotintool_calls_map:tool_calls_map[idx]={"id":tc_delta.id,"name":tc_delta.function.name,"arguments":"","type":"function",}iftc_delta.function.arguments:tool_calls_map[idx]["arguments"]+=tc_delta.function.argumentsprint()# 换行# 没有tool_calls → 回答完毕ifnottool_calls_map:returncontent# 有tool_calls → 执行工具tool_calls=[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]messages.append({"role":"assistant","content":content,"tool_calls":tool_calls,})fortcintool_calls:func_name=tc["name"]func_args=json.loads(tc["arguments"])logger.info(f" → 调用工具:{func_name}({json.dumps(func_args,ensure_ascii=False)})")result=executor.execute(func_name,func_args)logger.info(f" ← 工具结果:{result[:80]}{'...'iflen(result)>80else''}")messages.append({"role":"tool","tool_call_id":tc["id"],"content":result,})logger.warning(f"Agent Loop 达到最大迭代次数{MAX_ITERS}")return"[Agent] 抱歉,处理超时。"

关键改动:

  1. stream=True开启流式模式
  2. 遍历chunk,拼接content和tool_calls
  3. print(delta.content, end="", flush=True)实时输出
  4. 最后把拼接好的tool_calls转成列表,传给executor.execute()

踩坑实录:那些流式模式里的坑

流式模式有几个容易踩的坑,我全踩过了:

坑1:第一个chunk的delta.content可能是空字符串

有些模型在第一个chunk只返回{"role": "assistant"},content是空字符串。如果不判断直接拼接,会多一个空行。

# 错误写法content+=delta.content# 如果delta.content是"",会多一个空行# 正确写法ifdelta.content:content+=delta.content

坑2:finish_reason可能为null

只有最后一个chunk的finish_reason有值,其他chunk是null。如果不判断直接赋值,会把null覆盖掉之前的值。

# 错误写法finish_reason=chunk.choices[0].finish_reason# 可能被null覆盖# 正确写法ifchunk.choices[0].finish_reason:finish_reason=chunk.choices[0].finish_reason

坑3:tool_calls的index可能不连续

如果模型同时调三个工具,index可能是0, 1, 2,但也可能是0, 2, 5(虽然少见)。所以要用index作为key,不能用列表下标。

# 错误写法tool_calls_list=[]fortc_deltaindelta.tool_calls:tool_calls_list.append(...)# 如果index不连续,顺序会乱# 正确写法tool_calls_map={}fortc_deltaindelta.tool_calls:idx=tc_delta.index tool_calls_map[idx]=...

坑4:流式模式下usage信息可能缺失

有些API在流式模式下不返回usage(prompt_tokens、completion_tokens),只在最后一个chunk返回。如果需要在每次调用后统计token,流式模式可能拿不到。

解决方案:

  1. 接受这个限制,流式模式下不统计token
  2. 或者在流式结束后再调一次非流式API获取usage(浪费一次调用)
  3. 或者用tiktoken自己算(不准确,但能用)

我们选择方案1:流式模式下不统计token,只在非流式模式下统计。

完整实现:streaming_agent.py

把流式Agent Loop整合成完整模块:

# deep_pilot/streaming_agent.py —— 流式Agent Loop v0.3from__future__importannotationsimportjsonfromtypingimportAnyfromdeep_pilot.clientimportclientfromdeep_pilot.loggerimportget_loggerfromdeep_pilot.tool_registryimportregistryfromdeep_pilot.tool_executorimportget_executorfromdeep_pilot.configimportsettingsimportdeep_pilot.tools# noqa: F401logger=get_logger(__name__)executor=get_executor(registry)MAX_ITERS=10defrun_streaming(user_query:str)->str:""" 流式Agent Loop:逐token渲染回答,实时执行工具。 返回最终的文本回答。 """messages=[{"role":"system","content":"你是一个有用的助手。当需要查询实时数据时,使用提供的工具。"},{"role":"user","content":user_query},]tools=registry.to_openai_tools()foriterationinrange(1,MAX_ITERS+1):logger.info(f"Loop 第{iteration}轮 ↻")# 流式调用stream=client.chat.completions.create(model=settings.deepseek_model,messages=messages,tools=tools,stream=True,)# 拼接content和tool_callscontent=""tool_calls_map:dict[int,dict[str,Any]]={}forchunkinstream:delta=chunk.choices[0].delta# 拼接contentifdelta.content:content+=delta.contentprint(delta.content,end="",flush=True)# 拼接tool_callsifdelta.tool_calls:fortc_deltaindelta.tool_calls:idx=tc_delta.indexifidxnotintool_calls_map:tool_calls_map[idx]={"id":tc_delta.id,"name":tc_delta.function.name,"arguments":"","type":"function",}iftc_delta.function.arguments:tool_calls_map[idx]["arguments"]+=tc_delta.function.argumentsprint()# 换行# 没有tool_calls → 回答完毕ifnottool_calls_map:returncontent# 有tool_calls → 执行工具tool_calls=[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]messages.append({"role":"assistant","content":content,"tool_calls":tool_calls,})fortcintool_calls:func_name=tc["name"]func_args=json.loads(tc["arguments"])logger.info(f" → 调用工具:{func_name}({json.dumps(func_args,ensure_ascii=False)})")result=executor.execute(func_name,func_args)logger.info(f" ← 工具结果:{result[:80]}{'...'iflen(result)>80else''}")messages.append({"role":"tool","tool_call_id":tc["id"],"content":result,})logger.warning(f"Agent Loop 达到最大迭代次数{MAX_ITERS}")return"[Agent] 抱歉,处理超时。"

实测:流式回答+工具调用

uv run python-c" from deep_pilot.streaming_agent import run_streaming # 测试1:纯文本回答(流式渲染) print('=== 测试1:纯文本回答 ===') answer = run_streaming('写一首关于春天的诗') print(f'\n最终答案长度: {len(answer)} 字符') print() # 测试2:工具调用(流式+工具执行) print('=== 测试2:工具调用 ===') answer = run_streaming('北京今天天气怎么样?') print(f'\n最终答案: {answer}') "

控制台输出:

=== 测试1:纯文本回答 === 2026-09-12 19:00:01 | INFO | streaming_agent | Loop 第 1 轮 ↻ 春风轻拂柳丝长, 桃李芬芳满院香。 燕子归来寻旧巷, 莺歌婉转绕池塘。 最终答案长度: 48 字符 === 测试2:工具调用 === 2026-09-12 19:00:02 | INFO | streaming_agent | Loop 第 1 轮 ↻ 2026-09-12 19:00:02 | INFO | streaming_agent | → 调用工具: get_weather({"city": "北京"}) 2026-09-12 19:00:02 | INFO | streaming_agent | ← 工具结果: 晴,28°C,湿度 45%,北风 3 级 2026-09-12 19:00:02 | INFO | streaming_agent | Loop 第 2 轮 ↻ 北京今天天气晴朗,气温28°C,湿度45%,北风3级。适合户外活动。 最终答案: 北京今天天气晴朗,气温28°C,湿度45%,北风3级。适合户外活动。

注意测试1的输出是实时流出的,不是一次性打印。你在终端里会看到文字一个字一个字地出现,体验跟ChatGPT一样。

测试2里,模型先调工具查天气,工具执行完后模型继续生成回答,回答也是流式输出的。


小结

  1. 流式模式大幅提升用户体验:0.5秒看到第一个字,而不是等10秒看完整回答。
  2. chunk结构:delta.content是文本片段,需要拼接;delta.tool_calls是工具调用片段,需要按index分组拼接。
  3. tool_calls分片重组:用index作为key,第一个chunk提取id和name,后续chunk只拼接arguments。
  4. 踩坑实录:空字符串判断、finish_reasonnull判断、index不连续、流式模式下usage缺失。
  5. 流式Agent Loop:遍历chunk → 拼接content和tool_calls → 实时print → 执行工具 → 回填 → 循环。
  6. DeepPilot v0.3流式Agent完成——从同步等待到流式渲染,用户体验质的飞跃。

下节预告

流式Agent搞定了,但到现在我们还没写过一行测试代码。Agent Loop逻辑复杂(多轮循环、工具调用、异常处理),手动测试覆盖不全。下一节做单元测试:用MockClient替身测试Agent Loop(不发真实请求零成本),pytest fixture设计,循环逻辑边界用例(零工具/多工具/异常)。从此改代码不怕改坏。


如果觉得本文对你有帮助,欢迎点赞、收藏、关注三连!
本系列持续更新中,关注不迷路~

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询