☰
SKILL.state:用显式执行状态替代对话历史
2026/10/10 9:09:43 网站建设 项目流程

原论文:SKILL.state: Scalable Long-Horizon Agent Skills,
Google LLC + Purdue University,2026-08-26 首版,2026-09-02 v3。论文标注 accepted at EMNLP。

一、整体思想:多轮对话变成状态提交

SKILL.state 改的是多轮对话的默认模式。

传统 Agent 下一轮通常会继续看到这些东西:

用户最初的需求 上一轮、上上轮、更早的用户补充 模型之前的推理 所有工具调用 所有工具输出 模型之前说过的话

SKILL.state 把这些历史从下一轮 prompt 里拿掉。下一轮模型只看三块:

P Policy / Skill:稳定规则。告诉模型这类任务怎么做。 Sigma State:当前状态。告诉模型现在已经确定了什么、做到哪一步。 O Observation:最新观察。告诉模型刚刚发生了什么。

写成论文里的形式就是:

下一步输入 = (P, Sigma_t, O_t) 下一步输出 = reasoning + state_patch + action

这里的P通常来自SKILL.md,比如“修改前先读现有实现,保持 API 向后兼容,完成后运行测试”。
Sigma_t通常来自state.json,比如“正在给订单列表增加 CSV 导出;仅管理员可用;已定位到订单路由和权限中间件”。
O_t是最新一条工具结果或用户补充,比如“搜索结果显示订单接口在routes/orders.ts”。

之前的用户输入、工具输出、模型推理不会直接出现在下一轮。它们要想继续生效,必须被写进Sigma_t。
用户最初的任务可以写进state.task;用户中途追加的约束先作为本轮O_t出现,下一轮还要用,就提交进state_patch。

所以这套思想一句话就能说清:把多轮聊天记录改成一份当前状态;模型每轮只读当前状态和最新观察。

二、用一个开发需求看多轮怎么跑

假设用户提出需求:

给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。

这句话是第 0 轮的最新观察O_0。处理开发任务的SKILL.md是稳定的P:

# Implement Feature - 修改前先搜索并阅读现有实现。 - 保持现有 API 向后兼容。 - 权限校验复用项目已有机制。 - 修改后运行相关测试。

任务刚开始时,Sigma_0只表示还没有完成探索:

{"task":"","facts":[],"progress":"not_started","open_questions":[],"done":false}

第 1 轮:接收需求

模型收到的 prompt 是:

{"P":"修改前先搜索并阅读现有实现;保持现有 API 向后兼容;权限校验复用已有机制;修改后运行测试。","Sigma":{"task":"","facts":[],"progress":"not_started","open_questions":[],"done":false},"O":{"type":"user_message","content":"给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。"}}

模型把需求中下一轮仍要使用的信息提交进 state,同时选择搜索动作:

{"reasoning":"先定位订单接口、权限实现和现有测试。","state_patch":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation","open_questions":["where is the order list implemented?","how does the project enforce admin access?"],"done":false},"action":{"tool":"search_code","args":{"query":"orders route admin middleware"}}}

runtime 提交 patch,再执行搜索。搜索结果成为下一轮的O_1:

routes/orders.ts: GET /orders middleware/require-admin.ts: requireAdmin() tests/orders.test.ts: order list tests

第 2 轮:根据搜索结果继续

这一轮不携带用户原话和第 1 轮 reasoning。模型看到同一份P、已提交的Sigma_1,以及刚得到的O_1:

{"P":"修改前先搜索并阅读现有实现;保持现有 API 向后兼容;权限校验复用已有机制;修改后运行测试。","Sigma":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation","open_questions":["where is the order list implemented?","how does the project enforce admin access?"],"done":false},"O":{"type":"tool_result","content":"Order route: routes/orders.ts; admin middleware: middleware/require-admin.ts; tests: tests/orders.test.ts"}}

模型把搜索结论写回 state,并读取相关文件:

{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts"],"progress":"reading implementation","open_questions":["what response shape and query parameters already exist?"],"done":false},"action":{"tool":"read_files","args":{"paths":["routes/orders.ts","middleware/require-admin.ts","tests/orders.test.ts"]}}}

用户中途补充约束

读取结果也会在下一轮提炼进 state,例如“现有GET /orders返回 JSON,应新增独立导出路由”。随后用户又说:

CSV 里的时间统一用 UTC。

这条消息成为新的O。模型仍能从当前Sigma读到原需求、权限要求和代码位置,于是把 UTC 约束并入已有facts:

{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}

几轮之后,完整聊天、搜索命令和文件内容都留在审计日志里。下一轮只读取当前state.json和最新编辑结果。
用户原始需求已经离开 prompt,其中的长期约束仍保存在 state 中。

三、runtime 和 LLM 自己 patch 的区别

普通 OpenAI tool calling 的流程是:

LLM 返回 tool_call → runtime 执行工具 → 工具结果作为下一条 tool message 发回 LLM

SKILL.state 只改第一步的输出内容:

LLM 同时返回 state_patch + action → runtime 校验并提交 state_patch → runtime 执行 action → 工具结果成为下一轮 O

论文附录中的action是字符串命令,没有使用 OpenAI 原生tool_calls。接入 OpenAI 时不必修改 API 协议,
可以定义一个propose_transition工具,让一次 function call 同时携带两部分:

{"state_patch":{"progress":"ready to implement"},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts"]}}}

不要拆成update_state和edit_files两个 tool call:两者可能并行或只成功一个。runtime 应先校验整个 transition,
再提交 state 和执行工具。需求怎样实现由模型判断;这次状态更新能否生效,由 runtime 判断。

四、少上下文是结果,单源事实才是因

完整历史的输入会随步数增长。第 100 步要读前 99 步,第 101 步又要读前 100 步,累计成本接近O(T^2)。
SKILL.state 每一步只读P + Sigma + O;只要 Skill、state、最新 observation 的大小受控,累计成本就是O(T)。

这个成本变化来自一条规则:

下一步要依赖的事实,必须先通过state_patch提交进state.json。

开发例子里,“仅管理员可用”来自用户第一条消息。后续实现和测试都要用它,所以模型必须把它写进facts。
用户后来补充的 UTC 约束也要提交:

{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}

到了实现阶段,模型看到的是“目标需求 + 权限和兼容性约束 + 已定位文件 + 最新工具结果”。
任务进度、关键事实、未解问题都有一个固定入口。

五、审计日志和 state 分工不同

生产系统需要两份记录:

state.json 给下一步决策用,只保留当前充分状态 audit/events.jsonl 给人、测试和故障复盘用,保留完整提交记录

state.json要短、稳定、可放进 prompt。它回答“下一步现在知道什么”。

{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false}

audit/events.jsonl可以长。它回答“这份 state 是怎么来的”:

{"type":"step","observation":"给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。","state_patch":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation"},"action":{"tool":"search_code","args":{"query":"orders route admin middleware"}}} {"type":"step","observation":"Order route: routes/orders.ts; admin middleware: middleware/require-admin.ts","state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts"],"progress":"reading implementation"},"action":{"tool":"read_files","args":{"paths":["routes/orders.ts","middleware/require-admin.ts","tests/orders.test.ts"]}}} {"type":"step","observation":"CSV 里的时间统一用 UTC。","state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement"},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}

下一轮 prompt 使用state.json和最新 observation。审计日志留给ss check、故障复盘和人读。
这条边界让 operational state 保持小,也让旧轨迹仍然可查。

六、Token 收益要降温,执行语义更重要

论文报告的几个结果很强:

场景历史基线SKILL.state
仓储 T=1000.91 / 1,062,387 tokens0.94 / 65,408
仓储 T=2000.88 Stateful0.94
软件仓库 T=1000.63 / 2,308,0000.78 / 90,200
InterCode CTF46.4%54.2%
tau-Bench Retail51.7%58.3%
tau-Bench Airline28.1%32.4%

仓储 T=100 的16.2xToken 降幅最醒目,这个数字不能直接换算成生产 Agent 的成本或速度收益。实际 coding agent 常会裁剪大型 tool output、压缩旧轮次,而稳定的历史前缀还可能命中服务端 prompt cache。论文统计的是累计 Token,没有报告缓存命中率、缓存后的实际费用和wall-clock latency。因此,16.2x更接近“相对完整重放历史的理论收益上限”。

如果现有 Agent 已经丢弃旧工具输出,SKILL.state 能继续节省的 Token 确实会少很多。更关键的比较是论文的同预算实验:

仓储 T=100,约 1,800 字符/轮得分
滑动窗口0.18
LLMLingua0.22
受限摘要0.52
SKILL.state0.94

几种方法的累计 Token 都在约 6.2 万到 6.5 万之间,差异主要体现在正确率。这个实验支持的结论是:
关键不只是删掉历史,而是把仍然影响后续动作的事实写进有领域语义的状态。

除了 Token,显式状态还有三个价值:

  1. 旧事实不会与新事实同时留在 prompt。当前状态覆盖或删除旧值,模型不必从多轮记录中判断哪个版本有效。
  2. 环境变化可以立即成为当前事实。论文的状态漂移实验中,SKILL.state 恢复步数为 0;历史方案需要多个步骤摆脱旧信息。
  3. 状态提交可以校验和回滚。schema、类型检查和原子写入能阻止非法 patch 成为下一轮事实;状态文件也自然形成暂停和恢复的检查点。

前两项有论文实验支撑;第三项主要是 runtime 设计带来的工程收益。

论文结论仍有几个边界:

  1. 论文没有控制 prompt cache,也没有与具体生产 Agent 的 output masking、自动 compaction 做直接对照。
  2. Stateful (LangGraph-style)是作者构造的state + full historybaseline;LangGraph 本身允许调用方自行决定 prompt 内容。
  3. 开源模型收益弱很多。错误分析中,68% 来自提前覆盖或删除状态,20% 是 schema/type 错误,12% 是 JSON 格式错误。

所以,这项工作的意义不宜概括成“找到一种更省 Token 的压缩方法”。它真正提出的是一种执行契约:
Agent 的下一步只依赖经过提交的当前事实,不再依赖模型从聊天记录中重建当前事实。

这对状态结构稳定的长流程很有价值,例如订单、审批、发布和运维工作流;对探索式编码、开放研究以及需要保留用户语境的任务,
固定 schema 容易漏掉当时没有意识到的重要信息。实际 coding agent 更适合采用有界 operational state、独立审计日志和按需历史检索的组合。

七、和 Ledger 的分歧更值得看

同月还有一篇更工程化的相邻工作:Ledger: Turning Interaction History into Execution State。

Ledger 不让模型维护状态。runtime 根据已经完成的工具调用,确定性地记录三类事实:

观察记录:成功读取了哪个文件、哪些行,以及读取时的版本计数 修改状态:哪些文件被修改,以及文件级和仓库级的变更计数 命令记录:标准化后的命令、命令类别和最近执行位置

例如 Agent 已经读取src/order.ts:1-120,期间文件没有变化,又提出相同读取。Ledger 在命令真正执行前检查:

相同行范围以前成功读取过? → 是 读取后文件或仓库发生过修改? → 没有 旧结果仍在当前模型上下文中? → 是

此时 runtime 返回Reuse:跳过读取,引用旧结果。文件已经修改时返回Allow,重新读取。测试命令更保守:
没有修改代码却重复运行测试时返回Nudge,测试仍会执行,只在结果后附加“可能重复”的提示。Ledger 只强制复用确定安全的读取和搜索结果。

因此,所谓“拦截”不是让另一个模型判断两个动作是否相似,而是根据标准化命令、读取范围和变更计数做规则匹配。
解析失败、不支持的命令以及环境配置等操作默认Allow。

理解这个机制后,两者的分歧就很直接:

SKILL.state 下一次输入 = Skill + 当前 state + 最新 observation 历史消息不再发送 Ledger 下一次输入 = 原有对话历史 + 确定性执行账本 历史消息继续发送

Ledger 容易接进现有 coding agent,也能在执行前阻止一部分重复工作,但无法获得 SKILL.state 的固定 prompt 大小。

SKILL.state 能彻底停止历史增长,代价是必须重新组织 agent loop,并确保 schema 足以承载后续需要的事实。
本复现选择 SKILL.state 的执行方式,同时保留一份不进入 prompt 的审计日志,供测试和故障复盘使用。

八、这套东西该怎么迁移到真实 Agent

最小可用版本应该长这样:

runtime/ prompt_builder 只拼 P + Sigma + O patch_validator schema 校验,返回 Path + Hint state_store 原子提交,失败回滚 audit_log append-only,不进 prompt retry_policy patch 无效时重试,不执行 action tests/ no_history_leak invalid_patch_rollback resume_round_trip bounded_prompt_growth

最先写的测试应该是no_history_leak。因为这条一破,整个系统表面上仍能跑,实际上已经退回完整历史 Agent。

其次是invalid_patch_rollback。模型迟早会写错 JSON、写错字段、删掉还需要的事实。
runtime 要按 schema 裁决每一次提交。

最后才是成本曲线。prompt 变短只是结果;如果状态契约不对,短 prompt 只会更快地把任务跑错。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询