原论文:SKILL.state: Scalable Long-Horizon Agent Skills,
Google LLC + Purdue University,2026-08-26 首版,2026-09-02 v3。论文标注 accepted at EMNLP。
一、整体思想:多轮对话变成状态提交
SKILL.state 改的是多轮对话的默认模式。
传统 Agent 下一轮通常会继续看到这些东西:
用户最初的需求 上一轮、上上轮、更早的用户补充 模型之前的推理 所有工具调用 所有工具输出 模型之前说过的话SKILL.state 把这些历史从下一轮 prompt 里拿掉。下一轮模型只看三块:
P Policy / Skill:稳定规则。告诉模型这类任务怎么做。 Sigma State:当前状态。告诉模型现在已经确定了什么、做到哪一步。 O Observation:最新观察。告诉模型刚刚发生了什么。写成论文里的形式就是:
下一步输入 = (P, Sigma_t, O_t) 下一步输出 = reasoning + state_patch + action这里的P通常来自SKILL.md,比如“修改前先读现有实现,保持 API 向后兼容,完成后运行测试”。Sigma_t通常来自state.json,比如“正在给订单列表增加 CSV 导出;仅管理员可用;已定位到订单路由和权限中间件”。O_t是最新一条工具结果或用户补充,比如“搜索结果显示订单接口在routes/orders.ts”。
之前的用户输入、工具输出、模型推理不会直接出现在下一轮。它们要想继续生效,必须被写进Sigma_t。
用户最初的任务可以写进state.task;用户中途追加的约束先作为本轮O_t出现,下一轮还要用,就提交进state_patch。
所以这套思想一句话就能说清:把多轮聊天记录改成一份当前状态;模型每轮只读当前状态和最新观察。
二、用一个开发需求看多轮怎么跑
假设用户提出需求:
给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。
这句话是第 0 轮的最新观察O_0。处理开发任务的SKILL.md是稳定的P:
# Implement Feature - 修改前先搜索并阅读现有实现。 - 保持现有 API 向后兼容。 - 权限校验复用项目已有机制。 - 修改后运行相关测试。任务刚开始时,Sigma_0只表示还没有完成探索:
{"task":"","facts":[],"progress":"not_started","open_questions":[],"done":false}第 1 轮:接收需求
模型收到的 prompt 是:
{"P":"修改前先搜索并阅读现有实现;保持现有 API 向后兼容;权限校验复用已有机制;修改后运行测试。","Sigma":{"task":"","facts":[],"progress":"not_started","open_questions":[],"done":false},"O":{"type":"user_message","content":"给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。"}}模型把需求中下一轮仍要使用的信息提交进 state,同时选择搜索动作:
{"reasoning":"先定位订单接口、权限实现和现有测试。","state_patch":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation","open_questions":["where is the order list implemented?","how does the project enforce admin access?"],"done":false},"action":{"tool":"search_code","args":{"query":"orders route admin middleware"}}}runtime 提交 patch,再执行搜索。搜索结果成为下一轮的O_1:
routes/orders.ts: GET /orders middleware/require-admin.ts: requireAdmin() tests/orders.test.ts: order list tests第 2 轮:根据搜索结果继续
这一轮不携带用户原话和第 1 轮 reasoning。模型看到同一份P、已提交的Sigma_1,以及刚得到的O_1:
{"P":"修改前先搜索并阅读现有实现;保持现有 API 向后兼容;权限校验复用已有机制;修改后运行测试。","Sigma":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation","open_questions":["where is the order list implemented?","how does the project enforce admin access?"],"done":false},"O":{"type":"tool_result","content":"Order route: routes/orders.ts; admin middleware: middleware/require-admin.ts; tests: tests/orders.test.ts"}}模型把搜索结论写回 state,并读取相关文件:
{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts"],"progress":"reading implementation","open_questions":["what response shape and query parameters already exist?"],"done":false},"action":{"tool":"read_files","args":{"paths":["routes/orders.ts","middleware/require-admin.ts","tests/orders.test.ts"]}}}用户中途补充约束
读取结果也会在下一轮提炼进 state,例如“现有GET /orders返回 JSON,应新增独立导出路由”。随后用户又说:
CSV 里的时间统一用 UTC。
这条消息成为新的O。模型仍能从当前Sigma读到原需求、权限要求和代码位置,于是把 UTC 约束并入已有facts:
{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}几轮之后,完整聊天、搜索命令和文件内容都留在审计日志里。下一轮只读取当前state.json和最新编辑结果。
用户原始需求已经离开 prompt,其中的长期约束仍保存在 state 中。
三、runtime 和 LLM 自己 patch 的区别
普通 OpenAI tool calling 的流程是:
LLM 返回 tool_call → runtime 执行工具 → 工具结果作为下一条 tool message 发回 LLMSKILL.state 只改第一步的输出内容:
LLM 同时返回 state_patch + action → runtime 校验并提交 state_patch → runtime 执行 action → 工具结果成为下一轮 O论文附录中的action是字符串命令,没有使用 OpenAI 原生tool_calls。接入 OpenAI 时不必修改 API 协议,
可以定义一个propose_transition工具,让一次 function call 同时携带两部分:
{"state_patch":{"progress":"ready to implement"},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts"]}}}不要拆成update_state和edit_files两个 tool call:两者可能并行或只成功一个。runtime 应先校验整个 transition,
再提交 state 和执行工具。需求怎样实现由模型判断;这次状态更新能否生效,由 runtime 判断。
四、少上下文是结果,单源事实才是因
完整历史的输入会随步数增长。第 100 步要读前 99 步,第 101 步又要读前 100 步,累计成本接近O(T^2)。
SKILL.state 每一步只读P + Sigma + O;只要 Skill、state、最新 observation 的大小受控,累计成本就是O(T)。
这个成本变化来自一条规则:
下一步要依赖的事实,必须先通过
state_patch提交进state.json。
开发例子里,“仅管理员可用”来自用户第一条消息。后续实现和测试都要用它,所以模型必须把它写进facts。
用户后来补充的 UTC 约束也要提交:
{"state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}到了实现阶段,模型看到的是“目标需求 + 权限和兼容性约束 + 已定位文件 + 最新工具结果”。
任务进度、关键事实、未解问题都有一个固定入口。
五、审计日志和 state 分工不同
生产系统需要两份记录:
state.json 给下一步决策用,只保留当前充分状态 audit/events.jsonl 给人、测试和故障复盘用,保留完整提交记录state.json要短、稳定、可放进 prompt。它回答“下一步现在知道什么”。
{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement","open_questions":[],"next_action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}},"done":false}audit/events.jsonl可以长。它回答“这份 state 是怎么来的”:
{"type":"step","observation":"给订单列表增加 CSV 导出。仅管理员可用,不要改变现有 JSON 接口。","state_patch":{"task":"add CSV export to order list","facts":["CSV export is admin-only","existing JSON API must remain compatible"],"progress":"locating implementation"},"action":{"tool":"search_code","args":{"query":"orders route admin middleware"}}} {"type":"step","observation":"Order route: routes/orders.ts; admin middleware: middleware/require-admin.ts","state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts"],"progress":"reading implementation"},"action":{"tool":"read_files","args":{"paths":["routes/orders.ts","middleware/require-admin.ts","tests/orders.test.ts"]}}} {"type":"step","observation":"CSV 里的时间统一用 UTC。","state_patch":{"facts":["CSV export is admin-only","existing JSON API must remain compatible","CSV timestamps must use UTC","order list is implemented in routes/orders.ts","admin access uses requireAdmin()","order tests are in tests/orders.test.ts","GET /orders returns JSON and must remain unchanged","CSV export should use a separate route"],"progress":"ready to implement"},"action":{"tool":"edit_files","args":{"paths":["routes/orders.ts","tests/orders.test.ts"]}}}下一轮 prompt 使用state.json和最新 observation。审计日志留给ss check、故障复盘和人读。
这条边界让 operational state 保持小,也让旧轨迹仍然可查。
六、Token 收益要降温,执行语义更重要
论文报告的几个结果很强:
| 场景 | 历史基线 | SKILL.state |
|---|---|---|
| 仓储 T=100 | 0.91 / 1,062,387 tokens | 0.94 / 65,408 |
| 仓储 T=200 | 0.88 Stateful | 0.94 |
| 软件仓库 T=100 | 0.63 / 2,308,000 | 0.78 / 90,200 |
| InterCode CTF | 46.4% | 54.2% |
| tau-Bench Retail | 51.7% | 58.3% |
| tau-Bench Airline | 28.1% | 32.4% |
仓储 T=100 的16.2xToken 降幅最醒目,这个数字不能直接换算成生产 Agent 的成本或速度收益。实际 coding agent 常会裁剪大型 tool output、压缩旧轮次,而稳定的历史前缀还可能命中服务端 prompt cache。论文统计的是累计 Token,没有报告缓存命中率、缓存后的实际费用和wall-clock latency。因此,16.2x更接近“相对完整重放历史的理论收益上限”。
如果现有 Agent 已经丢弃旧工具输出,SKILL.state 能继续节省的 Token 确实会少很多。更关键的比较是论文的同预算实验:
| 仓储 T=100,约 1,800 字符/轮 | 得分 |
|---|---|
| 滑动窗口 | 0.18 |
| LLMLingua | 0.22 |
| 受限摘要 | 0.52 |
| SKILL.state | 0.94 |
几种方法的累计 Token 都在约 6.2 万到 6.5 万之间,差异主要体现在正确率。这个实验支持的结论是:
关键不只是删掉历史,而是把仍然影响后续动作的事实写进有领域语义的状态。
除了 Token,显式状态还有三个价值:
- 旧事实不会与新事实同时留在 prompt。当前状态覆盖或删除旧值,模型不必从多轮记录中判断哪个版本有效。
- 环境变化可以立即成为当前事实。论文的状态漂移实验中,SKILL.state 恢复步数为 0;历史方案需要多个步骤摆脱旧信息。
- 状态提交可以校验和回滚。schema、类型检查和原子写入能阻止非法 patch 成为下一轮事实;状态文件也自然形成暂停和恢复的检查点。
前两项有论文实验支撑;第三项主要是 runtime 设计带来的工程收益。
论文结论仍有几个边界:
- 论文没有控制 prompt cache,也没有与具体生产 Agent 的 output masking、自动 compaction 做直接对照。
Stateful (LangGraph-style)是作者构造的state + full historybaseline;LangGraph 本身允许调用方自行决定 prompt 内容。- 开源模型收益弱很多。错误分析中,68% 来自提前覆盖或删除状态,20% 是 schema/type 错误,12% 是 JSON 格式错误。
所以,这项工作的意义不宜概括成“找到一种更省 Token 的压缩方法”。它真正提出的是一种执行契约:
Agent 的下一步只依赖经过提交的当前事实,不再依赖模型从聊天记录中重建当前事实。
这对状态结构稳定的长流程很有价值,例如订单、审批、发布和运维工作流;对探索式编码、开放研究以及需要保留用户语境的任务,
固定 schema 容易漏掉当时没有意识到的重要信息。实际 coding agent 更适合采用有界 operational state、独立审计日志和按需历史检索的组合。
七、和 Ledger 的分歧更值得看
同月还有一篇更工程化的相邻工作:Ledger: Turning Interaction History into Execution State。
Ledger 不让模型维护状态。runtime 根据已经完成的工具调用,确定性地记录三类事实:
观察记录:成功读取了哪个文件、哪些行,以及读取时的版本计数 修改状态:哪些文件被修改,以及文件级和仓库级的变更计数 命令记录:标准化后的命令、命令类别和最近执行位置例如 Agent 已经读取src/order.ts:1-120,期间文件没有变化,又提出相同读取。Ledger 在命令真正执行前检查:
相同行范围以前成功读取过? → 是 读取后文件或仓库发生过修改? → 没有 旧结果仍在当前模型上下文中? → 是此时 runtime 返回Reuse:跳过读取,引用旧结果。文件已经修改时返回Allow,重新读取。测试命令更保守:
没有修改代码却重复运行测试时返回Nudge,测试仍会执行,只在结果后附加“可能重复”的提示。Ledger 只强制复用确定安全的读取和搜索结果。
因此,所谓“拦截”不是让另一个模型判断两个动作是否相似,而是根据标准化命令、读取范围和变更计数做规则匹配。
解析失败、不支持的命令以及环境配置等操作默认Allow。
理解这个机制后,两者的分歧就很直接:
SKILL.state 下一次输入 = Skill + 当前 state + 最新 observation 历史消息不再发送 Ledger 下一次输入 = 原有对话历史 + 确定性执行账本 历史消息继续发送Ledger 容易接进现有 coding agent,也能在执行前阻止一部分重复工作,但无法获得 SKILL.state 的固定 prompt 大小。
SKILL.state 能彻底停止历史增长,代价是必须重新组织 agent loop,并确保 schema 足以承载后续需要的事实。
本复现选择 SKILL.state 的执行方式,同时保留一份不进入 prompt 的审计日志,供测试和故障复盘使用。
八、这套东西该怎么迁移到真实 Agent
最小可用版本应该长这样:
runtime/ prompt_builder 只拼 P + Sigma + O patch_validator schema 校验,返回 Path + Hint state_store 原子提交,失败回滚 audit_log append-only,不进 prompt retry_policy patch 无效时重试,不执行 action tests/ no_history_leak invalid_patch_rollback resume_round_trip bounded_prompt_growth最先写的测试应该是no_history_leak。因为这条一破,整个系统表面上仍能跑,实际上已经退回完整历史 Agent。
其次是invalid_patch_rollback。模型迟早会写错 JSON、写错字段、删掉还需要的事实。
runtime 要按 schema 裁决每一次提交。
最后才是成本曲线。prompt 变短只是结果;如果状态契约不对,短 prompt 只会更快地把任务跑错。