Kairos03:面向 Physical AI 的具备后悔感知能力的原生世界-动作模型栈
2026/8/3 6:37:37 网站建设 项目流程

7 Related Work

7相关工作

7.1 Video Generation Models

7.1 视频生成模型

Diffusion-based Video Generation. The success of Diffusion Models (DMs) [1o] in image synthesis has catalyzed their extension to the video domain. Early pioneers like Video Diffusion Models (VDM) [11] first extended the standard 2D U-Net to a 3D structure [156] by replacing 2D convolutions with space-time factorized convolutions. To alleviate the heavy computational burden of 3D operators, many subsequent works [12-15, 155] adopted a “spatial-then-temporal” paradigm, inserting iD temporal attention layers after 2D spatial blocks to capture dynamic dependencies. A significant architectural shift occurred with the introduction of Diffusion Transformers (DiT) [6

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询