ByteDance Seed
首页模型博客&论文Seed Edge加入我们
首页模型博客&论文Seed Edge加入我们
体验豆包AI 体验中心

2024-11-04

How Far is Video Generation from World Model: A Physical Law Perspective

Download PDF
上一篇下一篇

摘要

OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. However, the ability of video generation models to discover such laws purely from visual data without human priors can be questioned. A world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios. In this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization. We developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws. This provides an unlimited supply of data for large-scale experimentation and enables quantitative evaluation of whether the generated videos adhere to physical laws. We trained diffusion-based video generation models to predict object movements based on initial frames. Our scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios. Further experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit "case-based" generalization behavior, i.e., mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color > size > velocity > shape. Our study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws, despite its role in Sora's broader success. See our project page at https://phyworld.github.io/

作者

Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, Jiashi Feng

期刊/会议

ICML 2025

追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
模型成果
Seed2.1
Seedance 2.5
Seedream 5.0 Pro
SeedRealtime
Seed Audio 1.0
Seed GR-RL
了解更多
论文
Seed Edge
Seed STEM 科学家计划
校园招聘
Copyright© 2026 Bytedance Seed
网站声明联系我们 : seed.feedback@bytedance.com
Seed2.1Seedance 2.5Seedream 5.0 ProSeedRealtimeSeed Audio 1.0Seed GR-RL
论文Seed EdgeSeed STEM 科学家计划校园招聘
追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
Copyright© 2026 Bytedance Seed
网站声明
联系我们 : seed.feedback@bytedance.com
追求智能上限,创造社会价值
联系我们 : seed.feedback@bytedance.com
Copyright© 2026 Bytedance Seed网站声明