GEN3C：具备精确相机控制的三维感知世界一致性视频生成

摘要

我们推出GEN3C，一款具备精确相机控制与时间维度三维一致性的生成式视频模型。现有视频模型虽能生成逼真视频，却较少利用三维信息，导致诸如物体突然出现或消失等不一致现象。即便实现了相机控制，也往往不够精确，因为相机参数仅作为神经网络的输入，网络需自行推断视频如何依赖于相机。相比之下，GEN3C由三维缓存引导：通过预测种子图像或先前生成帧的逐像素深度获取点云数据。在生成后续帧时，GEN3C以用户提供的新相机轨迹对三维缓存进行二维渲染作为条件。关键在于，这意味着GEN3C既无需记忆先前生成的内容，也不必从相机姿态推断图像结构。相反，模型可集中其全部生成能力于先前未观察到的区域，并将场景状态推进至下一帧。我们的成果展示了比以往工作更精确的相机控制，以及在稀疏视角新视图合成上的领先表现，即便在驾驶场景和单目动态视频等挑战性设置下亦如此。最佳效果请观看视频。访问我们的网页了解更多！https://research.nvidia.com/labs/toronto-ai/GEN3C/

English

We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if implemented at all, is imprecise, because camera parameters are mere inputs to the neural network which must then infer how the video depends on the camera. In contrast, GEN3C is guided by a 3D cache: point clouds obtained by predicting the pixel-wise depth of seed images or previously generated frames. When generating the next frames, GEN3C is conditioned on the 2D renderings of the 3D cache with the new camera trajectory provided by the user. Crucially, this means that GEN3C neither has to remember what it previously generated nor does it have to infer the image structure from the camera pose. The model, instead, can focus all its generative power on previously unobserved regions, as well as advancing the scene state to the next frame. Our results demonstrate more precise camera control than prior work, as well as state-of-the-art results in sparse-view novel view synthesis, even in challenging settings such as driving scenes and monocular dynamic video. Results are best viewed in videos. Check out our webpage! https://research.nvidia.com/labs/toronto-ai/GEN3C/

GEN3C：具备精确相机控制的三维感知世界一致性视频生成

GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

摘要

Summary

Support

Support