将视频生成模型转化为4D具身世界模型
RT @lixin4ever: Thx @_akhaliq for sharing 🚀🚀
Tri-branch DiT && Joint Cross-Modal Attention && 250M+ RGB frames with dense depth and optical flow annotations
Thats how we turn a video generation model into a 4D embodied world model 💪💪
More details available at https://t.co/kZd74lZAYN
Tri-branch DiT && Joint Cross-Modal Attention && 250M+ RGB frames with dense depth and optical flow annotations
Thats how we turn a video generation model into a 4D embodied world model 💪💪
More details available at https://t.co/kZd74lZAYN