LongCat-Video
美团 LongCat 团队开源的 13.6B 参数视频生成基础模型,把文生视频、图生视频和视频续写统一进单个框架,原生支持分钟级长视频生成。亮点是长视频不易偏色掉质,靠块稀疏注意力与从粗到细的时空生成策略,几分钟内可产出 720p/30fps 成片,另有音频驱动的数字人 Avatar 1.5 分支,权重为 MIT 协议。官方提示尚未覆盖所有下游场景,高风险用途需自行评估。
README
LongCat-Video
模型简介
我们推出 LongCat-Video,这是一个拥有 13.6B 参数的基础视频生成模型,在文生视频、图生视频和视频续写生成任务上均有强劲表现。它尤其在高效、高质量的长视频生成方面表现出色,是我们迈向世界模型的第一步。
核心特性
- 🌟 多任务统一架构:LongCat-Video 在单一视频生成框架内统一了文生视频、图生视频和视频续写任务。它通过单一模型原生支持所有这些任务,并在每个单独任务上持续保持强劲表现。
- 🌟 长视频生成:LongCat-Video 在视频续写任务上进行了原生预训练,使其能够生成长达数分钟的视频,而不会出现色彩漂移或质量下降。
- 🌟 高效推理:LongCat-Video 通过在时间轴和空间轴上采用由粗到细的生成策略,可在数分钟内生成 $720p$、$30fps$ 的视频。Block Sparse Attention(块稀疏注意力)进一步提升效率,尤其是在高分辨率下。
- 🌟 多奖励 RLHF 带来强劲表现:在多奖励 Group Relative Policy Optimization (GRPO) 的加持下,内部和公开基准上的全面评估表明,LongCat-Video 的性能可与领先的开源视频生成模型以及最新的商业解决方案相媲美。
更多细节,请参阅完整的 LongCat-Video 技术报告。
🎥 预告视频
🔥 最新动态!!
- 2026 年 5 月 21 日:🚀 我们发布了 LongCat-Video-Avatar-1.5,这是一个升级版开源框架,用于音频驱动的人像视频生成。v1.5 用 Whisper-Large 替代了 Wav2Vec2,以实现更精准的口型同步,达成生产级可用的物理合理性和时间稳定性,具备稳健的长视频生成能力,可泛化到风格化领域(动漫、动物、复杂真实世界条件),同时支持单路和多路音频输入,并通过步数蒸馏将推理加速到 8 步。[ 代码 | 🤗 权重 | 项目主页 ]
- 2025 年 12 月 16 日:🚀 我们很高兴宣布发布 LongCat-Video-Avatar,这是一个统一模型,可提供富有表现力且高度动态的音频驱动角色动画,原生支持音频-文本到视频、音频-文本-图像到视频以及视频续写等任务,并可与单路和多路音频输入无缝兼容。本次发布包括我们的 技术报告、推理代码、🤗 模型权重,以及 项目主页。
- 2025 年 10 月 25 日:🚀 我们发布了 LongCat-Video,一个基础视频生成模型。技术报告和模型可在 LongCat-Video 技术报告 和 🤗 Huggingface 获取!
快速开始
安装
克隆仓库:
git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video
安装依赖:
# create conda environment
conda create -n longcat-video python=3.10
conda activate longcat-video
# install torch (configure according to your CUDA version)
pip install torch==2.6.0+cu124 torchvision==0.21.0+cu124 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
# install flash-attn-2
pip install ninja
pip install psutil
pip install packaging
pip install flash_attn==2.7.4.post1
# install other requirements
pip install -r requirements.txt
# install longcat-video-avatar requirements
conda install -c conda-forge librosa
conda install -c conda-forge ffmpeg
pip install -r requirements_avatar.txt
模型配置中默认启用 FlashAttention-2;安装后,你也可以修改模型配置("./weights/LongCat-Video/dit/config.json")以使用 FlashAttention-3 或 xformers。
模型下载
| 模型 | 描述 | 下载链接 |
|---|---|---|
| LongCat-Video | 基础视频生成 | 🤗 Huggingface |
| LongCat-Video-Avatar | 单角色和多角色音频驱动视频生成(wav2vec2) | 🤗 Huggingface |
| LongCat-Video-Avatar-1.5 | 升级版 avatar 模型,使用 Whisper-large-v3 音频编码器,基于蒸馏的快速推理 | 🤗 Huggingface |
使用 huggingface-cli 下载模型:
pip install "huggingface_hub[cli]"
huggingface-cli download meituan-longcat/LongCat-Video --local-dir ./weights/LongCat-Video
huggingface-cli download meituan-longcat/LongCat-Video-Avatar --local-dir ./weights/LongCat-Video-Avatar
huggingface-cli download meituan-longcat/LongCat-Video-Avatar-1.5 --local-dir ./weights/LongCat-Video-Avatar-1.5
运行文生视频
# Single-GPU inference
torchrun run_demo_text_to_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_text_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行图生视频
# Single-GPU inference
torchrun run_demo_image_to_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_image_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行视频续写
# Single-GPU inference
torchrun run_demo_video_continuation.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_video_continuation.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行长视频生成
# Single-GPU inference
torchrun run_demo_long_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_long_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行交互式视频生成
# Single-GPU inference
torchrun run_demo_interactive_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_interactive_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行 LongCat-Video-Avatar
💡 1.5 使用提示💡 1.0 使用提示
- 口型同步准确度: Audio CFG 在 3–5 之间效果最佳。提高 audio CFG 值可获得更好的同步效果。
- 提示词增强: 更长、更具体的提示词比短提示词能带来更好的一致性和自然度。我们建议包含丰富细节,如角色外观、动作和场景上下文(例如*"一位长黑发的年轻女子正在说话并微笑,身穿白色衬衫,坐在明亮的咖啡馆里"*),以获得最佳效果。
- 缓解重复动作: 将参考图像索引(--ref_img_index,默认为 10)设置在 0 到 24 之间可确保更好的一致性;将其设置为 30 有助于减少重复动作。此外,增大掩码帧范围(--mask_frame_range,默认为 3)可进一步帮助缓解重复动作,但过大的值可能引入伪影。
- 超分辨率: 我们的模型兼容 480P 和 720P,可通过 --resolution 控制。
- 双音频模式: 合并模式(将 audio_type 设为 para)需要两段等长音频,生成的音频由两段音频相加得到;拼接模式(将 audio_type 设为 add)不要求输入等长,生成的音频通过依次拼接两段音频形成,任何间隔用静音填充,默认 person1 先说话,person2 后说话。
- 模型版本:
--model_type avatar-v1.0使用 wav2vec2 音频编码器(默认);--model_type avatar-v1.5使用 Whisper-large-v3 音频编码器,以获得更好的口型同步质量。- 蒸馏模式: 添加
--use_distill以启用蒸馏采样(步数更少,推理更快)。使用--model_type avatar-v1.5时必须启用。- INT8 量化: 添加
--use_int8以加载 INT8 量化的 DiT 模型,从而降低显存占用。仅支持--model_type avatar-v1.5。
- 口型同步准确度:Audio CFG 在 3–5 之间效果最佳。提高 audio CFG 值可获得更好的同步效果。
- 提示词增强:在提示词中包含清晰的口部动作线索(例如 talking、speaking),以实现更自然的口型动作。
- 缓解重复动作:将参考图像索引(--ref_img_index,默认为 10)设置在 0 到 24 之间可确保更好的一致性,而选择其他范围(例如 -10 或 30)有助于减少重复动作。此外,增大掩码帧范围(--mask_frame_range,默认为 3)可进一步帮助缓解重复动作,但过大的值可能引入伪影。
- 超分辨率:我们的模型兼容 480P 和 720P,可通过 --resolution 控制。
- 双音频模式:合并模式(将 audio_type 设为 para)需要两段等长音频,生成的音频由两段音频相加得到;拼接模式(将 audio_type 设为 add)不要求输入等长,生成的音频通过依次拼接两段音频形成,任何间隔用静音填充,默认 person1 先说话,person2 后说话。
LongCat-Video-Avatar-1.5
- 单音频到视频生成
# Audio-Text-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=at2v --input_json=assets/avatar/single_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=ai2v --input_json=assets/avatar/single_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Text-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=at2v --input_json=assets/avatar/single_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=ai2v --input_json=assets/avatar/single_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
- 多音频到视频生成
# Audio-Image-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_multi_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --input_json=assets/avatar/multi_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_multi_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --input_json=assets/avatar/multi_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
运行 Streamlit
# Single-GPU inference
streamlit run ./run_streamlit.py --server.fileWatcherType none --server.headless=false
评测结果
文生视频
我们在内部基准上的文生视频 MOS 评测结果。
| MOS 分数 | Veo3 | PixVerse-V5 | Wan 2.2-T2V-A14B | LongCat-Video |
|---|---|---|---|---|
| 可获取性 | 专有 | 专有 | 开源 | 开源 |
| 架构 | - | - | MoE | Dense |
| 总参数量 | - | - | 28B | 13.6B |
| 激活参数量 | - | - | 14B | 13.6B |
| Text-Alignment↑ | 3.99 | 3.81 | 3.70 | 3.76 |
| Visual Quality↑ | 3.23 | 3.13 | 3.26 | 3.25 |
| Motion Quality↑ | 3.86 | 3.81 | 3.78 | 3.74 |
| Overall Quality↑ | 3.48 | 3.36 | 3.35 | 3.38 |
图生视频
我们在内部基准上的图生视频 MOS 评测结果。
| MOS 分数 | Seedance 1.0 | Hailuo-02 | Wan 2.2-I2V-A14B | LongCat-Video |
|---|---|---|---|---|
| 可获取性 | 专有 | 专有 | 开源 | 开源 |
| 架构 | - | - | MoE | Dense |
| 总参数量 | - | - | 28B | 13.6B |
| 激活参数量 | - | - | 14B | 13.6B |
| Image-Alignment↑ | 4.12 | 4.18 | 4.18 | 4.04 |
| Text-Alignment↑ | 3.70 | 3.85 | 3.33 | 3.49 |
| Visual Quality↑ | 3.22 | 3.18 | 3.23 | 3.27 |
| Motion Quality↑ | 3.77 | 3.80 | 3.79 | 3.59 |
| Overall Quality↑ | 3.35 | 3.27 | 3.26 | 3.17 |
社区作品
欢迎社区贡献!请提交 PR 或在 Issue 中通知我们,以添加你的作品。
- CacheDiT 为 LongCat-Video 提供了基于 DBCache 和 TaylorSeer 的全量缓存加速支持,在精度无明显损失的情况下实现了近 1.7 倍的加速。访问他们的示例了解更多详情。
许可协议
模型权重基于 MIT License 发布。
对本仓库的任何贡献均依据 MIT License 授权,除非另有说明。本许可不授予使用美团商标或专利的任何权利。
完整许可文本请参阅 LICENSE 文件。
使用注意事项
本模型并未针对所有可能的下游应用进行专门设计或全面评估。
开发者应考虑大型语言模型的已知局限性,包括在不同语言间的性能差异,并在敏感或高风险场景中部署模型前,仔细评估准确性、安全性和公平性。 开发者及下游用户有责任理解并遵守与其使用场景相关的所有适用法律法规,包括但不限于数据保护、隐私和内容安全要求。
本模型卡中的任何内容均不应被解释为更改或限制模型发布所依据的 MIT License 条款。
引用
如果您觉得我们的工作有用,我们诚挚地鼓励您引用。
@misc{meituanlongcatteam2025longcatvideotechnicalreport,
title={LongCat-Video Technical Report},
author={Meituan LongCat Team and Xunliang Cai and Qilong Huang and Zhuoliang Kang and Hongyu Li and Shijun Liang and Liya Ma and Siyu Ren and Xiaoming Wei and Rixu Xie and Tong Zhang},
year={2025},
eprint={2510.22200},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.22200},
}
@misc{meituanlongcatteam2026longcatvideoavatar15technicalreport,
title={LongCat-Video-Avatar 1.5 Technical Report},
author={Meituan LongCat Team and Xunliang Cai and Meng Cheng and Feng Gao and Zhe Kong and Jiamu Li and Le Li and Weiheng Li and Hongyu Liu and Shuai Tan and Xiaoming Wei and Tianyu Yang and Yong Zhang},
year={2026},
eprint={2605.26486},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.26486},
}
@misc{meituanlongcatteam2025longcatvideoavatartechnicalreport,
title={LongCat-Video-Avatar Technical Report},
author={Meituan LongCat Team},
year={2025},
eprint={},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={},
}
致谢
我们感谢 Wan、UMT5-XXL、Diffusers 和 HuggingFace 仓库的贡献者,感谢他们的开放研究。
联系方式
如有任何问题,请通过 longcat-team@meituan.com 联系我们,或扫描二维码加入我们的微信群。
