⭐️ Star to follow our team's projects !
🚀🚀🚀 Official implementation of ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.
Here is a video for introducing ShareGPT4Video clearly:
demo_clip_v2.mp4
- Authors: Lin Chen*, Xilin Wei* Jinsong Li*, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao📧, Jiaqi Wang 📧
- Institutes: University of Science and Technology of China; Shanghai AI Laboratory; Peking University;
- Resources: [Paper] [Project Page] [ShareGPT4Video Dataset]
- Models: [🤗ShareGPT4Video-8B] [ShareCaptioner-Video]
- Demo: [🤗ShareGPT4Video-8B] [🤗ShareCaptioner-Video]
- 🔥 A large-scale highly descriptive video-text dataset, 40K GPT4-Vision-generated video captions, around 400K implicit video split captions
- 🔥 A general video captioner for various video durations, resolutions, aspect ratios, approaching GPT4-Vision's caption capability, featuring two inference mode targeted for quality and efficiency, separately.
- 🔥 A superior large video-language model ShareGPT4Video-8B, lasting 4 hours on 8xA100 GPUs of training respectively.
- 🔥 Improving Text-to-Video performance with high-quality video captions generate by our ShareCaptioner-Video
[2024/5/27] The ShareGPT4Video-8B model is released!
[2024/5/26] The ShareGPT4Video dataset and project page are released!
- Training and evaluation code for ShareGPT4Video-8B
- Local ShareCaptioner-Video
- Web demo and local demo of ShareGPT4V-8B
- Checkpoints of ShareGPT4Video-8B
- LLaVA: the codebase we built upon. Thanks for their wonderful work.
- Open-Sora-Plan: an excellent open-source codebase for Sora-like text-to-video implementation. Thanks for their wonderful work.
- Open-LLaVA-NeXT: an open-source codebase for re-producing the training procedure of LLaVA-NeXT series.