当前位置:首页 >英文主页 >中英对照 > 报告详情

阿里视频模型万相2.1技术报告(英文版)(60页).pdf

上传人: 白**** 编号:619009 2025-03-24 60页 20.83MB

下载:

1、WAN:OPEN ANDADVANCEDLARGE-SCALEVIDEOGENERATIVEMODELSWan Team,Alibaba GroupABSTRACTThis report presentsWan,a comprehensive and open suite of video foundationmodels designed to push the boundaries of video generation.Built upon the main-stream diffusion transformer paradigm,Wanachieves signifi cant ad

2、vancementsin generative capabilities through a series of innovations,including our novelspatio-temporal variational autoencoder(VAE),scalable pre-training strategies,large-scale data curation,and automated evaluation metrics.These contributionscollectively enhance the models performance and versatil

3、ity.Specifi cally,Wanischaracterized by four key features:Leading Performance:The 14B model ofWan,trained on a vast dataset comprising billions of images and videos,demonstratesthe scaling laws of video generation with respect to both data and model size.It consistently outperforms the existing open

4、-source models as well as state-of-the-art commercial solutions across multiple internal and external benchmarks,demonstrating a clear and signifi cant performance superiority.Comprehensive-ness:Wanoffers two capable models,i.e.,1.3B and 14B parameters,for effi ciencyand effectiveness respectively.I

5、t also covers multiple downstream applications,including image-to-video,instruction-guided video editing,and personal videogeneration,encompassing up to eight tasks.Meanwhile,Wanis the fi rst modelthat can generate visual text in both Chinese and English,signifi cantly enhancingits practical value.C

6、onsumer-Grade Efficiency:The 1.3B model demonstratesexceptional resource effi ciency,requiring only 8.19 GB VRAM,making it com-patible with a wide range of consumer-grade GPUs.It also exhibits superiorperformance compared to larger open-source models,showcasing remarkable ef-fi ciency for text-to-vi

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文介绍了Wan,一种全面且开放的视频基础模型套件,旨在推动视频生成技术的边界。Wan基于主流的扩散转换器范式,通过一系列创新,包括新颖的时空变分自动编码器(VAE)、可扩展的预训练策略、大规模数据策展和自动化的评估指标,实现了生成能力的显著提升。Wan的14B模型在包含数十亿图像和视频的大规模数据集上进行训练,展示了视频生成在数据和模型大小方面的扩展定律。Wan在多个内部和外部基准测试中均优于现有的开源模型和最先进的商业解决方案,展示了明显的性能优势。Wan提供两种能力模型,包括1.3B和14B参数模型,分别用于效率和效果。它还覆盖了多个下游应用,包括图像到视频、指令引导的视频编辑和个人视频生成,涵盖多达八个任务。Wan是第一个能够生成中英文视觉文本的视频生成模型,显著提高了其实际价值。Wan的1.3B模型展示了卓越的资源效率,仅需8.19 GB VRAM,使其与广泛的消费级GPU兼容。Wan的整个系列,包括源代码和所有模型,都开源,以促进视频生成社区的成长。
"Wan模型如何处理视频中的文本内容?" "Wan模型的训练过程包括哪些阶段?" "Wan模型的视频变分自编码器有哪些创新之处?"
客服
商务合作
小程序
服务号
折叠