当前位置:首页 >英文主页 >中英对照 > 报告详情

字节跳动:2025豆包大模型Seedream 2.0技术报告:原生中英双语图像生成模型(英文版).pdf

上传人: Kell****reet 编号:617256 2025-03-13 33页 40.68MB

下载:

1、Seedream 2.0:A Native Chinese-English Bilingual ImageGeneration Foundation ModelSeed Vision Team,ByteDanceAbstractRapid advancement of diffusion models has catalyzed remarkable progress in the field of imagegeneration.However,prevalent models such as Flux,SD3.5 and Midjourney,still grapple withissue

2、s like model bias,limited text rendering capabilities,and insufficient understanding of Chinesecultural nuances.To address these limitations,we present Seedream 2.0,a native Chinese-Englishbilingual image generation foundation model that excels across diverse dimensions,which adeptlymanages text pro

3、mpt in both Chinese and English,supporting bilingual image generation and textrendering.We develop a powerful data system that facilitates knowledge integration,and a captionsystem that balances the accuracy and richness for image description.Particularly,Seedream isintegrated with a self-developed

4、bilingual large language model(LLM)as a text encode,allowingit to learn native knowledge directly from massive data.This enable it to generate high-fidelityimages with accurate cultural nuances and aesthetic expressions described in either Chinese orEnglish.Beside,Glyph-Aligned ByT5 is applied for f

5、lexible character-level text rendering,while aScaled ROPE generalizes well to untrained resolutions.Multi-phase post-training optimizations,including SFT and RLHF iterations,further improve the overall capability.Through extensiveexperimentation,we demonstrate that Seedream 2.0 achieves state-of-the

6、-art performance acrossmultiple aspects,including prompt-following,aesthetics,text rendering,and structural correctness.Furthermore,Seedream 2.0 has been optimized through multiple RLHF iterations to closely alignits output with human preferences,as revealed by its outstanding ELO score.In addition,

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文介绍了Seedream 2.0,一种先进的中文-英文双语图像生成基础模型。该模型在多个方面表现出色,包括遵循提示、美学、文本渲染和结构正确性。Seedream 2.0通过多阶段优化,包括数据构建、模型预训练和后训练,实现了强大的模型能力。它还具有出色的文本渲染能力,特别是在处理包含复杂中文字符的长文本时。此外,通过与自研的多语言大语言模型(LLM)集成,Seedream 2.0能够直接从大量高质量的中英文数据中学习,从而生成具有准确文化内涵和审美表达的图像。经过多次RLHF优化,该模型输出的图像与人类偏好高度一致。在人类评估中,Seedream 2.0在多个方面均优于其他模型,包括文本-图像对齐、结构正确性和美学质量。
"Seedream 2.0如何处理中英文提示?" "Seedream 2.0在哪些方面表现出色?" "Seedream 2.0如何优化以提高性能?"
客服
商务合作
小程序
服务号
折叠