当前位置:首页 > 报告详情

罗震霄(3).pdf

上传人: 拾亿 编号:751796 2025-07-29 35页 2.60MB

1、Zhenxiao Luo(罗罗震霄震霄)Sr.Staff EngineerPinterestLast Mile Data Processing for LLM at Pinterest目目录录 What is Pinterest How Pinterest Leverages Large Language Models Legacy data processing pipelines Painpoints Ray Introduction Batch Inference using Ray Multi-Model Inference CarryOver Columns Accumulators

2、 LLM Inference results Pinterest Integration Next StepsZhenxiao LuoSr.Staff Software Engineer PinterestSr.Staff Software Engineer PinterestPresto Committer&Technical Steering Committee member since 2019Work on Data at Uber,Twitter,Facebook,Netflix,Cloudera,Vertica etc.Bachelor from Fudan UniversityP

3、h.D.(on leave)from the University of Wisconsin Madison讲师简讲师简介介Pinterest#1 Image Sharing Social NetworkMAU:500+MillionPublish and discover recipes,home,style,motivation,and inspiration on the Internetdouble digit growth in both MAU and RevenueWhat is Large Language Model(LLM)Large-scale transformer m

4、odels Ability to recognize input patterns and generate text output Billions of parametersPinterest has in house LLM Data Privacy Cost Service AvailabilityThrough batch inference,we leverage OSS LLM at Pinterest to:Enable new use cases Alternative service to OpenAI ChatGPT APILLM Batch Inference Plat

5、form+Inference Backend Platform-Ray Batch Inference Inference optimization Flash Attention vLLMInference Optimization-Flash Attention Good for non-sequence generation use cases-e.g.embedding extraction Flash attention accelerate model forward pass during inference Transformers consist of attention o

6、perations,Attention is memory bandwidth-bound,flash attention reduces memory complexity from quadratic to linearvLLM A community open-source project for efficient LLM sequence generation Fast PagedAttention for KV cache Continuous batching Quantization Optimized CUDA kernels Flexible Python-based Hu

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要介绍了Pinterest如何利用大型语言模型(LLM)进行数据处理,并采用Ray框架优化批处理推断。关键点如下: 1. Pinterest是一个图片分享社交网络,月活跃用户数超过5亿。 2. 大型语言模型具有大规模Transformer模型,能识别输入模式并生成文本输出,参数量达数十亿。 3. Pinterest采用内部LLM,关注数据隐私、成本和服务可用性。 4. 通过Ray Batch Inference平台和优化技术(如Flash Attention和vLLM),提高推断效率,降低成本。 5. Ray框架支持异构资源管理,提高开发速度,支持多种机器学习库。 6. Ray Data SDK有助于批处理推断与其它数据操作结合,提高作业吞吐量。 7. 引入Ray后,相关团队实现了显著的性能提升和成本降低,如Related Pins团队吞吐量提高4.7倍,成本降低20%;搜索质量团队年成本降低30倍。 8. Pinterest已整合Kubernetes、日志和监控、认证授权审计等功能,并使用Presto分析数据。 9. 未来的步骤包括Ray自动扩展、调优以及提高数据加载性能。 核心数据:月活跃用户数500+百万,Related Pins团队吞吐量提高4.7倍,成本降低20%;搜索质量团队年成本降低30倍。
"Pinterest如何利用大语言模型?" "Ray框架在Pinterest有哪些神操作?" "LLM数据处理中的痛点有哪些?"
客服
商务合作
小程序
服务号
折叠