当前位置:首页 > 报告详情

使用 Spark Streaming 和 Delta Lake 将身份图谱提取扩展到每秒 100 万个事件.pdf

上传人: Fl****zo 编号:718756 2025-06-22 43页 2.07MB

1、Scaling Identity Graph Ingestion to 1M events/sec using Spark structured streaming&Delta LakeAkanksha NagpalJianmei Ye2021 Adobe.All Rights Reserved.Adobe Confidential.2021 Adobe.All Rights Reserved.Adobe Confidential.IntroductionAkanksha NagpalSr.Software Engineer,AdobeLinkedIn:akanksha_nagpalJianm

2、ei YeSr.Software Engineer,AdobeLinkedIn:jianmeiye2020 Adobe.All Rights Reserved.Adobe Confidential.Agenda Overview Adobe Identity Graph Journey to 10 x Scaling&Optimization techniques Privacy Compliance Strategies Custom deployment workflows Lessons Learned&Takeaways Q/A Overview2020 Adobe.All Right

3、s Reserved.Adobe Confidential.Adobe Experience Platform:Real time CDPUnified Customer ProfilesActionable AudiencesActivationPersonalization at scale2020 Adobe.All Rights Reserved.Adobe Confidential.Adobe Identity GraphWhat is Identity Graph?Unifies fragmented identifiers(e.g.,emails,device IDs,cooki

4、es)into a single viewEnables consistent consumer recognition across channels and devices2020 Adobe.All Rights Reserved.Adobe Confidential.Identity Service Petabytes data processed daily1M records/sec50+billion identitiesThese capabilities are built from the ground up to link disconnected identities

5、into a single,unified profile to deliver consistent,connected experiences.70B+billion records/dayJourney to 10 x Evolution Phases2020 Adobe.All Rights Reserved.Adobe Confidential.Data Collection Streaming Topic 1PipelineStream Processing PipelineStreaming Topic 2Streaming Topic 3Initial Architecture

6、Identity Graph storeJob ManagerJob ServiceTask ManagerKubernetes Cluster2020 Adobe.All Rights Reserved.Adobe Confidential.Challenges Why we need to evolve?Operational OverheadSingle pipeline Coupled processing+StorageFragmented logic across batch&streaming Noisy Neighbor Multi-tenant pipelineWe init

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文介绍了Adobe如何使用Spark结构化流和Delta Lake将身份图表摄入规模扩展到每秒100万事件。关键点如下: 1. **Adobe身份图表**:统一了分散的标识符,如电子邮件、设备ID和Cookie,以实现跨渠道和设备的持续消费者识别。 2. **挑战与演进**:从初始的Apache Flink和Kubernetes架构转变为使用Databricks集群和Spark结构化流,以解决操作开销、处理存储耦合、逻辑碎片等问题。 3. **优化技术**:采用微批处理、去重、异步任务执行和多线程处理等技术,提高了资源利用率和吞吐量一致性。 4. **多租户公平性**:通过速率限制解决“噪声邻居”问题,确保多租户在规模上的公平性。 5. **隐私合规**:采用Delta Lake的安全清理策略,确保数据保留和隐私合规。 6. **部署工作流**:采用蓝绿部署机制,实现25+次Databricks部署,并跨AWS和Azure多云环境操作。 7. **核心数据**:每日处理PB级数据,每秒处理100万条记录,70亿+记录/天。 文章强调了可扩展的身份图表摄入、内置去重与数据倾斜处理、异步执行模型和元数据解析、异常检测与噪声邻居处理、隐私安全的清理策略以及可靠的蓝绿部署等关键学习点。
如何实现1M事件/秒的规模? Spark结构流如何优化? 如何确保多租户公平性?
客服
商务合作
小程序
服务号
折叠