当前位置:首页 > 报告详情

优化批处理和流式聚合.pdf

上传人: 2*** 编号:139024 2023-06-04 28页 648.58KB

1、Jacek Laskowski/jaceklaskowskiOptimizing Batch and Streaming AggregationsData+AI Summit 2023About the SpeakerJacek Laskowski is a Freelance IT ConsultantSpecializing in Apache Spark,Delta Lake,Databricks,Apache Kafka(incl.Kafka Streams and ksqlDB)Best known by The Internals Of online booksContact me

2、 at jacekjapila.plFollow me at JacekLaskowskiConnect on LinkedInTable of Contents1.The Intro to The Internals of Structured Queries2.The Internals of Aggregate Queries3.Scala UDAFs and Aggregators4.Streaming Aggregates5.Streaming Aggregates Performance Tuning Gig6.Things to Watch Out For(Recap)The I

3、ntro toThe Internals ofStructured QueriesStructured Queries Apache Spark is a general-purpose distributed compute platform Spark SQL is a module of Apache Spark to describe batch queries over structured and semi-structured datasets(of any size)Spark Structured Streaming is a module of Apache Spark f

4、or streaming queries over unbounded data Queries are described using High-Level Query OperatorsDataFrame APISQL In most cases,optimizing streaming queries is to optimize corresponding batch queriesNo need to focus on streaming features(less to worry about)Caveat:streaming issues may really be relate

5、d to how streaming queries workHigh-Level Query Language-DataFrame APIHigh-Level Query Language-SQLQueryExecutionQueryExecution is the execution pipeline(workflow)of a structured queryMade up of execution phasesLogical and Physical OperatorsLogical Operators are building blocks of logical query plan

6、s in Spark SQLAggregateJoinLocalRelationLogicalRDDMergeIntoTableProjectSortPhysical Operators are executable nodes of physical query plans in Spark SQLAdaptiveSparkPlanExecBroadcastHashJoinExecHashAggregateExecObjectHashAggregateExecProjectExecSortAggregateExecThe Internals of Aggregate QueriesAggre

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要介绍了Apache Spark中结构化查询的内部机制,批处理和流处理聚合操作的优化方法。作者Jacek Laskowski是一位自由职业的IT顾问,专注于Apache Spark、Delta Lake、Databricks、Apache Kafka等领域。文章首先概述了Spark SQL的模块,用于描述针对结构化和半结构化数据集的批量查询,以及针对无界数据的流查询。接着,详细讲解了聚合查询的内部原理,包括逻辑和物理操作符,以及聚合函数的使用。文章还讨论了流处理聚合的性能调优,以及在使用过程中需要关注的问题。最后,作者给出了一系列优化建议,如避免使用Scala UDAFs,使用整数类型作为分组键,观察sort fallback tasks Metric等。
"Spark SQL中聚合查询的内部机制是什么?" "如何优化Spark Structured Streaming的聚合查询?" "在Spark中使用UDAF时,有哪些需要注意的性能问题?"
客服
商务合作
小程序
服务号
折叠