当前位置:首页 > 报告详情

Apache Paimon:面向数据和多模态人工智能的统一湖存储兼容 Apache Iceberg.pdf

上传人: 可*** 编号:991750 2025-12-07 11页 3.47MB

1、Data LakeFile FormatsReal-Time Lake Format200820162017202220232025The upgrade table format of Hive,high-performance format for huge analytic tablesIncubated Flink-Table-StoreStreaming+Lake FormatApache Paimon(Flink-Table-Store)Real-Time Lake FormatApache Paimon-1.0-1.2 ReleaseTable Format for Increm

2、ental updates with Apache SparkHive Metastore+Hive TablesBuilt on HDFS for Data lake2013File Format:EfficientCompression and scanningApache Paimon Real-Time Lake Format-Timeline of Industry DevelopmentLake FormatsApache FlinkApache HudiApache ORCApache ParquetApache HiveApache Paimon Accelerating th

3、e Data Analytics on Lakehouse ArchitectureData LakeApache Paimon(Real-Time Lake Format compatible with Iceberg)BronzeSilverGoldenStreaming&BatchEnd-to-End Minutes LatencyStreaming&BatchStreaming LakehouseReal-Time IngestionReal-Time UpdateOLAPOnline QueryExternal Data SourcesDatabaseMessage QueueLSM

4、+ParquetApache Paimon Key Designs for Streaming Updates Data is appended as a file and written to Level 0 of LSMAsynchronous Minor Compact Balanced Write and ReadBenchmarkKey DesignsLog Structured Merge TreeReal-TimeUpdate&ChangelogCDC Ingestion Schema Evolution Real-Time Update Benchmark for Paimon

5、,Hudi and Iceberg Using Serverless Flink in Alibaba Cloud,Lower is Better020040060080010001200PaimonHudiIcebergServerless Flink Streaming Update:TPC-H Update Time Cost1.0 X2.5 X3.75 XL0L1L2Minor/MajorAsync CompactionLSM in BucketData FileUsing Paimon CDC to Sync CDC from KafkaUsing Flink CDC to Sync

6、 Database DirectlyApache Paimon Ingestion with Schema Evolution Unified snapshot reading and incremental readingNested Schema EvolutionDatabaseODSFlink CDCDatabaseODSPaimon CDCCDCSnapshotFilesInitalNear Real-Time QueryApache Paimons flourishing development in the industryTaobao&Tmall(Alibaba Group)1

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
根据报告的内容,全文主要概括了Apache Paimon在数据湖架构中的应用和发展。以下是关键点: 1. **Paimon发展历程**:从2013年基于HDFS的文件格式,到支持实时湖格式,再到与Iceberg兼容,Paimon不断进化。 2. **性能优势**:Paimon在实时更新、数据压缩和扫描效率上表现出色,支持分钟级延迟的流批处理。 3. **应用案例**:Paimon被应用于阿里巴巴、vivo、小米、字节跳动等大型企业,处理海量数据,实现数据时效性提升。 4. **核心设计**:Paimon采用LSM树结构,支持异步压缩和读写,以及CDC数据同步。 5. **生态系统**:Paimon与Flink、Iceberg等工具集成,提供Python API和SQL查询支持。 6. **云服务**:Paimon集成于阿里云Data Lake Formation(DLF),提供全托管服务,优化计算和存储。
实时湖格式新宠" "湖仓架构,Paimon加速数据分析" 数据湖格式新篇章"
客服
商务合作
小程序
服务号
折叠