当前位置:首页 > 报告详情

Big Data:From Theory to Systems 樊文飞.pdf

上传人: 张** 编号:153179 2024-01-15 35页 3.81MB

1、Wenfei FanShenzhen Institute of Computing SciencesUniversity of EdinburghBeihang UniversityBig Data:From Theory to Systems1The 5 Vs of Big DataThe study has raised as many questions as it has answered2 Volume:The size of data grows rapidly and continuouslyChina generated 23.9 ZB business data in 202

2、2.It is expected to reach 76.6 ZB in 2027 Velocity:“You cannot afford to make decisions based on yesterdays data”Healthcare,retail,financial services,cyber security,Variety:Relational database D,transaction graph GCan we write a query across D and G in SQL?Veracity:The most challenging issue among t

3、he 5VsReal-life data is dirty:semantic inconsistencies,duplicates,stale data,missing linksValue:Killer APPs?What practical value can we get out of big data?Big Data:Volume,Variety,Velocity,Veracity,ValueThe challenges introduced by digital economyDigital Currency Heterogeneous queries on big data ac

4、ross different models Real-time transaction processing with consistency and reliability requirements Data-driven fraud detection and intelligent analysisChallenges:How to query big data with limited resources?Volume How to answer queries across heterogeneous data models?Variety How to query dynamic

5、data in response to updates?Velocity How to clean dirty data?Veracity What is benefit of big data analytics?Value Smart City Fusion of data from various models(historical BIM/CIM;and newly collected data)Massive data from unreliable data sources Real-time analysis in response to updatesThe need for

6、both theory and systems for big data analytics 3The challenges introduced by AIGC ChatGPT has led to a large number of AIGC startups 73%startups in China focus on application domains,and 14%on LLMs.Most LLMs are developed via fine-tuning of open-source pre-trained models.To make practical use of AIG

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要介绍了大数据的五个V特性:体积、速度、多样性、真实性和价值,以及深圳计算科学研究院在大数据处理方面的研究成果。 1. 体积:数据量快速增长,2022年中国产生了23.9ZB的商业数据,预计2027年将达到76.6ZB。 2. 速度:决策不能基于昨天的数据,如医疗、零售、金融服务等领域。 3. 多样性:数据类型多样,包括关系数据库和事务图等。 4. 真实性:真实数据往往存在语义不一致、重复、陈旧和缺失链接等问题。 5. 价值:大数据分析的实际价值,如杀手级应用。 深圳计算科学研究院开发了YashanDB数据库管理系统,支持混合工作负载,比ClickHouse快18%,比Oracle和MySQL快60倍。此外,还开发了Rock数据质量系统,通过规则学习和逻辑推理提高数据质量。Fishing Fort通过逻辑推理和机器学习进行大数据图分析。MedHunter用于帕金森病的药物再定位,Dream Creak用于锂铁电池制造,Mirror用于在线推荐,Dasan Pass用于预测网络攻击。
如何在大量数据中进行有效查询? 如何处理异构数据模型之间的查询? 如何提高机器学习模型的准确性和可解释性?
客服
商务合作
小程序
服务号
折叠