1、How does data-driven Al work?一O.归一,守.“,心.-.,一 一,.-,.,-.,.口一一 a 一)Explaining a joke Prompt Explain this joke:Joke:Did you see that Google just hired an eloquent whale for their TPU team?It showed them how to communicate between two different pods I Model Response Prediction:TPUs are a type of compute
2、r chip that Google uses for deep learning.A pod is a group of TPUs.A pod is also a group of whales.The joke is that the whale is able to communicate between two groups of whales,but the speaker is pretending that the whale 1s able to communicate between two groups of TPUs.Pe(x)Pe(yjx)0.06 o.os 0.04
3、0.03 0.02 0.01 0.00 10 20 30 40 50,.4 RIOD ATS Bee U0 LEB?2 2%&8 RERE TRIG tem AERER SIE:RIE:MBBARMIMI AAT PUBS ARAN EREIS?Crile T SOA EPA SAT BSP od Z al#t1T74i!AAO FM):TRUREARAFRESINMiTAAL.S“Pod”=#8TPU.*“Pod”th#0th&,iG SRS ASE ARS ZITSEM,(Wie RE 88 TE PRLATPU Z lalidtT 3 sito hk?What does reinforc
4、ement learning do?Mnhih et al.13 Schulman et al.14&15 am 1)EIT AB?Mnih#Ao 13+SchulmanSA.144015=aa Emergent behaviors with reinforcement learning Impressive because no person had Impressive because it looks like thought of it!something a person might draw!Pi.:7.“|):Google DeepMind “Move 37”in Lee Sed
5、ol AlphaGo match:reinforcement learning“discovers”a move that surprises everyone IT oR OF ERAT A SALIBA,AARAAR BW SAMBIRZ,AABHR REA ZEA RES BIH REY AR!eRe Oca i aan es os)Bt oe N ON Be K ey!”erect?oe x 4 4 ANY?a ee 0:St GAlphaGolb se Pay B37:sR)FS)IN FT TILER A ABR UB EE What does reinforcement lear
6、ning not do?Mnhih et al.13 Schulman et al.14&15 this is done many times enormous gulf ae MIT A?Mnih#A.13 SchulmanSA.144015 BARS So where does that leave us?Data-Driven Al Reinforcement Learning Explaining a joke 2 ee)-.Le are eas;f R a LEE SEDOL Hieawt y lox Explain this joke:y)F 01:33:54 Joke:Did y