10、Spark on YARN

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

●DatascientistatCloudera●RecentlyleadApacheSparkdevelopmentatCloudera●Beforethat,committingonApacheYARNandMapReduce●HadoopPMCmemberTraditionalOperatingSystemStorage:FileSystemExecution/Scheduling:Processes/KernelSchedulerHadoopStorage:HadoopDistributedFileSystem(HDFS)Execution/Scheduling:YARN!HDFSImpalaMapReduceSparkEngineering-50%Finance-30%Marketing-20%SparkMRImpalaSparkMRImpalaSparkMRImpalaHDFSHDFSImpalaMapReduceSparkYARN●RunSparkalongsideotherHadoopworkloads○Leverageexistingclusters○Datalocality●Manageworkloadsusingadvancedpolicies○Allocatesharestodifferentteamsandusers○Hierarchicalqueues○Queueplacementpolicies●TakeadvantageofHadoop’ssecurity○RunonKerberizedclusters●Late2012/Spark0.6-experimentalprojectatYahoo●Late2013/Spark0.8-pulledintoSpark,Hadoop-0.23only●Early2014/Spark0.9-Hadoop2.2lineaswell,supportforspark-shell●Early2014/Spark0.9.1/CDH5.0-Stable!●Mid2014/Spark1.0.0/CDH5.1-Easierappsubmissionwithspark-submitResourceManagerNodeManagerNodeManagerResourceManagerNodeManagerNodeManagerContainerContainerContainerResourceManagerNodeManagerNodeManagerClientResourceManagerNodeManagerNodeManagerContainerApplicationMasterClientResourceManagerNodeManagerNodeManagerContainerMapTaskContainerApplicationMasterContainerReduceTaskClientNodeManagerSparkExecutorAppMaster/ExecutorLauncherTaskTaskNodeManagerSparkExecutorTaskTaskResourceManagerClient/SparkDriverNodeManagerSparkExecutorAppMaster/SparkDriverTaskTaskNodeManagerSparkExecutorTaskTaskResourceManagerClient●Whenrunningajob,SparktriestoplacetasksalongsideHDFSblocks●Problem:SparkneedstoaskYARNforexecutorsbeforeitrunsitsjobs●Solution:TellSparkwhatfilesyou’regoingtotouchwhencreatingSparkContextvallocData=InputFormatInfo.computePreferredLocations(Seq(newInputFormatInfo(conf,classOf[TextInputFormat],newPath(“myfile.txt”)))valsc=newSparkContext(conf,locData)●Sparkholdsontofullresourcesevenwhenappisidle●Givebacktocluster●Requirescontainer-resizing(YARN-1197)●SparkHistoryServer●GenericYARNhistory(YARN-321)●Lookingatlogsshouldbeeasier●Betterdocumentationondatalocality●Removeunnecessarysleeps●Long-runningappsonsecureclusters(YARN-941)

1 / 31
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功