HiBench构造Hadoop集群

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

HadoopinChina2010HadoopinChina2010Performance,UtilizationandPowerPerformance,UtilizationandPowerCharacterizationofHadoopClustersusingCharacterizationofHadoopClustersusingHiBenchHiBenchJinquanJinquan(Jason)Dai(Jason)DaiCloudComputingArchitectCloudComputingArchitectIntelChinaSoftwareCenterIntelChinaSoftwareCenter20102010--99--7722HadoopinChina2010HadoopinChina2010Background&BiasBackground&Bias5yearsincompilerdevelopment5yearsincompilerdevelopment••LeadarchitectonIntelnetworkprocessorLeadarchitectonIntelnetworkprocessorcompilercompiler––IntelnetworkprocessorsIntelnetworkprocessors••16cores,8hardwarethreadspercore@16cores,8hardwarethreadspercore@Year2002Year2002••ForeshadowthegeneraltrendtomultiForeshadowthegeneraltrendtomulti--core,core,multimulti--threadarchitecturesthreadarchitectures––Focusedonparallelprocessing,performanceFocusedonparallelprocessing,performanceandscalabilityandscalabilityCurrentlyCloudComputingarchitectCurrentlyCloudComputingarchitect••LeadtheworkonmassivelydistributedcloudLeadtheworkonmassivelydistributedcloudplatformsplatforms––CloudstorageCloudstorage––BigDataanalyticsBigDataanalytics––Virtualizedutilitycloud/Virtualizedutilitycloud/IaaSIaaS––OnlinewebserviceOnlinewebservice––ClouddatacenterbuildingblocksClouddatacenterbuildingblocksIntelIXP2800IntelIXP280020102010--99--7733HadoopinChina2010HadoopinChina2010AgendaAgendaHiBenchHiBench••ArealisticandcomprehensiveHadoopbenchmarksuiteArealisticandcomprehensiveHadoopbenchmarksuite••DataflowDataflow--basedworkloadcharacterizationsbasedworkloadcharacterizationsBalancedHadoopClusterArchitectureBalancedHadoopClusterArchitecture••HyperThreadingHyperThreading••SSDSSD••ChannelbondingChannelbondingHadoopPowerCharacterizationHadoopPowerCharacterization••CPUpowerstatesCPUpowerstates••FrequencyscalingFrequencyscaling••ChannelbondingChannelbonding20102010--99--7744HadoopinChina2010HadoopinChina2010HiBenchHiBench:ARealisticandComprehensive:ARealisticandComprehensiveHadoopBenchmarkSuiteHadoopBenchmarkSuiteHiBenchHiBench–EnhancedDFSIOMicroBenchmarksWebSearch–Sort–WordCount–TeraSort–NutchIndexing–PageRankMachineLearning–BayesianClassification–K-MeansClusteringHDFSSeeourpaperSeeourpaper““TheTheHiBenchHiBenchSuite:CharacterizationoftheSuite:CharacterizationoftheMapReduceMapReduce--BasedDataBasedDataAnalysisAnalysis””inICDEinICDE’’10workshops(WISS10workshops(WISS’’10)10)20102010--99--7755HadoopinChina2010HadoopinChina2010CharacterizationofCharacterizationofHiBenchHiBenchWorkloadsWorkloads5WorkloadWorkloadSystemResourceSystemResourceUtilizationUtilizationDataAccessPatternsDataAccessPatternsMap/ReduceMap/ReduceStageTimeRatioStageTimeRatioSortI/OboundWordCountCPUboundTeraSortMapstage:CPU-bound;Redstage:I/O-boundNutchIndexingI/Obound,highCPUutilizationinmapstagePageRank(1st&2ndjob)CPU-boundinalljobsBayesianClassification(1st&2ndjob)I/Obound,withhighCPUutilizationinmapstageinthe1stjobK-meansClusteringCPUboundiniteration;I/OboundinclusteringEnhancedDFSIOI/O-boundtrivialtrivialMMMMRRMMMMMMMMMMRRRRRRRRRRMMRRMMMMRRRRnoreducerdatafewerdataevenfewerdatacompressed20102010--99--7766HadoopinChina2010HadoopinChina2010HadoopDataflowModelHadoopDataflowModelStreamingdataflowStreamingdataflowDDAATTAAMAPMAPMAPMAPMAPMAPMAPMAPreducereduceGroupedGroupedIntermediateIntermediateResultsResultsAggregatedAggregatedOutputOutputInputInputSlitsSlitsHttpHttpserverservercopiercopiercopiercopiercopiercopiersortsortsortsortsortsortmergemergeHttpHttpserverserverHttpHttpserverserverHttpHttpserverservermergemergemergemergereducereducereducereduceSequentialdataflowSequentialdataflowshuffleshuffleshuffleshuffleshuffleshufflespillspillspillspillMapTasksMapTasksReduceTasksReduceTasksStreamingdataflowStreamingdataflowDDDDAAAATTTTAAAAMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPMAPreducereducereducereduceGroupedGroupedIntermediateIntermediateResultsResultsAggregatedAggregatedOutputOutputInputInputSlitsSlitsHttpHttpserverserverHttpHttpserverservercopiercopiercopiercopiercopiercopiercopiercopiercopiercopiercopiercopiersortsortsortsortsortsortsortsortsortsortsortsortmergemergeHttpHttpserverserverHttpHttpserverserverHttpHttpserverserverHttpHttpserverserverHttpHttpserverserverHttpHttpserverservermergemergemergemergereducereducereducereducereducereducereducereduceSequentialdataflowSequentialdataflowshuffleshuffleshuffleshuffleshuffleshuffleshuffleshufflespillspillspillspillMapTasksMapTasksReduceTasksReduceTasks20102010--99--7777HadoopinChina2010HadoopinChina2010Sort,Sort,WordCountWordCountandandTeraSortTeraSortSortSortWordCountWordCountTeraSortTeraSort20102010--99--7788HadoopinChina2010HadoopinChina2010AgendaAgendaHiBenchHiBench••ArealisticandcomprehensiveHadoopbenchmarksuiteArealisticandcomprehensiveHadoopbenchmarksuite••DataflowDataflow--basedworkloadcharacterizationsbasedworkloadcharacterizationsBalancedHadoopClusterArchitectureBalancedHadoopClusterArchitecture••HyperThreadingHyperThreading••SSDSSD••ChannelbondingChannelbondingHadoopPowerCharacterizationHadoopPowerCharacterization••CPUpowerstatesCPUpowerst

1 / 20
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功