+VMware大数据解决方案2014-架构

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

©2014VMwareInc.Allrightsreserved.VMwareBigDataSIGReferenceArchitectureChoicesforVirtualizingHadooponvSphereVMwareBigDataProductTeamMarch27,2014Agenda•Introductions•HadooponVMwareUpdate•ArchitectureConsiderationsforHadooponVMware•Q&ACONFIDENTIAL2HadooponVMwareUpdate•vSphereBigDataExtensionsv1.1–ReleasedJan2014–Multi-networksupport–User-definedpasswords–Elasticityenhancements(min/max,scheduledelasticity)–Upgradecapability•Partnernews–SupportforInteldistributioninBDE1.1–CertifiedwithHortonworksHDP2.0–HadoopStarterKitusingIsilon•Visitusat–EMCWorldinLasVegas–HadoopSummitinSanJoseCONFIDENTIAL3©2014VMwareInc.Allrightsreserved.ArchitectureConsiderationsforHadooponVMwareJustinMurray,TechnicalMarketingManager,VMwareAgendaCONFIDENTIAL5•AQuickLookatvSphereBigDataExtensions•ExampleArchitectures•VirtualizingHadoop:theStandardModel•TheCompute–DataSeparationModel•ALookatStorageArchitectures•Configuration-Summary•QuestionsandAnswersvSphereBigDataExtensionsReferenceArchitecture:32-ServerPerformanceTestCONFIDENTIAL7UptofourVMsperservervCPUsperVMfitwithinsocketsize(e.g.4VMsx4vCPUs,2X8)MemoryperVM-fitwithinNUMAnodesizeNativevs.Virtual,32hosts,16disksperhostSource:(usingTeraSortasanexample)75%ofDiskBandwidthJobMapTaskMapTaskMapTaskMapTaskReduceReduceHDFSDFSInputDataDFSOutputData12%ofBandwidth12%ofBandwidthSpills&Logsspill*.outSpillsMapOutputfile.outShuffleMap_*.outSortCombineIntermediate.outLargerArchitecturewithDataComputeSeparatedCONFIDENTIAL10VirtualizedHadoop–ChoicesforPlacement•C=computenode(TaskTracker)•D=DataNodePhysicalHostsVirtualMachineCombinedModel:aStandardDeploymentVirtualizationHostVMDKOSImage–VMDKHadoopVirtualNode1DataNodeExt4TaskTrackerExt4Ext4Ext4SharedstorageSAN/NASLocaldisksOSImage–VMDKVMDKVMDKVMDKVMDKVMDKVMDKVMDKHadoopVirtualNode2DatanodeExt4TaskTrackerExt4Ext4Ext4TheData-ComputeSeparationDeploymentModelVirtualizationHostOSImage–VMDKHadoopVirtualNode1TaskTrackerSharedstorageSAN/NASLocaldisksOSImage–VMDKVMDKVMDKVMDKVMDKVMDKVMDKVMDKHadoopVirtualNode2DataNodeExt4Ext4Ext4Ext4Ext4Ext4Ext4Ext4Ext4Ext4Ext4VMDKVMDKVMDKVMDKVMDKVMDKVMDKVMDKVMDK…DataPaths:CombinedvsData-ComputeSeparationVirtualizationHostVirtualizationHostHadoopVirtualNode1HadoopVirtualNode2TaskTrackerVirtualSwitchDataNodeHadoopVirtualNodeVirtualSwitchDataNodeTaskTrackerVMDKIsolationModelCompute-onlyClustersWithIsilonSharedstorageSAN/NASHadoopVirtualNode2NNNNNNNNNNNNdatanodeIsilonVirtualizationHostVMDKOSImage–VMDKOSImage–VMDKVMDKVMDKHadoopVirtualNode1Ext4JobTrackerExt4TempOSImage–VMDKExt4TaskTrackerExt4HadoopVirtualNode3Ext4TaskTrackerExt4HybridStorageModel-BestofBothWorlds–Masternodes:•NameNode,JobTracker,ZooKeeperetc.onsharedstorage•LeveragevSphereHAandFT–Workernodes•TaskTracker/DataNodemakeuseoflocalstorage•Lowercost,scalablebandwidth•TaskTrackerTempdataiswrittentolocalstorageforbestperformance(alternatesmaybeontoSSDorNFSstorage)LocalStorageSharedStorageWorkerDAS/SSD/NFSforTemp/ShuffleDataProvisionthevirtualmachinesattherightsize•Reserve6%ofphysicalmemoryontheESXiServerforvSphereusage•Avoidover-commitmentofmemoryintheGuestOSandatthehostlevel•ThevirtualmachinememorysizeandvCPUcountfitwithintheNUMAnode(andsocketsize/logicalprocessorcount)UseLargeMemoryPagesattheJVMandguestOSlevel(vSpheredoesthisathostlevelbydefault,exceptwhenphysicalmemoryislow)Sizethehardwaretocontaintheappropriatenumberofthevirtualmachinesabove.Hadoopisaworkloadthatvirtualizeswell.Thereareplentyofgoodreasonstodothis.vSphereGeneralGuidelines-Summary•VMwarevSphereBDEwebsite://•VirtualizedHadoopPerformancewithVMwarevSphere5.1•ABenchmarkingCaseStudyofVirtualizedHadoopPerformanceonvSphere5•ScalingtheDeploymentofMultipleHadoopWorkloadsonaVirtualizedInfrastructure(Intel-Dell-VMware)•ApacheHadoopHighAvailabilitySolutiononVMwarevSphere5.1•HadoopVirtualizationExtensions(HVE):@vmware.comKevinLeong,ProductManagementkleong@vmware.comJuanNovella,ProductMarketingjnovella@vmware.comBackupSlidesHadoopworkloadsworkverywellonVMwarevSphere•Variousperformancestudieshaveshownthatanydifferencebetweenvirtualizedperformanceandnativeperformanceisminimal•FollowthegeneralbestpracticeguidelinesthatVMwarehaspublishedvSphereBigDataExtensionsenhancesyourHadoopexperienceontheVMwarevirtualizationplatform•RapidprovisioningtoolfordeploymentofHadoopcomponentsinvirtualmachines•AlgorithmsforbestlayoutofyourHadoopdataandclustercomponentsarebuiltintotheBDEVHMandHVEcomponents•Designpatternssuchasdata-computeseparationcanbeusedtoprovideelastic

1 / 28
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功