北风网Hadoop,The.Definitive.Guid 3rd Early.Release

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

THIRDEDITIONHadoop:TheDefinitiveGuideTomWhiteBeijing•Cambridge•Farnham•Köln•Sebastopol•TokyoHadoop:TheDefinitiveGuide,ThirdEditionbyTomWhiteRevisionHistoryforthe:2012-01-27Earlyreleaserevision1See=9781449311520forreleasedetails.ISBN:978-1-449-31152-01327616795ForEliane,Emilia,andLottieTableofContentsForeword..................................................................xiiiPreface.....................................................................xv1.MeetHadoop...........................................................1Data!1DataStorageandAnalysis3ComparisonwithOtherSystems4RDBMS4GridComputing6VolunteerComputing8ABriefHistoryofHadoop9ApacheHadoopandtheHadoopEcosystem12HadoopReleases13What’sCoveredinthisBook14Compatibility152.MapReduce...........................................................17AWeatherDataset17DataFormat17AnalyzingtheDatawithUnixTools19AnalyzingtheDatawithHadoop20MapandReduce20JavaMapReduce22ScalingOut30DataFlow31CombinerFunctions34RunningaDistributedMapReduceJob37HadoopStreaming37Ruby37Python40iiiHadoopPipes41CompilingandRunning423.TheHadoopDistributedFilesystem.......................................45TheDesignofHDFS45HDFSConcepts47Blocks47NamenodesandDatanodes48HDFSFederation49HDFSHigh-Availability50TheCommand-LineInterface51BasicFilesystemOperations52HadoopFilesystems54Interfaces55TheJavaInterface57ReadingDatafromaHadoopURL57ReadingDataUsingtheFileSystemAPI59WritingData62Directories64QueryingtheFilesystem64DeletingData69DataFlow69AnatomyofaFileRead69AnatomyofaFileWrite72CoherencyModel75ParallelCopyingwithdistcp76KeepinganHDFSClusterBalanced78HadoopArchives78UsingHadoopArchives79Limitations804.HadoopI/O...........................................................83DataIntegrity83DataIntegrityinHDFS83LocalFileSystem84ChecksumFileSystem85Compression85Codecs87CompressionandInputSplits91UsingCompressioninMapReduce92Serialization94TheWritableInterface95WritableClasses98iv|TableofContentsImplementingaCustomWritable105SerializationFrameworks110Avro112File-BasedDataStructures132SequenceFile132MapFile1395.DevelopingaMapReduceApplication....................................145TheConfigurationAPI146CombiningResources147VariableExpansion148ConfiguringtheDevelopmentEnvironment148ManagingConfiguration148GenericOptionsParser,Tool,andToolRunner151WritingaUnitTest154Mapper154Reducer156RunningLocallyonTestData157RunningaJobinaLocalJobRunner157TestingtheDriver161RunningonaCluster162Packaging162LaunchingaJob162TheMapReduceWebUI164RetrievingtheResults167DebuggingaJob169HadoopLogs173RemoteDebugging175TuningaJob176ProfilingTasks177MapReduceWorkflows180DecomposingaProblemintoMapReduceJobs180JobControl182ApacheOozie1826.HowMapReduceWorks................................................187AnatomyofaMapReduceJobRun187ClassicMapReduce(MapReduce1)188YARN(MapReduce2)194Failures200FailuresinClassicMapReduce200FailuresinYARN202JobScheduling204TableofContents|vTheFairScheduler205TheCapacityScheduler205ShuffleandSort205TheMapSide206TheReduceSide207ConfigurationTuning209TaskExecution212TheTaskExecutionEnvironment212SpeculativeExecution213OutputCommitters215TaskJVMReuse216SkippingBadRecords2177.MapReduceTypesandFormats..........................................221MapReduceTypes221TheDefaultMapReduceJob225InputFormats232InputSplitsandRecords232TextInput243BinaryInput247MultipleInputs248DatabaseInput(andOutput)249OutputFormats249TextOutput250BinaryOutput251MultipleOutputs251LazyOutput255DatabaseOutput2568.MapReduceFeatures..................................................257Counters257Built-inCounters257User-DefinedJavaCounters262User-DefinedStreamingCounters266Sorting266Preparation266PartialSort268TotalSort272SecondarySort276Joins281Map-SideJoins282Reduce-SideJoins284SideDataDistribution287vi|TableofContentsUsingtheJobConfiguration287DistributedCache288MapReduceLibraryClasses2949.SettingUpaHadoopCluster............................................295ClusterSpecification295NetworkTopology297ClusterSetupandInstallation299InstallingJava300CreatingaHadoopUser300InstallingHadoop300TestingtheInstallation301SSHConfiguration301HadoopConfiguration302ConfigurationManagement303EnvironmentSettings305ImportantHadoopDaemonProperties309HadoopDaemonAddressesandPorts314OtherHadoopProperties315UserAccountCreation318YARNConfiguration318ImportantYARNDaemonProperties319YARNDaemonAddressesandPorts322Security323KerberosandHadoop324DelegationTokens326OtherSecurityEnhancements327BenchmarkingaHadoopCluster329HadoopBenchmarks329UserJobs331HadoopintheCloud332HadooponAmazonEC233210.AdministeringHadoop.................................................337HDFS337PersistentDataStructures337SafeMode342AuditLogging344Tools344Monitoring349Logging349Metrics350JavaManagementExtensions353TableofContents|viiMaintenance355RoutineAdministrationProcedures355CommissioningandDecommissioningNodes357Upgrades36011.Pig.................................................................365InstallingandRunningPig366ExecutionTypes366Running

1 / 647
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功