• Fri frakt över 249 kr
  • •
  • Snabba leveranser
  • •
  • Billiga böcker
Kundservice

Du är på sajten för privatpersoner.

Företag, bibliotek eller offentlig verksamhet?

Du handlar på classic.bokus.com, där alla dina funktioner finns intakta.
Till classic.bokus.com
Bokus logotyp. Gå till startsidan.
  • Erbjudanden
  • Nyheter
  • Student
  • Topplistor
  • Barn & ungdom
  • Bokus Play
  • E-böcker
  • Pocketböcker
  • Spel & pussel

10% rabatt på allt med kod NYSTART10 →

Sidfot

Mina sidor

    Hjälp

    • Kundservice
    • Vanliga frågor och svar
    • Frakt och leverans
    • Retur vid ångerrätt
    • Reklamera vara
    • Betalning
    • Köpvillkor
    • Allmänna villkor
    • Information om webbplatsens tillgänglighet

    Om Bokus

    • Om oss
    • Pressrum
    • För studenter
    • För företag
    • För bibliotek och offentlig verksamhet
    • För leverantörer
    • Hållbarhet

    Populärt

    • Aktuella erbjudanden
    • Presentkort
    • Studentlitteratur
    • Nya böcker
    • Topplistor
    • Signerade böcker
    • Engelska böcker

    Inspiration

    • Boktips
    • BookTok
    • Populära bokserier
    • Barnbokskaraktärer
    • Populära författare
    Logotyp för Bokus
    Följ oss på Facebook (extern länk)Följ oss på Instagram (extern länk)Följ oss på YouTube (extern länk)Följ oss på TikTok (extern länk)
    bokus @ CookiesAnpassa cookiesIntegritetspolicyKöpvillkor
    Till Citymail hemsida (extern länk)Till Budbee hemsida (extern länk)Till Postnord hemsida (extern länk)Till Schenker hemsida (extern länk)Till Early Bird hemsida (extern länk)Till Walleys hemsida (extern länk)
    1. Data och IT
    2. Informationsteknik: allmänt

    Expert Hadoop Administration

    Managing, Tuning, and Securing Spark, YARN, and HDFS

    AvSam Alapati

    Häftad, Engelska, 2017

    Del i serien Addison-Wesley Data & Analytics Series

    345 kr

    Beställningsvara. Skickas inom 7-10 vardagar. Fri frakt över 249 kr.

    Beskrivning

    The Comprehensive, Up-to-Date Apache Hadoop Administration Handbook and Reference

    “Sam Alapati has worked with production Hadoop clusters for six years. His unique depth of experience has enabled him to write the go-to resource for all administrators looking to spec, size, expand, and secure production Hadoop clusters of any size.”

    –Paul Dix, Series Editor

    In Expert Hadoop® Administration, leading Hadoop administrator Sam R. Alapati brings together authoritative knowledge for creating, configuring, securing, managing, and optimizing production Hadoop clusters in any environment. Drawing on his experience with large-scale Hadoop administration, Alapati integrates action-oriented advice with carefully researched explanations of both problems and solutions. He covers an unmatched range of topics and offers an unparalleled collection of realistic examples.


    Alapati demystifies complex Hadoop environments, helping you understand exactly what happens behind the scenes when you administer your cluster. You’ll gain unprecedented insight as you walk through building clusters from scratch and configuring high availability, performance, security, encryption, and other key attributes. The high-value administration skills you learn here will be indispensable no matter what Hadoop distribution you use or what Hadoop applications you run.


    • Understand Hadoop’s architecture from an administrator’s standpoint
    • Create simple and fully distributed clusters
    • Run MapReduce and Spark applications in a Hadoop cluster
    • Manage and protect Hadoop data and high availability
    • Work with HDFS commands, file permissions, and storage management
    • Move data, and use YARN to allocate resources and schedule jobs
    • Manage job workflows with Oozie and Hue
    • Secure, monitor, log, and optimize Hadoop
    • Benchmark and troubleshoot Hadoop

    Produktinformation

    • Utgivningsdatum:2017-01-26
    • Mått:178 x 229 x 41 mm
    • Vikt:1 271 g
    • Format:Häftad
    • Språk:Engelska
    • Serie:Addison-Wesley Data & Analytics Series
    • Antal sidor:848
    • Upplaga:1
    • Förlag:Pearson Education
    • ISBN:9780134597195

    Utforska kategorier

    • Informationsteknik: allmänt inom Data och IT
    • Databaser inom Data och IT

    Mer om författaren

    Sam R. Alapati has been working with various aspects of the Hadoop environment for the past six years. He is currently the principal Hadoop administrator at Sabre Corporation in Westlake, Texas, and works on a daily basis with multiple large Hadoop 2 clusters. In addition to being the point person for all Hadoop administration at Sabre, Sam manages multiple critical data-science- and data-analysis-related Hadoop job flows and is also an expert Oracle Database Administrator. His vast knowledge of relational databases and SQL contributes to his work with Hadoop related projects. Sam’s recognition in the database and middleware area includes having published 18 well-received books over the past 14 years, mostly on Oracle Database Administration and Oracle Weblogic Server. His experience dealing with numerous configuration, architectural, and performance-related Hadoop issues over the years led him to the realization that many working Hadoop administrators and developers would appreciate having a handy reference such as this book to turn to when creating, managing, securing and optimizing their Hadoop infrastructure.

    Innehållsförteckning

    • Foreword xxviiPreface xxixAcknowledgments xxxvAbout the Author xxxvii  Part I: Introduction to Hadoop—Architecture and Hadoop Clusters 1  Chapter 1: Introduction to Hadoop and Its Environment 3Hadoop—An Introduction 4Cluster Computing and Hadoop Clusters 12Hadoop Components and the Hadoop Ecosphere 15What Do Hadoop Administrators Do? 18Key Differences between Hadoop 1 and Hadoop 2 21Distributed Data Processing: MapReduce and Spark, Hive and Pig 24Data Integration: Apache Sqoop, Apache Flume andApache Kafka 27Key Areas of Hadoop Administration 28Summary 31  Chapter 2: An Introduction to the Architecture of Hadoop 33Distributed Computing and Hadoop 33Hadoop Architecture 34Data Storage—The Hadoop Distributed File System 37Data Processing with YARN, the Hadoop Operating System 48Summary 57  Chapter 3: Creating and Configuring a Simple Hadoop Cluster 59Hadoop Distributions and Installation Types 60Setting Up a Pseudo-Distributed Hadoop Cluster 62Performing the Initial Hadoop Configuration 71Operating the New Hadoop Cluster 86Summary 90  Chapter 4: Planning for and Creating a Fully Distributed Cluster 91Planning Your Hadoop Cluster 92Going from a Single Rack to Multiple Racks 95Creating a Multinode Cluster 102Modifying the Hadoop Configuration 106Starting Up the Cluster 114Configuring Hadoop Services, Web Interfaces and Ports 119Summary 126  Part II: Hadoop Application Frameworks 127  Chapter 5: Running Applications in a Cluster—The MapReduce Framework (and Hive and Pig) 129The MapReduce Framework 129Apache Hive 141Apache Pig 144Summary 145  Chapter 6: Running Applications in a Cluster—The Spark Framework 147What Is Spark? 148Why Spark? 149The Spark Stack 153Installing Spark 155Spark Run Modes 158Understanding the Cluster Managers 159Spark and Data Access 164Summary 167  Chapter 7: Running Spark Applications 169The Spark Programming Model 169Spark Applications 173Architecture of a Spark Application 179Running Spark Applications Interactively 181Creating and Submitting Spark Applications 185Configuring Spark Applications 192Monitoring Spark Applications 194Handling Streaming Data with Spark Streaming 194Using Spark SQL for Handling Structured Data 198Summary 201  Part III: Managing and Protecting Hadoop Data and High Availability 203  Chapter 8: The Role of the NameNode and How HDFS Works 205HDFS—The Interaction between the NameNode and the DataNodes 205Rack Awareness and Topology 209HDFS Data Replication 212How Clients Read and Write HDFS Data 218Understanding HDFS Recovery Processes 224Centralized Cache Management in HDFS 227Hadoop Archival Storage, SSD and Memory (Heterogeneous Storage) 232Summary 241  Chapter 9: HDFS Commands, HDFS Permissions and HDFS Storage 243Managing HDFS through the HDFS Shell Commands 243Using the dfsadmin Utility to Perform HDFS Operations 251Managing HDFS Permissions and Users 255Managing HDFS Storage 260Rebalancing HDFS Data 267Reclaiming HDFS Space 274Summary 276  Chapter 10: Data Protection, File Formats and Accessing HDFS 277Safeguarding Data 278Data Compression 289Hadoop File Formats 295Using Hadoop WebHDFS and HttpFS 308Summary 315  Chapter 11: NameNode Operations, High Availability and Federation 317Understanding NameNode Operations 318The Checkpointing Process 323NameNode Safe Mode Operations 329Configuring HDFS High Availability 334HDFS Federation 349Summary 351  Part IV: Moving Data, Allocating Resources, Scheduling Jobs and Security 353  Chapter 12: Moving Data Into and Out of Hadoop 355Introduction to Hadoop Data Transfer Tools 355Loading Data into HDFS from the Command Line 356Copying HDFS Data between Clusters with DistCp 361Ingesting Data from Relational Databases with Sqoop 365Ingesting Data from External Sources with Flume 388Ingesting Data with Kafka 398Summary 406  Chapter 13: Resource Allocation in a Hadoop Cluster 407Resource Allocation in Hadoop 407The FIFO Scheduler 410The Capacity Scheduler 411The Fair Scheduler 426Comparing the Capacity Scheduler and the Fair Scheduler 435Summary 436  Chapter 14: Working with Oozie to Manage Job Workflows 437Using Apache Oozie to Schedule Jobs 437Oozie Architecture 439Deploying Oozie in Your Cluster 441Understanding Oozie Workflows 446How Oozie Runs an Action 449Creating an Oozie Workflow 454Running an Oozie Workflow Job 461Oozie Coordinators 464Managing and Administering Oozie 470Summary 475  Chapter 15: Securing Hadoop 477Hadoop Security—An Overview 478Hadoop Authentication with Kerberos 481Hadoop Authorization 505Auditing Hadoop 518Securing Hadoop Data 520Other Hadoop-Related Security Initiatives 524Summary 525  Part V: Monitoring, Optimization and Troubleshooting 527  Chapter 16: Managing Jobs, Using Hue and Performing Routine Tasks 529Using the YARN Commands to Manage Hadoop Jobs 530Decommissioning and Recommissioning Nodes 535ResourceManager High Availability 541Performing Common Management Tasks 545Managing the MySQL Database 548Backing Up Important Cluster Data 551Using Hue to Administer Your Cluster 553Implementing Specialized HDFS Features 562Summary 567  Chapter 17: Monitoring, Metrics and Hadoop Logging 569Monitoring Linux Servers 570Hadoop Metrics 576Using Ganglia for Monitoring 579Understanding Hadoop Logging 582Using Hadoop’s Web UIs for Monitoring 599Monitoring Other Hadoop Components 609Summary 610  Chapter 18: Tuning the Cluster Resources, Optimizing MapReduce Jobs and Benchmarking 611How to Allocate YARN Memory and CPU 612Configuring Efficient Performance 621Tuning Map and Reduce Tasks—What the Administrator Can Do 625Optimizing Pig and Hive Jobs 635Benchmarking Your Cluster 638Hadoop Counters 647Optimizing MapReduce 652Summary 658  Chapter 19: Configuring and Tuning Apache Spark on YARN 659Configuring Resource Allocation for Spark on YARN 659Dynamic Resource Allocation when Running Spark on YARN 676Storage Formats and Compressing Data 678Monitoring Spark Applications 681Tuning Garbage Collection 686Tuning Spark Streaming Applications 688Summary 689  Chapter 20: Optimizing Spark Applications 691Revisiting the Spark Execution Model 692Shuffle Operations and How to Minimize Them 694Partitioning and Parallelism (Number of Tasks) 703Optimizing Data Serialization and Compression 710Understanding Spark’s SQL Query Optimizer 712Caching Data 717Summary 723  Chapter 21: Troubleshooting Hadoop—A Sampler 725Space-Related Issues 725Handling YARN Jobs That Are Stuck 731JVM Memory-Allocation and Garbage-Collection Strategies 732Handling Different Types of Failures 737Troubleshooting Spark Jobs 739Debugging Spark Applications 740Summary 742  Chapter 22: Installing VirtualBox and Linux and Cloning the Virtual Machines 743Installing Oracle VirtualBox 744Installing Oracle Enterprise Linux 745Cloning the Linux Server 745  Index 747