• Fri frakt över 249 kr
  • •
  • Snabba leveranser
  • •
  • Billiga böcker
Kundservice

Du är på sajten för privatpersoner.

Företag, bibliotek eller offentlig verksamhet?

Du handlar på classic.bokus.com, där alla dina funktioner finns intakta.
Till classic.bokus.com
Bokus logotyp. Gå till startsidan.
  • Erbjudanden
  • Nyheter
  • Student
  • Topplistor
  • Barn & ungdom
  • Bokus Play
  • E-böcker
  • Pocketböcker
  • Spel & pussel

Upp till 20% på populära nyheter →

Sidfot

Mina sidor

    Hjälp

    • Kundservice
    • Vanliga frågor och svar
    • Frakt och leverans
    • Retur vid ångerrätt
    • Reklamera vara
    • Betalning
    • Köpvillkor
    • Allmänna villkor
    • Information om webbplatsens tillgänglighet

    Om Bokus

    • Om oss
    • Pressrum
    • För studenter
    • För företag
    • För bibliotek och offentlig verksamhet
    • För leverantörer
    • Hållbarhet

    Populärt

    • Aktuella erbjudanden
    • Presentkort
    • Studentlitteratur
    • Nya böcker
    • Topplistor
    • Signerade böcker
    • Engelska böcker

    Inspiration

    • Boktips
    • BookTok
    • Populära bokserier
    • Barnbokskaraktärer
    • Populära författare

    Mina sidor

      Hjälp

      • Kundservice
      • Vanliga frågor och svar
      • Frakt och leverans
      • Retur vid ångerrätt
      • Reklamera vara
      • Betalning
      • Köpvillkor
      • Allmänna villkor
      • Information om webbplatsens tillgänglighet

      Om Bokus

      • Om oss
      • Pressrum
      • För studenter
      • För företag
      • För bibliotek och offentlig verksamhet
      • För leverantörer
      • Hållbarhet

      Populärt

      • Aktuella erbjudanden
      • Presentkort
      • Studentlitteratur
      • Nya böcker
      • Topplistor
      • Signerade böcker
      • Engelska böcker

      Inspiration

      • Boktips
      • BookTok
      • Populära bokserier
      • Barnbokskaraktärer
      • Populära författare
      Logotyp för Bokus
      Följ oss på Facebook (extern länk)Följ oss på Instagram (extern länk)Följ oss på YouTube (extern länk)Följ oss på TikTok (extern länk)
      bokus @ CookiesAnpassa cookiesIntegritetspolicyKöpvillkor
      Till Citymail hemsida (extern länk)Till Budbee hemsida (extern länk)Till Postnord hemsida (extern länk)Till Schenker hemsida (extern länk)Till Early Bird hemsida (extern länk)Till Walleys hemsida (extern länk)
      1. Data och IT
      2. Systemvetenskap och AI

      Professional CUDA C Programming

      AvJohn Cheng,Max Grossman

      Häftad, Engelska, 2014

      559 kr

      Beställningsvara. Skickas inom 3-6 vardagar. Fri frakt över 249 kr.

      Beskrivning

      Break into the powerful world of parallel GPU programming with this down-to-earth, practical guideDesigned for professionals across multiple industrial sectors, Professional CUDA C Programming  presents CUDA -- a parallel computing platform and programming model designed to ease the development of GPU programming -- fundamentals in an easy-to-follow format, and teaches readers how to think in parallel and implement parallel algorithms on GPUs. Each chapter covers a specific topic, and includes workable examples that demonstrate the development process, allowing readers to explore both the "hard" and "soft" aspects of GPU programming.Computing architectures are experiencing a fundamental shift toward scalable parallel computing motivated by application requirements in industry and science. This book demonstrates the challenges of efficiently utilizing compute resources at peak performance, presents modern techniques for tackling these challenges, while increasing accessibility for professionals who are not necessarily parallel programming experts. The CUDA programming model and tools empower developers to write high-performance applications on a scalable, parallel computing platform: the GPU. However, CUDA itself can be difficult to learn without extensive programming experience. Recognized CUDA authorities John Cheng, Max Grossman, and Ty McKercher guide readers through essential GPU programming skills and best practices in Professional CUDA C Programming, including: CUDA Programming ModelGPU Execution ModelGPU Memory modelStreams, Event and ConcurrencyMulti-GPU ProgrammingCUDA Domain-Specific LibrariesProfiling and Performance TuningThe book makes complex CUDA concepts easy to understand for anyone with knowledge of basic software development with exercises designed to be both readable and high-performance. For the professional seeking entrance to parallel computing and the high-performance computing community, Professional CUDA C Programming is an invaluable resource, with the most current information available on the market.

      Produktinformation

      • Utgivningsdatum:2014-10-07
      • Mått:185 x 234 x 28 mm
      • Vikt:885 g
      • Format:Häftad
      • Språk:Engelska
      • Antal sidor:528
      • Förlag:John Wiley & Sons Inc
      • ISBN:9781118739327

      Utforska kategorier

      • Systemvetenskap och AI inom Data och IT

      Mer om författaren

      John Cheng, PHD, is a Research Scientist at BGP International in Houston. He has developed seismic imaging products with GPU technology and many high-performance parallel production applications on heterogeneous computing-platforms.Max Grossman is an expert in GPU computing with experience applying CUDA to problems in medical imaging, machine learning, geophysics, and more.Ty McKercher has been helping customers adopt GPU acceleration technologies while he has been employed at NVIDIA since 2008.

      Innehållsförteckning

      • Foreword xviiPreface xixIntroduction xxiChapter 1: Heterogeneous Parallel Computing with CUDA 1Parallel Computing 2Sequential and Parallel Programming 3Parallelism 4Computer Architecture 6Heterogeneous Computing 8Heterogeneous Architecture 9Paradigm of Heterogeneous Computing 12CUDA: A Platform for Heterogeneous Computing 14Hello World from GPU 17Is CUDA C Programming Difficult? 20Summary 21Chapter 2: CUDA Programming Model 23Introducing the CUDA Programming Model 23CUDA Programming Structure 25Managing Memory 26Organizing Threads 30Launching a CUDA Kernel 36Writing Your Kernel 37Verifying Your Kernel 39Handling Errors 40Compiling and Executing 40Timing Your Kernel 43Timing with CPU Timer 44Timing with nvprof 47Organizing Parallel Threads 49Indexing Matrices with Blocks and Threads 49Summing Matrices with a 2D Grid and 2D Blocks 53Summing Matrices with a 1D Grid and 1D Blocks 57Summing Matrices with a 2D Grid and 1D Blocks 58Managing Devices 60Using the Runtime API to Query GPU Information 61Determining the Best GPU 63Using nvidia-smi to Query GPU Information 63Setting Devices at Runtime 64Summary 65Chapter 3: CUDA Execution Model 67Introducing the CUDA Execution Model 67GPU Architecture Overview 68The Fermi Architecture 71The Kepler Architecture 73Profile-Driven Optimization 78Understanding the Nature of Warp Execution 80Warps and Thread Blocks 80Warp Divergence 82Resource Partitioning 87Latency Hiding 90Occupancy 93Synchronization 97Scalability 98Exposing Parallelism 98Checking Active Warps with nvprof 100Checking Memory Operations with nvprof 100Exposing More Parallelism 101Avoiding Branch Divergence 104The Parallel Reduction Problem 104Divergence in Parallel Reduction 106Improving Divergence in Parallel Reduction 110Reducing with Interleaved Pairs 112Unrolling Loops 114Reducing with Unrolling 115Reducing with Unrolled Warps 117Reducing with Complete Unrolling 119Reducing with Template Functions 120Dynamic Parallelism 122Nested Execution 123Nested Hello World on the GPU 124Nested Reduction 128Summary 132Chapter 4: Global Memory 135Introducing the CUDA Memory Model 136Benefits of a Memory Hierarchy 136CUDA Memory Model 137Memory Management 145Memory Allocation and Deallocation 146Memory Transfer 146Pinned Memory 148Zero-Copy Memory 150Unified Virtual Addressing 156Unified Memory 157Memory Access Patterns 158Aligned and Coalesced Access 158Global Memory Reads 160Global Memory Writes 169Array of Structures versus Structure of Arrays 171Performance Tuning 176What Bandwidth Can a Kernel Achieve? 179Memory Bandwidth 179Matrix Transpose Problem 180Matrix Addition with Unified Memory 195Summary 199Chapter 5: Shared Memory and Constant Memory 203Introducing CUDA Shared Memory 204Shared Memory 204Shared Memory Allocation 206Shared Memory Banks and Access Mode 206Configuring the Amount of Shared Memory 212Synchronization 214Checking the Data Layout of Shared Memory 216Square Shared Memory 217Rectangular Shared Memory 225Reducing Global Memory Access 232Parallel Reduction with Shared Memory 232Parallel Reduction with Unrolling 236Parallel Reduction with Dynamic Shared Memory 238Effective Bandwidth 239Coalescing Global Memory Accesses 239Baseline Transpose Kernel 240Matrix Transpose with Shared Memory 241Matrix Transpose with Padded Shared Memory 245Matrix Transpose with Unrolling 246Exposing More Parallelism 249Constant Memory 250Implementing a 1D Stencil with Constant Memory 250Comparing with the Read-Only Cache 253The Warp Shuffle Instruction 255Variants of the Warp Shuffle Instruction 256Sharing Data within a Warp 258Parallel Reduction Using the Warp Shuffle Instruction 262Summary 264Chapter 6: Streams and Concurrency 267Introducing Streams and Events 268CUDA Streams 269Stream Scheduling 271Stream Priorities 273CUDA Events 273Stream Synchronization 275Concurrent Kernel Execution 279Concurrent Kernels in Non-NULL Streams 279False Dependencies on Fermi GPUs 281Dispatching Operations with OpenMP 283Adjusting Stream Behavior Using Environment Variables 284Concurrency-Limiting GPU Resources 286Blocking Behavior of the Default Stream 287Creating Inter-Stream Dependencies 288Overlapping Kernel Execution and Data Transfer 289Overlap Using Depth-First Scheduling 289Overlap Using Breadth-First Scheduling 293Overlapping GPU and CPU Execution 294Stream Callbacks 295Summary 297Chapter 7: Tuning Instruction-Level Primitives 299Introducing CUDA Instructions 300Floating-Point Instructions 301Intrinsic and Standard Functions 303Atomic Instructions 304Optimizing Instructions for Your Application 306Single-Precision vs. Double-Precision 306Standard vs. Intrinsic Functions 309Understanding Atomic Instructions 315Bringing It All Together 322Summary 324Chapter 8: GPU-Accelerated CUDA Libraries and OpenACC 327Introducing the CUDA Libraries 328Supported Domains for CUDA Libraries 329A Common Library Workflow 330The CUSPARSE Library 332cuSPARSE Data Storage Formats 333Formatting Conversion with cuSPARSE 337Demonstrating cuSPARSE 338Important Topics in cuSPARSE Development 340cuSPARSE Summary 341The cuBLAS Library 341Managing cuBLAS Data 342Demonstrating cuBLAS 343Important Topics in cuBLAS Development 345cuBLAS Summary 346The cuFFT Library 346Using the cuFFT API 347Demonstrating cuFFT 348cuFFT Summary 349The cuRAND Library 349Choosing Pseudo- or Quasi- Random Numbers 349Overview of the cuRAND Library 350Demonstrating cuRAND 354Important Topics in cuRAND Development 357CUDA Library Features Introduced in CUDA 6 358Drop-In CUDA Libraries 358Multi-GPU Libraries 359A Survey of CUDA Library Performance 361cuSPARSE versus MKL 361cuBLAS versus MKL BLAS 362cuFFT versus FFTW versus MKL 363CUDA Library Performance Summary 364Using OpenACC 365Using OpenACC Compute Directives 367Using OpenACC Data Directives 375The OpenACC Runtime API 380Combining OpenACC and the CUDA Libraries 382Summary of OpenACC 384Summary 384Chapter 9: Multi-GPU Programming 387Moving to Multiple GPUs 388Executing on Multiple GPUs 389Peer-to-Peer Communication 391Synchronizing across Multi-GPUs 392Subdividing Computation across Multiple GPUs 393Allocating Memory on Multiple Devices 393Distributing Work from a Single Host Thread 394Compiling and Executing 395Peer-to-Peer Communication on Multiple GPUs 396Enabling Peer-to-Peer Access 396Peer-to-Peer Memory Copy 396Peer-to-Peer Memory Access with Unified Virtual Addressing 398Finite Difference on Multi-GPU 400Stencil Calculation for 2D Wave Equation 400Typical Patterns for Multi-GPU Programs 4012D Stencil Computation with Multiple GPUs 403Overlapping Computation and Communication 405Compiling and Executing 406Scaling Applications across GPU Clusters 409CPU-to-CPU Data Transfer 410GPU-to-GPU Data Transfer Using Traditional MPI 413GPU-to-GPU Data Transfer with CUDA-aware MPI 416Intra-Node GPU-to-GPU Data Transfer with CUDA-Aware MPI 417Adjusting Message Chunk Size 418GPU to GPU Data Transfer with GPUDirect RDMA 419Summary 422Chapter 10: Implementation Considerations 425The CUDA C Development Process 426APOD Development Cycle 426Optimization Opportunities 429CUDA Code Compilation 432CUDA Error Handling 437Profile-Driven Optimization 438Finding Optimization Opportunities Using nvprof 439Guiding Optimization Using nvvp 443NVIDIA Tools Extension 446CUDA Debugging 448Kernel Debugging 448Memory Debugging 456Debugging Summary 462A Case Study in Porting C Programs to CUDA C 462Assessing crypt 463Parallelizing crypt 464Optimizing crypt 465Deploying Crypt 472Summary of Porting crypt 475Summary 476Appendix: Suggested Readings 477Index 481
      Hoppa över listan

      Mer från samma författare

      John Cheng - Astounding Wonder, E-bok

      Astounding Wonder

      John Cheng

      E-bok
      2012

      656 kr

      John Cheng - Astounding Wonder, Häftad

      Astounding Wonder

      John Cheng

      Häftad, 2013

      452 kr

      Ty McKercher, Max Grossman, John Cheng - Professional CUDA C Programming, E-bok

      Professional CUDA C Programming

      Ty McKercher, Max Grossman, John Cheng

      E-bok
      2014

      579 kr

      Ty McKercher, Max Grossman, John Cheng - Professional CUDA C Programming, E-bok

      Professional CUDA C Programming

      Ty McKercher, Max Grossman, John Cheng

      E-bok
      2014

      579 kr

      Hoppa över listan

      Du kanske också är intresserad av

      Ty McKercher, Max Grossman, John Cheng - Professional CUDA C Programming, E-bok

      Professional CUDA C Programming

      Ty McKercher, Max Grossman, John Cheng

      E-bok
      2014

      579 kr

      Ty McKercher, Max Grossman, John Cheng - Professional CUDA C Programming, E-bok

      Professional CUDA C Programming

      Ty McKercher, Max Grossman, John Cheng

      E-bok
      2014

      579 kr

      Paolo Alei, Max Grossman - Building Family Identity, Inbunden
      Del 5

      Building Family Identity

      Paolo Alei, Max Grossman

      Inbunden, 2019

      810 kr

      John Cheng - Astounding Wonder, Häftad

      Astounding Wonder

      John Cheng

      Häftad, 2013

      452 kr

      John Cheng - Astounding Wonder, E-bok

      Astounding Wonder

      John Cheng

      E-bok
      2012

      656 kr

      Brendan Gregg - Systems Performance, Häftad

      Systems Performance

      Brendan Gregg

      Häftad, 2021

      752 kr

      Mark A Richards, Mark A. Richards, William L. Melvin - Principles of Modern Radar, Inbunden

      Principles of Modern Radar

      Mark A Richards, Mark A. Richards, William L. Melvin

      Inbunden, 2023

      1 534 kr

      Brian Kernighan, Rob Pike - The Practice of Programming, Häftad

      The Practice of Programming

      Brian Kernighan, Rob Pike

      Häftad, 1999

      4,0 utav 5 stjärnor. Totalt antal röster:(1)

      650 kr

      Carola Häggkvist - SIGNERAD - Jag är Carola, Inbunden
      • Signerad!

      SIGNERAD - Jag är Carola

      Carola Häggkvist

      Inbunden, 2026

      269 kr

      Måns Petter Zelmerlöw - När allt faller, Inbunden
      • -12%

      När allt faller

      Måns Petter Zelmerlöw

      Inbunden, 2026

      229 kr259 kr