• Fri frakt över 249 kr
  • •
  • Snabba leveranser
  • •
  • Billiga böcker
Kundservice

Du är på sajten för privatpersoner.

Företag, bibliotek eller offentlig verksamhet?

Du handlar på classic.bokus.com, där alla dina funktioner finns intakta.
Till classic.bokus.com
Bokus logotyp. Gå till startsidan.
  • Erbjudanden
  • Nyheter
  • Student
  • Topplistor
  • Barn & ungdom
  • Bokus Play
  • E-böcker
  • Pocketböcker
  • Spel & pussel

10% rabatt på allt med kod: NYSTART10 →

Sidfot

Mina sidor

    Hjälp

    • Kundservice
    • Vanliga frågor och svar
    • Frakt och leverans
    • Retur vid ångerrätt
    • Reklamera vara
    • Betalning
    • Köpvillkor
    • Allmänna villkor
    • Information om webbplatsens tillgänglighet

    Om Bokus

    • Om oss
    • Pressrum
    • För studenter
    • För företag
    • För bibliotek och offentlig verksamhet
    • För leverantörer
    • Hållbarhet

    Populärt

    • Aktuella erbjudanden
    • Presentkort
    • Studentlitteratur
    • Nya böcker
    • Topplistor
    • Signerade böcker
    • Engelska böcker

    Inspiration

    • Boktips
    • BookTok
    • Populära bokserier
    • Barnbokskaraktärer
    • Populära författare
    Logotyp för Bokus
    Följ oss på Facebook (extern länk)Följ oss på Instagram (extern länk)Följ oss på YouTube (extern länk)Följ oss på TikTok (extern länk)
    bokus @ CookiesAnpassa cookiesIntegritetspolicyKöpvillkor
    Till Citymail hemsida (extern länk)Till Budbee hemsida (extern länk)Till Postnord hemsida (extern länk)Till Schenker hemsida (extern länk)Till Early Bird hemsida (extern länk)Till Walleys hemsida (extern länk)
    1. Data och IT
    2. Systemvetenskap och AI

    Large Language Model-Based Solutions

    How to Deliver Value with Cost-Effective Generative AI Applications

    AvShreyas Subramanian

    Häftad, Engelska, 2024

    Del i serien Tech Today

    557 kr

    Beställningsvara. Skickas inom 5-8 vardagar. Fri frakt över 249 kr.

    Fler format och utgåvor

    E-bok

    691 kr

    E-bok

    634 kr

    Beskrivning

    Learn to build cost-effective apps using Large Language Models In Large Language Model-Based Solutions: How to Deliver Value with Cost-Effective Generative AI Applications, Principal Data Scientist at Amazon Web Services, Shreyas Subramanian, delivers a practical guide for developers and data scientists who wish to build and deploy cost-effective large language model (LLM)-based solutions. In the book, you'll find coverage of a wide range of key topics, including how to select a model, pre- and post-processing of data, prompt engineering, and instruction fine tuning. The author sheds light on techniques for optimizing inference, like model quantization and pruning, as well as different and affordable architectures for typical generative AI (GenAI) applications, including search systems, agent assists, and autonomous agents. You'll also find: Effective strategies to address the challenge of the high computational cost associated with LLMsAssistance with the complexities of building and deploying affordable generative AI apps, including tuning and inference techniquesSelection criteria for choosing a model, with particular consideration given to compact, nimble, and domain-specific modelsPerfect for developers and data scientists interested in deploying foundational models, or business leaders planning to scale out their use of GenAI, Large Language Model-Based Solutions will also benefit project leaders and managers, technical support staff, and administrators with an interest or stake in the subject.

    Produktinformation

    • Utgivningsdatum:2024-04-29
    • Mått:185 x 234 x 15 mm
    • Vikt:318 g
    • Format:Häftad
    • Språk:Engelska
    • Serie:Tech Today
    • Antal sidor:224
    • Förlag:John Wiley & Sons Inc
    • ISBN:9781394240722

    Utforska kategorier

    • Systemvetenskap och AI inom Data och IT

    Mer om författaren

    SHREYAS SUBRAMANIAN, PhD, is a principal data scientist at AWS, one of the largest organizations building and providing large language models for enterprise use. He is currently advising both internal Amazon teams and large enterprise customers on building, tuning, and deploying Generative AI applications at scale. Shreyas runs machine learning-focused cost optimization workshops, helping them reduce the costs of machine learning applications on the cloud. Shreyas also actively participates in cutting-edge research and development of advanced training, tuning and deployment techniques for foundation models.

    Innehållsförteckning

    • Introduction xixChapter 1: Introduction 1Overview of GenAI Applications and Large Language Models 1The Rise of Large Language Models 1Neural Networks, Transformers, and Beyond 2GenAI vs. LLMs: What’s the Difference? 5The Three-Layer GenAI Application Stack 6The Infrastructure Layer 6The Model Layer 7The Application Layer 8Paths to Productionizing GenAI Applications 9Sample LLM-Powered Chat Application 11The Importance of Cost Optimization 12Cost Assessment of the Model Inference Component 12Cost Assessment of the Vector Database Component 19Benchmarking Setup and Results 20Other Factors to Consider 23Cost Assessment of the Large Language Model Component 24Summary 27Chapter 2: Tuning Techniques for Cost Optimization 29Fine-Tuning and Customizability 29Basic Scaling Laws You Should Know 30Parameter-Efficient Fine-Tuning Methods 32Adapters Under the Hood 33Prompt Tuning 34Prefix Tuning 36P-tuning 39IA3 40Low-Rank Adaptation 44Cost and Performance Implications of PEFT Methods 46Summary 48Chapter 3: Inference Techniques for Cost Optimization 49Introduction to Inference Techniques 49Prompt Engineering 50Impact of Prompt Engineering on Cost 50Estimating Costs for Other Models 52Clear and Direct Prompts 53Adding Qualifying Words for Brief Responses 53Breaking Down the Request 54Example of Using Claude for PII Removal 55Conclusion 59Providing Context 59Examples of Providing Context 60RAG and Long Context Models 60Recent Work Comparing RAG with Long Content Models 61Conclusion 62Context and Model Limitations 62Indicating a Desired Format 63Example of Formatted Extraction with Claude 63Trade-Off Between Verbosity and Clarity 66Caching with Vector Stores 66What Is a Vector Store? 66How to Implement Caching Using Vector Stores 66Conclusion 69Chains for Long Documents 69What Is Chaining? 69Implementing Chains 69Example Use Case 70Common Components 70Tools That Implement Chains 72Comparing Results 76Conclusion 76Summarization 77Summarization in the Context of Cost and Performance 77Efficiency in Data Processing 77Cost-Effective Storage 77Enhanced Downstream Applications 77Improved Cache Utilization 77Summarization as a Preprocessing Step 77Enhanced User Experience 77Conclusion 77Batch Prompting for Efficient Inference 78Batch Inference 78Experimental Results 80Using the accelerate Library 81Using the DeepSpeed Library 81Batch Prompting 82Example of Using Batch Prompting 83Model Optimization Methods 83Quantization 83Code Example 84Recent Advancements: GPTQ 85Parameter-Efficient Fine-Tuning Methods 85Recap of PEFT Methods 85Code Example 86Cost and Performance Implications 87Summary 88References 88Chapter 4: Model Selection and Alternatives 89Introduction to Model Selection 89Motivating Example: The Tale of Two Models 89The Role of Compact and Nimble Models 90Examples of Successful Smaller Models 91Quantization for Powerful but Smaller Models 91Text Generation with Mistral 7B 93Zephyr 7B and Aligned Smaller Models 94CogVLM for Language-Vision Multimodality 95Prometheus for Fine-Grained Text Evaluation 96Orca 2 and Teaching Smaller Models to Reason 98Breaking Traditional Scaling Laws with Gemini and Phi 99Phi 1, 1.5, and 2 B Models 100Gemini Models 102Domain-Specific Models 104Step 1 - Training Your Own Tokenizer 105Step 2 - Training Your Own Domain-Specific Model 107More References for Fine-Tuning 114Evaluating Domain-Specific Models vs. Generic Models 115The Power of Prompting with General-Purpose Models 120Summary 122Chapter 5: Infrastructure and Deployment Tuning Strategies 123Introduction to Tuning Strategies 123Hardware Utilization and Batch Tuning 124Memory Occupancy 126Strategies to Fit Larger Models in Memory 128KV Caching 130PagedAttention 131How Does PagedAttention Work? 131Comparisons, Limitations, and Cost Considerations 131AlphaServe 133How Does AlphaServe Work? 133Impact of Batching 134Cost and Performance Considerations 134S3: Scheduling Sequences with Speculation 134How Does S3 Work? 135Performance and Cost 135Streaming LLMs with Attention Sinks 136Fixed to Sliding Window Attention 137Extending the Context Length 137Working with Infinite Length Context 137How Does StreamingLLM Work? 138Performance and Results 139Cost Considerations 139Batch Size Tuning 140Frameworks for Deployment Configuration Testing 141Cloud-Native Inference Frameworks 142Deep Dive into Serving Stack Choices 142Batching Options 143Options in DJL Serving 144High-Level Guidance for Selecting Serving Parameters 146Automatically Finding Good Inference Configurations 146Creating a Generic Template 148Defining a HPO Space 149Searching the Space for Optimal Configurations 151Results of Inference HPO 153Inference Acceleration Tools 155TensorRT and GPU Acceleration Tools 156CPU Acceleration Tools 156Monitoring and Observability 157LLMOps and Monitoring 157Why Is Monitoring Important for LLMs? 159Monitoring and Updating Guardrails 160Summary 161Conclusion 163Index 181