본문

서브메뉴

Large Language Model-Based Solutions : How to Deliver Value with Cost-Effective Generative AI Applications
Large Language Model-Based Solutions  : How to Deliver Value with Cost-Effective Generativ...
Large Language Model-Based Solutions : How to Deliver Value with Cost-Effective Generative AI Applications

상세정보

자료유형  
 전자책 국외
최종처리일시  
20260202073946.0
ISBN  
9781394240746 (electronic bk.)
ISBN  
9781394240722
DDC  
006.35
저자명  
Subramanian, Shreyas.
서명/저자  
Large Language Model-Based Solutions : How to Deliver Value with Cost-Effective Generative AI Applications
판사항  
1st ed.
형태사항  
1 online resource (221 pages)
총서명  
Tech Today Series
내용주기  
완전내용Cover -- Contents At A Glance -- Title Page -- Copyright Page -- Dedication Page -- About the Author -- About the Technical Editor -- Contents -- Introduction -- GenAI Applications and Large Language Models -- Importance of Cost Optimization -- Challenges and Opportunities -- Micro Case Studies -- OpenAI: Leading the Way -- Hugging Face: Open-Source Community Building -- Bloomberg GPT: LLMs in Large Commercial Institutions -- Who Is This Book For? -- Summary -- Chapter 1 Introduction -- Overview of GenAI Applications and Large Language Models -- The Rise of Large Language Models -- Neural Networks, Transformers, and Beyond -- GenAI vs. LLMs: What's the Difference? -- The Three-Layer GenAI Application Stack -- The Infrastructure Layer -- The Model Layer -- The Application Layer -- Paths to Productionizing GenAI Applications -- Sample LLM-Powered Chat Application -- The Importance of Cost Optimization -- Cost Assessment of the Model Inference Component -- Cost Assessment of the Vector Database Component -- Benchmarking Setup and Results -- Other Factors to Consider -- Cost Assessment of the Large Language Model Component -- Summary -- Chapter 2 Tuning Techniques for Cost Optimization -- Fine-Tuning and Customizability -- Basic Scaling Laws You Should Know -- Parameter-Efficient Fine-Tuning Methods -- Adapters Under the Hood -- Prompt Tuning -- Prefix Tuning -- P-tuning -- IA3 -- Low-Rank Adaptation -- Cost and Performance Implications of PEFT Methods -- Summary -- Chapter 3 Inference Techniques for Cost Optimization -- Introduction to Inference Techniques -- Prompt Engineering -- Impact of Prompt Engineering on Cost -- Estimating Costs for Other Models -- Clear and Direct Prompts -- Adding Qualifying Words for Brief Responses -- Breaking Down the Request -- Example of Using Claude for PII Removal -- Conclusion -- Providing Context.
내용주기  
완전내용Examples of Providing Context -- RAG and Long Context Models -- Recent Work Comparing RAG with Long Content Models -- Conclusion -- Context and Model Limitations -- Indicating a Desired Format -- Example of Formatted Extraction with Claude -- Trade-Off Between Verbosity and Clarity -- Caching with Vector Stores -- What Is a Vector Store? -- How to Implement Caching Using Vector Stores -- Conclusion -- Chains for Long Documents -- What Is Chaining? -- Implementing Chains -- Example Use Case -- Common Components -- Tools That Implement Chains -- Comparing Results -- Conclusion -- Summarization -- Summarization in the Context of Cost and Performance -- Efficiency in Data Processing -- Cost-Effective Storage -- Enhanced Downstream Applications -- Improved Cache Utilization -- Summarization as a Preprocessing Step -- Enhanced User Experience -- Conclusion -- Batch Prompting for Efficient Inference -- Batch Inference -- Experimental Results -- Using the accelerate Library -- Using the DeepSpeed Library -- Batch Prompting -- Example of Using Batch Prompting -- Model Optimization Methods -- Quantization -- Code Example -- Recent Advancements: GPTQ -- Parameter-Efficient Fine-Tuning Methods -- Recap of PEFT Methods -- Code Example -- Cost and Performance Implications -- Summary -- References -- Chapter 4 Model Selection and Alternatives -- Introduction to Model Selection -- Motivating Example: The Tale of Two Models -- The Role of Compact and Nimble Models -- Examples of Successful Smaller Models -- Quantization for Powerful but Smaller Models -- Text Generation with Mistral 7B -- Zephyr 7B and Aligned Smaller Models -- CogVLM for Language-Vision Multimodality -- Prometheus for Fine-Grained Text Evaluation -- Orca 2 and Teaching Smaller Models to Reason -- Breaking Traditional Scaling Laws with Gemini and Phi -- Phi 1, 1.5, and 2 B Models -- Gemini Models.
내용주기  
완전내용Domain-Specific Models -- Step 1 - Training Your Own Tokenizer -- Step 2 - Training Your Own Domain-Specific Model -- More References for Fine-Tuning -- Evaluating Domain-Specific Models vs. Generic Models -- The Power of Prompting with General-Purpose Models -- Summary -- Chapter 5 Infrastructure and Deployment Tuning Strategies -- Introduction to Tuning Strategies -- Hardware Utilization and Batch Tuning -- Memory Occupancy -- Strategies to Fit Larger Models in Memory -- KV Caching -- PagedAttention -- How Does PagedAttention Work? -- Comparisons, Limitations, and Cost Considerations -- AlphaServe -- How Does AlphaServe Work? -- Impact of Batching -- Cost and Performance Considerations -- S3: Scheduling Sequences with Speculation -- How Does S3 Work? -- Performance and Cost -- Streaming LLMs with Attention Sinks -- Fixed to Sliding Window Attention -- Extending the Context Length -- Working with Infinite Length Context -- How Does StreamingLLM Work? -- Performance and Results -- Cost Considerations -- Batch Size Tuning -- Frameworks for Deployment Configuration Testing -- Cloud-NativeInference Frameworks -- Deep Dive into Serving Stack Choices -- Batching Options -- Options in DJL Serving -- High-Level Guidance for Selecting Serving Parameters -- Automatically Finding Good Inference Configurations -- Creating a Generic Template -- Defining a HPO Space -- Searching the Space for Optimal Configurations -- Results of Inference HPO -- Inference Acceleration Tools -- TensorRT and GPU Acceleration Tools -- CPU Acceleration Tools -- Monitoring and Observability -- LLMOps and Monitoring -- Why Is Monitoring Important for LLMs? -- Monitoring and Updating Guardrails -- Summary -- Conclusion -- Index -- EULA.
기타형태저록  
Print version / Subramanian, ShreyasLarge Language Model-Based Solutions. Newark : John Wiley & Sons, Incorporated,c2024. 9781394240722
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260202s2024        xx            o                  0  eng  d
■001EBC31246938
■003MiAaPQ
■00520260202073946.0
■006m          o    d  |            
■007cr  cnu||||||||
■020    ▼a9781394240746▼q(electronic  bk.)
■020    ▼z9781394240722
■035    ▼a(MiAaPQ)EBC31246938
■035    ▼a(Au-PeEL)EBL31246938
■035    ▼a(OCoLC)1428902446
■040    ▼aMiAaPQ▼beng▼erda▼epn▼cMiAaPQ▼dMiAaPQ
■0820  ▼a006.35
■1001  ▼aSubramanian,  Shreyas.
■24510▼aLarge  Language  Model-Based  Solutions  ▼bHow  to  Deliver  Value  with  Cost-Effective  Generative  AI  Applications
■250    ▼a1st  ed.
■264  1▼aNewark▼bJohn  Wiley  &  Sons,  Incorporated▼c2024.
■264  4▼c?024.
■300    ▼a1  online  resource  (221  pages)
■336    ▼atext▼btxt▼2rdacontent
■337    ▼acomputer▼bc▼2rdamedia
■338    ▼aonline  resource▼bcr▼2rdacarrier
■4900  ▼aTech  Today  Series
■5050  ▼aCover  --  Contents  At  A  Glance  --  Title  Page  --  Copyright  Page  --  Dedication  Page  --  About  the Author  --  About  the Technical  Editor  --  Contents  --  Introduction  --  GenAI  Applications  and  Large  Language  Models  --  Importance  of Cost  Optimization  --  Challenges  and  Opportunities  --  Micro  Case  Studies  --  OpenAI:  Leading  the  Way  --  Hugging  Face:  Open-Source  Community  Building  --  Bloomberg  GPT:  LLMs  in  Large  Commercial  Institutions  --  Who  Is  This  Book  For?  --  Summary  --  Chapter  1  Introduction  --  Overview  of  GenAI  Applications  and  Large  Language  Models  --  The  Rise  of Large  Language  Models  --  Neural  Networks,  Transformers,  and  Beyond  --  GenAI  vs.  LLMs:  What's  the  Difference?  --  The  Three-Layer  GenAI  Application  Stack  --  The  Infrastructure  Layer  --  The  Model  Layer  --  The  Application  Layer  --  Paths  to  Productionizing  GenAI  Applications  --  Sample  LLM-Powered  Chat  Application  --  The  Importance  of Cost  Optimization  --  Cost  Assessment  of the  Model  Inference  Component  --  Cost  Assessment  of the  Vector  Database  Component  --  Benchmarking  Setup  and  Results  --  Other  Factors  to  Consider  --  Cost  Assessment  of the  Large  Language  Model  Component  --  Summary  --  Chapter  2  Tuning  Techniques  for Cost  Optimization  --  Fine-Tuning  and  Customizability  --  Basic  Scaling  Laws  You Should  Know  --  Parameter-Efficient  Fine-Tuning  Methods  --  Adapters  Under  the Hood  --  Prompt  Tuning  --  Prefix  Tuning  --  P-tuning  --  IA3  --  Low-Rank  Adaptation  --  Cost  and  Performance  Implications  of  PEFT  Methods  --  Summary  --  Chapter  3  Inference  Techniques  for Cost  Optimization  --  Introduction  to Inference  Techniques  --  Prompt  Engineering  --  Impact  of Prompt  Engineering  on Cost  --  Estimating  Costs  for  Other  Models  --  Clear  and  Direct  Prompts  --  Adding  Qualifying  Words  for  Brief  Responses  --  Breaking  Down  the  Request  --  Example  of  Using  Claude  for  PII  Removal  --  Conclusion  --  Providing  Context.
■5058  ▼aExamples  of  Providing  Context  --  RAG  and  Long  Context  Models  --  Recent  Work  Comparing  RAG  with  Long  Content  Models  --  Conclusion  --  Context  and  Model  Limitations  --  Indicating  a Desired  Format  --  Example  of  Formatted  Extraction  with  Claude  --  Trade-Off  Between  Verbosity  and  Clarity  --  Caching  with Vector  Stores  --  What  Is  a  Vector  Store?  --  How  to  Implement  Caching  Using  Vector  Stores  --  Conclusion  --  Chains  for Long  Documents  --  What  Is  Chaining?  --  Implementing  Chains  --  Example  Use  Case  --  Common  Components  --  Tools  That  Implement  Chains  --  Comparing  Results  --  Conclusion  --  Summarization  --  Summarization  in the  Context  of Cost  and  Performance  --  Efficiency  in  Data  Processing  --  Cost-Effective  Storage  --  Enhanced  Downstream  Applications  --  Improved  Cache  Utilization  --  Summarization  as  a  Preprocessing  Step  --  Enhanced  User  Experience  --  Conclusion  --  Batch  Prompting  for Efficient  Inference  --  Batch  Inference  --  Experimental  Results  --  Using  the  accelerate  Library  --  Using  the  DeepSpeed  Library  --  Batch  Prompting  --  Example  of  Using  Batch  Prompting  --  Model  Optimization  Methods  --  Quantization  --  Code  Example  --  Recent  Advancements:  GPTQ  --  Parameter-Efficient  Fine-Tuning  Methods  --  Recap  of  PEFT  Methods  --  Code  Example  --  Cost  and  Performance  Implications  --  Summary  --  References  --  Chapter  4  Model  Selection  and  Alternatives  --  Introduction  to  Model  Selection  --  Motivating  Example:  The Tale  of Two  Models  --  The  Role  of Compact  and  Nimble  Models  --  Examples  of Successful  Smaller  Models  --  Quantization  for Powerful  but  Smaller  Models  --  Text  Generation  with  Mistral  7B  --  Zephyr  7B  and  Aligned  Smaller  Models  --  CogVLM  for  Language-Vision  Multimodality  --  Prometheus  for Fine-Grained  Text  Evaluation  --  Orca  2  and  Teaching  Smaller  Models  to Reason  --  Breaking  Traditional  Scaling  Laws  with Gemini  and  Phi  --  Phi  1,  1.5,  and  2  B  Models  --  Gemini  Models.
■5058  ▼aDomain-Specific  Models  --  Step 1 -  Training  Your  Own  Tokenizer  --  Step 2 -  Training  Your  Own  Domain-Specific  Model  --  More  References  for  Fine-Tuning  --  Evaluating  Domain-Specific  Models  vs.  Generic  Models  --  The  Power  of Prompting  with General-Purpose  Models  --  Summary  --  Chapter  5  Infrastructure  and  Deployment  Tuning  Strategies  --  Introduction  to Tuning  Strategies  --  Hardware  Utilization  and  Batch  Tuning  --  Memory  Occupancy  --  Strategies  to Fit  Larger  Models  in Memory  --  KV  Caching  --  PagedAttention  --  How  Does  PagedAttention  Work?  --  Comparisons,  Limitations,  and  Cost  Considerations  --  AlphaServe  --  How  Does  AlphaServe  Work?  --  Impact  of  Batching  --  Cost  and  Performance  Considerations  --  S3:  Scheduling  Sequences  with  Speculation  --  How  Does  S3  Work?  --  Performance  and  Cost  --  Streaming  LLMs  with  Attention  Sinks  --  Fixed  to  Sliding  Window  Attention  --  Extending  the  Context  Length  --  Working  with  Infinite  Length  Context  --  How  Does  StreamingLLM  Work?  --  Performance  and  Results  --  Cost  Considerations  --  Batch  Size  Tuning  --  Frameworks  for  Deployment  Configuration  Testing  --  Cloud-NativeInference  Frameworks  --  Deep  Dive  into  Serving  Stack  Choices  --  Batching  Options  --  Options  in  DJL  Serving  --  High-Level  Guidance  for  Selecting  Serving  Parameters  --  Automatically  Finding  Good  Inference  Configurations  --  Creating  a  Generic  Template  --  Defining  a  HPO  Space  --  Searching  the  Space  for  Optimal  Configurations  --  Results  of  Inference  HPO  --  Inference  Acceleration  Tools  --  TensorRT  and  GPU  Acceleration  Tools  --  CPU  Acceleration  Tools  --  Monitoring  and  Observability  --  LLMOps  and  Monitoring  --  Why  Is  Monitoring  Important  for  LLMs?  --  Monitoring  and  Updating  Guardrails  --  Summary  --  Conclusion  --  Index  --  EULA.
■588    ▼aDescription  based  on  publisher  supplied  metadata  and  other  sources.
■590    ▼aElectronic  reproduction.  Ann  Arbor,  Michigan  :  ProQuest  Ebook  Central,  2026.  Available  via  World  Wide  Web.  Access  may  be  limited  to  ProQuest  Ebook  Central  affiliated  libraries.  
■655  4▼aElectronic  books.
■77608▼iPrint  version▼aSubramanian,  Shreyas▼tLarge  Language  Model-Based  Solutions▼dNewark  :  John  Wiley  &  Sons,  Incorporated,c2024▼z9781394240722
■7972  ▼aProQuest  (Firm)
■85640▼uhttps://ebookcentral.proquest.com/lib/baekseok-ebooks/detail.action?docID=31246938▼zClick  to  View

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    BE67233 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.