서브메뉴
검색
The Cloud Data Lake : A Guide to Building Robust Cloud Data Architecture
The Cloud Data Lake : A Guide to Building Robust Cloud Data Architecture
상세정보
- 자료유형
- 전자책 국외
- 최종처리일시
- 20260202073946.0
- ISBN
- 9781098116552 (electronic bk.)
- ISBN
- 9781098116583
- DDC
- 004.6782
- 서명/저자
- The Cloud Data Lake : A Guide to Building Robust Cloud Data Architecture
- 판사항
- 1st ed.
- 형태사항
- 1 online resource (247 pages)
- 내용주기
- 완전내용Intro -- Copyright -- Table of Contents -- Preface -- Why I Wrote This Book -- Who Should Read This Book? -- Introducing Klodars Corporation -- Navigating the Book -- Conventions Used in This Book -- O'Reilly Online Learning -- How to Contact Us -- Acknowledgments -- Chapter 1. Big Data-Beyond the Buzz -- What Is Big Data? -- Elastic Data Infrastructure-The Challenge -- Cloud Computing Fundamentals -- Cloud Computing Terminology -- Value Proposition of the Cloud -- Cloud Data Lake Architecture -- Limitations of On-Premises Data Warehouse Solutions -- What Is a Cloud Data Lake Architecture? -- Benefits of a Cloud Data Lake Architecture -- Defining Your Cloud Data Lake Journey -- Summary -- Chapter 2. Big Data Architectures on the Cloud -- Why Klodars Corporation Moves to the Cloud -- Fundamentals of Cloud Data Lake Architectures -- A Word on Variety of Data -- Cloud Data Lake Storage -- Big Data Analytics Engines -- Cloud Data Warehouses -- Modern Data Warehouse Architecture -- Reference Architecture -- Sample Use Case for a Modern Data Warehouse Architecture -- Benefits and Challenges of Modern Data Warehouse Architecture -- Data Lakehouse Architecture -- Reference Architecture for the Data Lakehouse -- Sample Use Case for Data Lakehouse Architecture -- Benefits and Challenges of the Data Lakehouse Architecture -- Data Warehouses and Unstructured Data -- Data Mesh -- Reference Architecture -- Sample Use Case for a Data Mesh Architecture -- Challenges and Benefits of a Data Mesh Architecture -- What Is the Right Architecture for Me? -- Know Your Customers -- Know Your Business Drivers -- Consider Your Growth and Future Scenarios -- Design Considerations -- Hybrid Approaches -- Summary -- Chapter 3. Design Considerations for Your Data Lake -- Setting Up the Cloud Data Lake Infrastructure -- Identify Your Goals.
- 내용주기
- 완전내용Plan Your Architecture and Deliverables -- Implement the Cloud Data Lake -- Release and Operationalize -- Organizing Data in Your Data Lake -- A Day in the Life of Data -- Data Lake Zones -- Organization Mechanisms -- Introduction to Data Governance -- Actors Involved in Data Governance -- Data Classification -- Metadata Management, Data Catalog, and Data Sharing -- Data Access Management -- Data Quality and Observability -- Data Governance at Klodars Corporation -- Data Governance Wrap-Up -- Manage Data Lake Costs -- Demystifying Data Lake Costs on the Cloud -- Data Lake Cost Strategy -- Summary -- Chapter 4. Scalable Data Lakes -- A Sneak Peek into Scalability -- What Is Scalability? -- Scale in Our Day-to-Day Life -- Scalability in Data Lake Architectures -- Internals of Data Lake Processing Systems -- Data Copy Internals -- ELT/ETL Processing Internals -- A Note on Other Interactive Queries -- Considerations for Scalable Data Lake Solutions -- Pick the Right Cloud Offerings -- Plan for Peak Capacity -- Data Formats and Job Profile -- Summary -- Chapter 5. Optimizing Cloud Data Lake Architectures for Performance -- Basics of Measuring Performance -- Goals and Metrics for Performance -- Measuring Performance -- Optimizing for Faster Performance -- Cloud Data Lake Performance -- SLAs, SLOs, and SLIs -- Example: How Klodars Corporation Managed Its SLAs, SLOs, and SLIs -- Drivers of Performance -- Performance Drivers for a Copy Job -- Performance Drivers for a Spark Job -- Optimization Principles and Techniques for Performance Tuning -- Data Formats -- Data Organization and Partitioning -- Choosing the Right Configurations on Apache Spark -- Minimize Overheads with Data Transfer -- Premium Offerings and Performance -- The Case of Bigger Virtual Machines -- The Case of Flash Storage -- Summary -- Chapter 6. Deep Dive on Data Formats.
- 내용주기
- 완전내용Why Do We Need These Open Data Formats? -- Why Do We Need to Store Tabular Data? -- Why Is It a Problem to Store Tabular Data in a Cloud Data Lake Storage? -- Delta Lake -- Why Was Delta Lake Founded? -- How Does Delta Lake Work? -- When Do You Use Delta Lake? -- Apache Iceberg -- Why Was Apache Iceberg Founded? -- How Does Apache Iceberg Work? -- When Do You Use Apache Iceberg? -- Apache Hudi -- Why Was Apache Hudi Founded? -- How Does Apache Hudi Work? -- When Do You Use Apache Hudi? -- Summary -- Chapter 7. Decision Framework for Your Architecture -- Cloud Data Lake Assessment -- Cloud Data Lake Assessment Questionnaire -- Analysis for Your Cloud Data Lake Assessment -- Starting from Scratch -- Migrating an Existing Data Lake or Data Warehouse to the Cloud -- Improving an Existing Cloud Data Lake -- Phase 1 of Decision Framework: Assess -- Understand Customer Requirements -- Understand Opportunities for Improvement -- Know Your Business Drivers -- Complete the Assess Phase by Prioritizing the Requirements -- Phase 2 of Decision Framework: Define -- Finalize the Design Choices for the Cloud Data Lake -- Plan Your Cloud Data Lake Project Deliverables -- Phase 3 of Decision Framework: Implement -- Phase 4 of Decision Framework: Operationalize -- Summary -- Chapter 8. Six Lessons for a Data Informed Future -- Lesson 1: Focus on the How and When, Not the If and Why, When It Comes to Cloud Data Lakes -- Lesson 2: With Great Power Comes Great Responsibility-Data Is No Exception -- Lesson 3: Customers Lead Technology, Not the Other Way Around -- Lesson 4: Change Is Inevitable, so Be Prepared -- Lesson 5: Build Empathy and Prioritize Ruthlessly -- Lesson 6: Big Impact Does Not Happen Overnight -- Summary -- Appendix A. Cloud Data Lake Decision Framework Template -- Phase 1: Assess Framework -- Phase 2: Define Framework.
- 내용주기
- 완전내용Planning the Cloud Data Lake Deliverables -- Phase 3: Implement Framework -- Index -- About the Author -- Colophon.
- 초록/해제
- 요약More organizations than ever understand the importance of data lake architectures for deriving value from their data.Building a robust, scalable, and performant data lake remains a complex proposition, however, with a buffet of tools and options that need to work together to provide a seamless end-to-end pipeline from data to insights.
- 기타형태저록
- Print version / Gopalan, RukmaniThe Cloud Data Lake. Sebastopol : O'Reilly Media, Incorporated,c2023. 9781098116583
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260202s2023 xx o 0 eng d■001EBC30291632
■003MiAaPQ
■00520260202073946.0
■006m o d |
■007cr cnu||||||||
■020 ▼a9781098116552▼q(electronic bk.)
■020 ▼z9781098116583
■035 ▼a(MiAaPQ)EBC30291632
■035 ▼a(Au-PeEL)EBL30291632
■035 ▼a(OCoLC)1355222127
■040 ▼aMiAaPQ▼beng▼erda▼epn▼cMiAaPQ▼dMiAaPQ
■0820 ▼a004.6782
■1001 ▼aGopalan, Rukmani.
■24514▼aThe Cloud Data Lake ▼bA Guide to Building Robust Cloud Data Architecture
■250 ▼a1st ed.
■264 1▼aSebastopol▼bO'Reilly Media, Incorporated▼c2023.
■264 4▼c?023.
■300 ▼a1 online resource (247 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■5050 ▼aIntro -- Copyright -- Table of Contents -- Preface -- Why I Wrote This Book -- Who Should Read This Book? -- Introducing Klodars Corporation -- Navigating the Book -- Conventions Used in This Book -- O'Reilly Online Learning -- How to Contact Us -- Acknowledgments -- Chapter 1. Big Data-Beyond the Buzz -- What Is Big Data? -- Elastic Data Infrastructure-The Challenge -- Cloud Computing Fundamentals -- Cloud Computing Terminology -- Value Proposition of the Cloud -- Cloud Data Lake Architecture -- Limitations of On-Premises Data Warehouse Solutions -- What Is a Cloud Data Lake Architecture? -- Benefits of a Cloud Data Lake Architecture -- Defining Your Cloud Data Lake Journey -- Summary -- Chapter 2. Big Data Architectures on the Cloud -- Why Klodars Corporation Moves to the Cloud -- Fundamentals of Cloud Data Lake Architectures -- A Word on Variety of Data -- Cloud Data Lake Storage -- Big Data Analytics Engines -- Cloud Data Warehouses -- Modern Data Warehouse Architecture -- Reference Architecture -- Sample Use Case for a Modern Data Warehouse Architecture -- Benefits and Challenges of Modern Data Warehouse Architecture -- Data Lakehouse Architecture -- Reference Architecture for the Data Lakehouse -- Sample Use Case for Data Lakehouse Architecture -- Benefits and Challenges of the Data Lakehouse Architecture -- Data Warehouses and Unstructured Data -- Data Mesh -- Reference Architecture -- Sample Use Case for a Data Mesh Architecture -- Challenges and Benefits of a Data Mesh Architecture -- What Is the Right Architecture for Me? -- Know Your Customers -- Know Your Business Drivers -- Consider Your Growth and Future Scenarios -- Design Considerations -- Hybrid Approaches -- Summary -- Chapter 3. Design Considerations for Your Data Lake -- Setting Up the Cloud Data Lake Infrastructure -- Identify Your Goals.
■5058 ▼aPlan Your Architecture and Deliverables -- Implement the Cloud Data Lake -- Release and Operationalize -- Organizing Data in Your Data Lake -- A Day in the Life of Data -- Data Lake Zones -- Organization Mechanisms -- Introduction to Data Governance -- Actors Involved in Data Governance -- Data Classification -- Metadata Management, Data Catalog, and Data Sharing -- Data Access Management -- Data Quality and Observability -- Data Governance at Klodars Corporation -- Data Governance Wrap-Up -- Manage Data Lake Costs -- Demystifying Data Lake Costs on the Cloud -- Data Lake Cost Strategy -- Summary -- Chapter 4. Scalable Data Lakes -- A Sneak Peek into Scalability -- What Is Scalability? -- Scale in Our Day-to-Day Life -- Scalability in Data Lake Architectures -- Internals of Data Lake Processing Systems -- Data Copy Internals -- ELT/ETL Processing Internals -- A Note on Other Interactive Queries -- Considerations for Scalable Data Lake Solutions -- Pick the Right Cloud Offerings -- Plan for Peak Capacity -- Data Formats and Job Profile -- Summary -- Chapter 5. Optimizing Cloud Data Lake Architectures for Performance -- Basics of Measuring Performance -- Goals and Metrics for Performance -- Measuring Performance -- Optimizing for Faster Performance -- Cloud Data Lake Performance -- SLAs, SLOs, and SLIs -- Example: How Klodars Corporation Managed Its SLAs, SLOs, and SLIs -- Drivers of Performance -- Performance Drivers for a Copy Job -- Performance Drivers for a Spark Job -- Optimization Principles and Techniques for Performance Tuning -- Data Formats -- Data Organization and Partitioning -- Choosing the Right Configurations on Apache Spark -- Minimize Overheads with Data Transfer -- Premium Offerings and Performance -- The Case of Bigger Virtual Machines -- The Case of Flash Storage -- Summary -- Chapter 6. Deep Dive on Data Formats.
■5058 ▼aWhy Do We Need These Open Data Formats? -- Why Do We Need to Store Tabular Data? -- Why Is It a Problem to Store Tabular Data in a Cloud Data Lake Storage? -- Delta Lake -- Why Was Delta Lake Founded? -- How Does Delta Lake Work? -- When Do You Use Delta Lake? -- Apache Iceberg -- Why Was Apache Iceberg Founded? -- How Does Apache Iceberg Work? -- When Do You Use Apache Iceberg? -- Apache Hudi -- Why Was Apache Hudi Founded? -- How Does Apache Hudi Work? -- When Do You Use Apache Hudi? -- Summary -- Chapter 7. Decision Framework for Your Architecture -- Cloud Data Lake Assessment -- Cloud Data Lake Assessment Questionnaire -- Analysis for Your Cloud Data Lake Assessment -- Starting from Scratch -- Migrating an Existing Data Lake or Data Warehouse to the Cloud -- Improving an Existing Cloud Data Lake -- Phase 1 of Decision Framework: Assess -- Understand Customer Requirements -- Understand Opportunities for Improvement -- Know Your Business Drivers -- Complete the Assess Phase by Prioritizing the Requirements -- Phase 2 of Decision Framework: Define -- Finalize the Design Choices for the Cloud Data Lake -- Plan Your Cloud Data Lake Project Deliverables -- Phase 3 of Decision Framework: Implement -- Phase 4 of Decision Framework: Operationalize -- Summary -- Chapter 8. Six Lessons for a Data Informed Future -- Lesson 1: Focus on the How and When, Not the If and Why, When It Comes to Cloud Data Lakes -- Lesson 2: With Great Power Comes Great Responsibility-Data Is No Exception -- Lesson 3: Customers Lead Technology, Not the Other Way Around -- Lesson 4: Change Is Inevitable, so Be Prepared -- Lesson 5: Build Empathy and Prioritize Ruthlessly -- Lesson 6: Big Impact Does Not Happen Overnight -- Summary -- Appendix A. Cloud Data Lake Decision Framework Template -- Phase 1: Assess Framework -- Phase 2: Define Framework.
■5058 ▼aPlanning the Cloud Data Lake Deliverables -- Phase 3: Implement Framework -- Index -- About the Author -- Colophon.
■520 ▼aMore organizations than ever understand the importance of data lake architectures for deriving value from their data.Building a robust, scalable, and performant data lake remains a complex proposition, however, with a buffet of tools and options that need to work together to provide a seamless end-to-end pipeline from data to insights.
■588 ▼aDescription based on publisher supplied metadata and other sources.
■590 ▼aElectronic reproduction. Ann Arbor, Michigan : ProQuest Ebook Central, 2026. Available via World Wide Web. Access may be limited to ProQuest Ebook Central affiliated libraries.
■655 4▼aElectronic books.
■77608▼iPrint version▼aGopalan, Rukmani▼tThe Cloud Data Lake▼dSebastopol : O'Reilly Media, Incorporated,c2023▼z9781098116583
■7972 ▼aProQuest (Firm)
■85640▼uhttps://ebookcentral.proquest.com/lib/baekseok-ebooks/detail.action?docID=30291632▼zClick to View


