Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Marco Tranquillin, Valliappa Lakshmanan, Firat Tekiner

Rating No ratings yet

All cloud architects need to know how to build data platforms that enable businesses to make data-driven decisions and deliver enterprise-wide intelligence in a fast and efficient way. This handbook shows you how to design, build, and modernize cloud native data and machine learning platforms using AWS, Azure, Google Cloud, and multicloud tools like Snowflake and Databricks. Authors Marco Tranquillin, Valliappa Lakshmanan, and Firat Tekiner cover the entire data lifecycle from ingestion to activation in a cloud environment using real-world enterprise architectures. You’ll learn how to transform, secure, and modernize familiar solutions like data warehouses and data lakes, and you’ll be able to leverage recent AI/ML patterns to get accurate and quicker insights to drive competitive advantage. You’ll learn how to: • Design a modern and secure cloud native or hybrid data analytics and machine learning platform • Accelerate data-led innovation by consolidating enterprise data in a governed, scalable, and resilient data platform • Democratize access to enterprise data and govern how business teams extract insights and build AI/ML capabilities • Enable your business to make decisions in real time using streaming pipelines • Build an MLOps platform to move to a predictive and prescriptive analytics approach

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical field guide for cloud and data architects who need to design, modernize, and govern enterprise data and ML platforms across AWS, Azure, Google Cloud, and multicloud tools. It suits architects, data engineering leads, and technical decision-makers moving from siloed analytics toward governed, real-time, AI-driven platforms. 【Book Arc】 - **Opening (~0%–10%)**: Frames why traditional, siloed data ecosystems fail and introduces the modern data lifecycle—ingestion, storage, processing, analysis, and activation—using a water-pipes analogy to orient readers. - **Early (~10%–30%)**: Traces the evolution from data warehouses to Hadoop-era data lakes and their governance pitfalls, then argues for a unified analytics platform and introduces data mesh as a way to decentralize ownership while retaining accountability. - **Early–Middle (~30%–45%)**: Covers strategic and organizational foundations—buy-versus-build decisions for AI, prebuilt ML building blocks, real-time analytics needs, and the people-process-technology balance anchored by a Center of Excellence. - **Middle (~45%–60%)**: Gets concrete about storage and data choices: when to use SQL-optimized warehouses versus open table formats like Iceberg or Delta Lake on object storage, and how to handle structured, semistructured, and unstructured data. - **Late (~60%–85%)**: Moves into platform architecture—data lakes, warehouses, and lakehouses as the three common patterns, plus ML platform activities such as labeling, development, training, deployment, and automation. - **Ending (~85%–100%)**: Addresses operational maturity: MLOps, training-serving skew, continuous evaluation, responsible AI principles, and evaluation processes like PoCs and RFPs. 【Key Takeaways】 - **Siloed data is silenced data** (Early): Task-specific stores fragment enterprise intelligence; a unified platform is the prerequisite for broad, governed insight. - **The data lifecycle is the organizing spine** (Early): Ingestion → storage → processing → analysis → activation gives architects a repeatable mental model for platform design. - **Lakehouses resolve the lake-versus-warehouse tension** (Middle): They can be reached by evolving from either a lake or a warehouse, and the book helps you choose between those two paths. - **Data mesh federates ownership without losing governance** (Early): Domain teams own data as a product, while cloud IAM, encryption, and masking enforce org-wide security centrally. - **Storage choice is a deliberate trade-off** (Middle): SQL-optimized warehouses favor performance; open formats like Parquet with Iceberg or Delta Lake favor cost and non-SQL workloads like ML. - **Buy versus build is a strategic AI decision** (Middle): Prebuilt models and building blocks let non-experts adopt AI quickly, while custom models deliver differentiation where it matters. - **Real-time integration is where ML value lives** (Middle): The platform must ingest, process, and serve data fast enough for inference before the customer context changes. - **MLOps closes the loop from predictive to prescriptive** (Late): Automation, orchestration pipelines, and continuous evaluation combat training-serving skew and keep models useful. 【Reading Tips】 - Deep-read the early chapters on the data lifecycle and data mesh—they anchor every later architectural decision. - Skim the cloud-provider comparisons if you already work in one ecosystem; focus instead on the trade-off reasoning that transfers across AWS, Azure, and Google Cloud. - Treat the storage and lakehouse chapters as reference material to revisit when making concrete platform choices. - Pay attention to the organizational and CoE discussion; the book repeatedly stresses that data transformation is cultural, not just technological. - Use the MLOps and responsible AI sections as a checklist when moving from experimentation to production. 【Coverage Limits】 The excerpts cover the book's structure, lifecycle framing, architecture patterns, and strategic themes, but do not include detailed code, full chapter contents, or specific implementation walkthroughs. Some late-chapter material is only visible through table-of-contents entries.
Page 7
ry Overview. . . . . . . . . . . . . . . . . . . . . . . 1 The Data Lifecycle 2 The Journey to Wisdom 2 Water Pipes Analogy 3 Collect 4 Store 5 Process/Trans...
View in text
Excerpt 2
data engineers collect and store data in an analytics store. The stored data is then processed using a variety of tools. If the tools involve programming, th...
View in text
Excerpt 3
mation of data across the organization without clear owner‐ ship responsibilities over the newly created data. In the data mesh, the authoritative data sourc...
View in text
Excerpt 4
aged by different teams but as part of the same larger DWH. Another option is to store the structured or semistructured data in an open format such as Parque...
View in text
Excerpt 5
se with new cloud DWHs. The role of ingestion is now simply to bring data close to the cloud, and the transformation and processing part moves back to the cl...
View in text
Excerpt 6
p 3: low business value, low effort to migrate (Priority 2) • Group 4: low business value, high effort to migrate (Priority 3) In Figure 4-2 you can see an e...
View in text
Excerpt 7
he preceding description of offline versus online transfer. Third-party commercial off-the-shelf (COTS) solutions These could provide more features like netw...
View in text
Excerpt 8
re potentially ingesting petabytes of data every single day. Additionally, to guarantee performance it is crucial to select a compression algorithm that is f...
View in text
Tags
AI categories
Cloud NativeArtificial IntelligenceBig Data
ISBN: 1098151615
Publisher: O'Reilly Media
Publish Year: 2023
Language: English
Pages: 362
File Format: PDF
File Size: 7.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…