Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Jean-Georges Perrin & Eric Broda

Rating No ratings yet

As data continues to grow and become more complex, organizations seek innovative solutions to manage their data effectively. Data Mesh is one solution that provides a new approach to managing data in complex organizations. This practical guide offers step-by-step guidance on how to implement data mesh in your organization. Authors Jean-Georges Perrin and Eric Broda focus on the key components of data mesh and provide practical advice supported by code.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Implementing Data Mesh — Reading Guide ## 【One-Line Pitch】 A practical, code-supported guide for architects, data engineers, and technical leaders who want to move beyond centralized data platforms and build a decentralized, domain-owned data ecosystem. If you've heard the buzzwords but need concrete patterns, contracts, and operating models, this book bridges theory and implementation. ## 【Book Arc】 - **Opening (~0%–11%)**: Sets the stage with the core promise of Data Mesh—decentralized ownership, self-serve infrastructure, and federated governance—and outlines the full journey from principles to practice, including a chapter roadmap covering data products, experience planes, and GenAI's role. - **Early (~11%–25%)**: Introduces the four foundational principles (domain ownership, data as a product, self-serve platform, federated computational governance) and grounds them in a detailed case study—Climate Quantum Inc.—showing how a global climate data mesh handles massive, diverse, and constantly changing data. - **Early (~25%–36%)**: Dives into the anatomy of data products: artifacts (datasets, programs, models), tags, policies, access rights, licenses, and endpoints. Also covers ingestion strategies—streaming, bulk, and pipeline-based (Airflow, dbt)—and introduces data contracts with semantic versioning. - **Middle (~36%–50%)**: Explores data quality beyond traditional metrics, introducing the seven dimensions (CACTAR plus alignment with EDM Council), service-level indicators like latency, retention, and end-of-life, and the concept of human lineage alongside data lineage. Then shifts to the three experience planes: infrastructure, data product, and mesh. - **Middle (~50%–end of excerpts)**: Focuses on meshing data products together—registration via data contracts (ODCS/ODPS), the data marketplace as the "entrance door," feedback loops, and the sidecar pattern that isolates implementation while keeping APIs consistent across all data products. ## 【Key Takeaways】 - **Data Mesh is a socio-technical shift, not just a technology stack** (Early): The four principles—domain ownership, data as a product, self-serve platform, and federated computational governance—require changing team topology and operating models, not just deploying new tools. - **Data products are more than datasets** (Early): They bundle artifacts (including programs and models), tags, policies, access rights, licenses, and endpoints into a complete, ready-to-use package. This makes them "more than just a data store." - **Federated governance replaces top-down policing with certification** (Early): Domain product owners (DPOs) closest to the data manage compliance and publish certification statuses, avoiding the bottlenecks of centralized governance while maintaining enterprise-wide standards. - **Data contracts are the backbone of interoperability** (Middle): Semantic versioning (patch/minor/major) ensures backward compatibility, and contracts track both data lineage and human lineage—who owned the product and when—creating a rich single source of truth. - **Data quality is necessary but insufficient** (Middle): The book expands beyond quality to include service-level indicators like latency, retention, and end-of-life dates, arguing that trust requires understanding the full lifecycle of data, not just its accuracy. - **The sidecar pattern standardizes data product implementation** (Middle): All data products share the same APIs for observability, discovery, and control, with implementation isolated in a "running library" sidecar—so you never learn a new API when switching between products. - **The mesh experience plane unlocks new capabilities** (Middle): Combining data products enables eight new services, including the data marketplace (the essential "entrance door" for discovery), feedback loops, and knowledge graph representations of lineage. ## 【Reading Tips】 - **Skim the foreword and early principle chapters** (~0%–11%) if you already know the Data Mesh theory—the real value starts with the Climate Quantum case study, which makes abstract principles concrete. - **Deep-read the data contract chapters** (~36%–46%): The semantic versioning examples and human lineage scenarios are practical and immediately applicable. This is where the book earns its "code-supported" promise. - **Pay attention to the sidecar pattern discussion** (~46%–50%): It's a genuinely useful architectural insight that you can apply even outside a full Data Mesh implementation. - **Don't skip the ingestion methods section** (~29%–36%): The streaming vs. bulk vs. pipeline comparison is a practical decision framework you'll reference when building your first data products. - **The excerpts don't cover the GenAI chapters or the operating model/roadmap sections**—if those are your primary interest, you'll need to read the full book. ## 【Coverage Limits】 This guide synthesizes the first ~50% of the book. The excerpts do not cover the later chapters on generative AI applications, team topology, operating models, or the implementation roadmap—all of which are promised in the table of contents but not included in the source material. ##
Page 18
what and open standards like the ones promoted by the Bitol project to streamline development and operations. Chapter 7, “Aligning with the Experience Planes...
View in text
Excerpt 2
ther it’s scaling to accommodate growth or integrating with new systems and applications. Underpinning all these attributes is the role of comprehensive docu...
View in text
Excerpt 3
he ingestion, transformation, and loading of data in a more controlled, systematic manner. Pipelines are particularly useful when the data-ingestion process ...
View in text
Excerpt 4
tractId: af12347d-b730-48e5-a369-33a2c70fd version: 2.0.0 Consider the group of input data sources and data output ports as a single set. The data produc...
View in text
Excerpt 5
aps most important, these challenges foster a costly, slow, bureaucratic, and complex approach that stands in stark contrast to business’s demand for speed a...
View in text
Excerpt 6
prises, as the costs and time involved in training LLMs are extensive, measured in the tens and hundreds of millions of dollars and many months of time. Toda...
View in text
Excerpt 7
quired to build, secure, and deploy data products at scale. To summarize, the factory stream emphasizes an agile, iterative approach to developing and scalin...
View in text
Excerpt 8
ce plane, Capabilities of the Data Product Experience Plane enterprise grade data products, Defining an Enterprise- Grade Data Product infrastructure experie...
View in text
Tags
AI categories
DataBackendCloud Native
Publish Year: 2024
Language: English
File Format: PDF
File Size: 9.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…