Data lakes and warehouses have become increasingly fragile, costly, and difficult to maintain as data gets bigger and moves faster. Data meshes can help your organization decentralize data, giving ownership back to the engineers who produced it. This book provides a concise yet comprehensive overview of data mesh patterns for streaming and real-time data services.
Authors Hubert Dulay and Stephen Mooney examine the vast differences between streaming and batch data meshes. Data engineers, architects, data product owners, and those in DevOps and MLOps roles will learn steps for implementing a streaming data mesh, from defining a data domain to building a good data product. Through the course of the book, you'll create a complete self-service data platform and devise a data governance system that enables your mesh to work seamlessly.
With this book, you will
• Design a streaming data mesh using Kafka
• Learn how to identify a domain
• Build your first data product using self-service tools
• Apply data governance to the data products you create
• Learn the differences between synchronous and asynchronous data services
• Implement self-services that support decentralized data
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Streaming Data Mesh: A Model for Optimizing Real-Time Data Services
## 【One-Line Pitch】
A practical guide for data engineers, architects, and product owners who want to move beyond fragile batch-oriented data lakes and build decentralized, real-time data platforms using streaming data mesh patterns with Kafka and related technologies. If you're wrestling with data governance, domain ownership, or the gap between batch ETL thinking and streaming reality, this book gives you a concrete path forward.
## 【Book Arc】
- **Opening (~0%–9%)**: Introduces the core problem—data lakes and warehouses are becoming too costly and fragile as data scales—and positions the data mesh as a decentralized alternative. The authors set expectations: this is a streaming-focused take on Zhamak Dehghani's data mesh pillars, not a general data architecture book.
- **Early (~9%–25%)**: Establishes the foundational concepts: the four pillars of data mesh (domain ownership, data as a product, federated computational governance, self-service platform), and contrasts streaming vs. batch approaches. Covers key architectural patterns including reverse ETL, Lambda, and Kappa architectures, plus the CAP theorem's implications for real-time data stores.
- **Early–Middle (~25%–38%)**: Dives into domain identification—the hardest and most important step. Covers practical patterns like active-active vs. active-passive disaster recovery across cloud providers, data replication for cross-region SLAs, and introduces Domain-Driven Design (DDD) concepts like the ubiquitous language and bounded contexts as tools for defining domains.
- **Middle (~38%–47%)**: Focuses on domain roles and operational concerns. Introduces the data product engineer vs. data product owner split, and tackles the messy reality of charge-backs—how to allocate costs fairly across domains using usage-based, cost-splitting, or data-product-level approaches.
- **Middle (~47% onward)**: Shifts to building data products. Uses real-world examples like Apache Spark wrappers (Big Data Integrator, Sparknado, Apache Envelope) to show how self-service tools make complex data engineering accessible to generalist engineers, then lays out data product requirements for a smooth consumer experience.
## 【Key Takeaways】
- **Data mesh is not full decentralization** (Early): While data lives in domains, the "mesh" itself—governance, security, interoperability—must be centrally coordinated. Federated computational governance is what makes the mesh work, not an afterthought.
- **Streaming changes the governance conversation** (Early): Batch-oriented governance assumes you can inspect data at rest. Streaming requires thinking about schemas, serialization, and tokenization/encryption as first-class concerns, because data is always in motion.
- **Reverse ETL is a symptom of architectural inversion** (Early): Moving data from the analytical plane back to operational systems is an anti-pattern that puts data warehouses in the hot path of tier-1 applications. A streaming data mesh should eliminate this need by serving data directly from source domains.
- **CAP theorem forces a choice in streaming** (Early): For real-time views, consistency matters more than availability. The Lambda architecture's dual batch/streaming branches create inconsistency problems—the Kappa architecture with tiered storage (Kafka, Redpanda, Pulsar) offers a cleaner alternative.
- **Domain identification is the critical first step** (Middle): Use DDD's ubiquitous language and bounded contexts to define domains. The goal is a shared vernacular so domain experts and engineers can communicate issues across the full spectrum from business logic to technical exceptions.
- **Cross-region SLAs require data replication, not stretched pipelines** (Middle): When a consuming domain is far from the source, replicate data products locally rather than trying to maintain a long-distance pipeline. The replication should be transparent to consumers and keep data in motion in real time.
- **Charge-backs need a cost model that fits your governance style** (Middle): Usage-based charge-backs are precise but complex; cost-splitting is simple but can lead to over- or under-billing; data-product-level charge-backs require monitoring but give the clearest picture. Choose based on your domain structure and consumer relationships.
- **Self-service tools are the "easy buttons" that make mesh adoption possible** (Middle): The goal is not to reduce data engineers' workload but to enable domain engineers to solve complex data problems independently. Wrappers like Sparknado and Apache Envelope show how to make powerful tools accessible without dumbing them down.
## 【Reading Tips】
- **Skim the preface and Chapter 1** if you're already familiar with data mesh concepts—the authors themselves point to Dehghani's book for deeper pillar-level detail. Focus instead on where they diverge: streaming-specific implications.
- **Deep-read Chapter 3 (domain identification)**—this is where the book earns its keep. The DDD material and the active-active vs. active-passive DR patterns are immediately actionable and rarely covered well in other data mesh resources.
- **Pay attention to the tables** (e.g., Spark wrapper projects, charge-back models). They compress a lot of practical comparison into scannable form—use them as decision aids rather than reading every surrounding paragraph.
- **Don't get bogged down in the CAP theorem and Lambda/Kappa architecture review** if you've seen it before—it's standard distributed systems material. The novel content is in how these constraints play out in a mesh context.
- **Read Chapter 4's data product requirements with your own use case in mind**—the authors keep scope tight to streaming products, so use their checklist as a starting point for your own product definition rather than a universal spec.
## 【Coverage Limits】
The excerpts cover roughly the first half of the book (through ~47%), focusing on foundations, domain identification, roles, and early data product definition. Later chapters on self-service platform implementation, data governance in depth, and the full streaming mesh build-out are not covered in this guide.
##
h skills that are accessible to a more generalist engineer. When building a data mesh, it is necessary to enable existing engineers in a domain to perform th...
orei—.e., Amazon S3, Google Cloud Storage, Azure Blob Store. The top tier is storage within the Kafka brokers. When applications request data that is outside...
p features deliver reports in a business intelligence tool. Sparknado—an Apache Spark wrapper that used Airflow engineers build Spark applications to move da...
data may not provide the numbers in that format. We need to make sure all the formatting standards are applied to the data before it is published as a stream...
d OpenID that are supported by streaming platforms. We will not go over each implementation in detail because it is beyond the scope of this book (it would b...
ssumption that it is computational. We call the overarching data governance, federal data governance. Likewise for the domains: domain data governance. In th...
ngineering simpler for domains. These easy buttons hide the hard parts behind a streaming data mesh. It is the responsibility of the central team to implemen...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Streaming Data Mesh A Model for Optimizing Real-Time Data Services (Hubert Dulay, Stephen Mooney)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Streaming Data Mesh A Model for Optimizing Real-Time Data Services (Hubert Dulay, Stephen Mooney)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment