Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Bartosz Konieczny

Rating No ratings yet

Data projects are an intrinsic part of an organization's technical ecosystem, but data engineers in many companies continue to work on problems that others have already solved. This hands-on guide shows you how to provide valuable data by focusing on various aspects of data engineering, including data ingestion, data quality, idempotency, and more. Author Bartosz Konieczny guides you through the process of building reliable end-to-end data engineering projects, from data ingestion to data observability, focusing on data engineering design patterns that solve common business problems in a secure and storage-optimized manner. Each pattern includes a user-facing description of the problem, solutions, and consequences that place the pattern into the context of real-life scenarios. Throughout this journey, you'll use open source data tools and public cloud services to apply each pattern. You'll learn: Challenges data engineers face and their impact on data systems How these challenges relate to data system components Useful applications of data engineering patterns How to identify and fix issues with your current data components Technology-agnostic solutions to new and existing data projects, with open source implementation examples Bartosz Konieczny is a freelance data engineer who's been coding since 2010. He's held various senior hands-on positions that allowed him to work on many data engineering problems in batch and stream processing.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Engineering Design Patterns: Recipes for Solving the Most Common Data Engineering Problems ## 【One-Line Pitch】 A practical pattern catalog for working data engineers who want to stop reinventing the wheel—covering ingestion, error management, and data quality with technology-agnostic solutions and open-source examples. Read this if you have at least six months of commercial data engineering experience and want battle-tested recipes for recurring problems. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces the concept of data engineering design patterns as a shared language and reusable solutions, similar to cooking recipes. Sets prerequisites (ETL/ELT, data warehousing, cloud basics) and points readers to *Fundamentals of Data Engineering* for background. Code examples are available on GitHub with Docker Compose setups. - **Early (~9%–28%)**: Dives into data ingestion patterns—Full Loader for slowly changing reference data, Incremental Loader for partition-based ingestion with Airflow sensors, Replication for cross-environment data copying, and Compactor for solving the small files problem in lakehouses. - **Early (~28%–34%)**: Covers data readiness and event-driven ingestion, including the Readiness Marker pattern for signaling data availability and the External Trigger pattern for pull- and push-based event detection. - **Middle (~34%–47%)**: Moves to error management—the Dead-Letter pattern for handling unprocessable records (poison pills), the Windowed Deduplicator for at-least-once delivery semantics, and the Dynamic Late Data Integrator for handling late-arriving data without full partition backfills. - **Late (~47%+)**: Discusses consequences of patterns—concurrency issues with dynamic late data integration, limitations of declarative languages like SQL, and the overhead of making stateless streaming jobs stateful for filter interception. ## 【Key Takeaways】 - **Design patterns give data teams a shared vocabulary** (Early): Instead of describing solutions ad hoc, patterns like Dead-Letter or Compactor let engineers communicate complex ideas quickly. This saves time and improves collaboration with teammates and new hires. - **Full Loader is the simplest ingestion pattern but has hidden pitfalls** (Early): It works best for small, slowly evolving datasets without update timestamps. The two-step extract-load approach is ideal for homogeneous data stores but requires careful handling to avoid data quality issues from transformations. - **Incremental Loader requires partition-aware orchestration** (Early): Using Airflow sensors (FileSensor, partition sensors for AWS Glue, BigQuery, Databricks) prevents loading partial data. The DAG pattern—sensor → trigger → job sensor—ensures data readiness before processing. - **Replication should be as dumb as possible** (Early): Keep copy jobs simple—use native database copy commands or raw text APIs instead of JSON I/O to avoid unintended type conversions. Distributed frameworks can introduce file count and naming inconsistencies. - **The small files problem is still alive in modern lakehouses** (Early): Metadata listing can consume 70% of execution time. The Compactor pattern solves this via table format commands—Iceberg's rewrite data files, Delta Lake's OPTIMIZE, or Hudi's merge-on-read configuration. - **Dead-Letter pattern keeps pipelines running but adds complexity** (Middle): Separating valid from invalid records (poison pills) prevents job failures, but implementing it in declarative SQL is verbose and hard to maintain. Accept the extra code complexity as a trade-off. - **Deduplication requires bounded data windows** (Middle): For streaming jobs, define time limits for deduplication state. The Windowed Deduplicator pattern ensures each record processes once, but you must handle the trade-off between state size and accuracy. - **Dynamic late data integration fixes lookback windows but introduces concurrency risks** (Middle): Using state tables to track last processed times avoids reprocessing full partitions, but parallel job executions can cause duplicate runs. Declarative languages and stateless streaming jobs struggle with this pattern. ## 【Reading Tips】 - **Skim Chapter 1** if you're already familiar with pattern concepts—the recipe analogy and prerequisites are useful but not the core value. - **Deep-read the ingestion chapter** (Early section) for the Full Loader, Incremental Loader, and Compactor patterns—these are the most universally applicable and have concrete Airflow/Spark examples. - **Pay attention to the "Consequences" sections**—the author is honest about trade-offs (complexity, concurrency, declarative language limitations), which is where the real learning happens. - **Use the GitHub repo alongside the book**—each chapter has working examples with Docker Compose, so run them to see patterns in action rather than just reading code snippets. - **Skip the SQL implementation examples** if you're not a Spark SQL user—the programmatic API examples are more instructive for understanding pattern logic. ## 【Coverage Limits】 Excerpts cover ingestion and error management patterns in depth, but do not cover later chapters on data quality, data observability, or security patterns mentioned in the blurb. The guide focuses on the first half of the book (~0–47%). ##
Page 7
. . . ix 1. Introducing Data Engineering Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 What Are Design Patterns? 1 Y...
View in text
Excerpt 2
a few times a week. It’s also a very slowly evolving entity with the total number of rows not exceeding one million. Unfortunately, the data pro‐ vider doesn...
View in text
Excerpt 3
can be configured as a merge-on-read (MoR) table where the dataset is written in columnar format and any subsequent changes are written in row 4 There’s a de...
View in text
Excerpt 4
k window, but it also has its own shortcomings. Concurrency. If your pipeline supports concurrent executions, dynamic late data inte‐ gration may generate du...
View in text
Excerpt 5
u should slightly adapt the implementation to your use case. In that scenario, the pipeline will be composed of the steps in Figure 4-6. Figure 4-6. The Stat...
View in text
Excerpt 6
ched dataset or vice versa. To mitigate this issue, dynamic joins are often completed with additional time conditions. Defining these time conditions implies...
View in text
Excerpt 7
e state in case of failure or restart, the job synchronizes the state regularly to a more resilient fault tolerance storage. The data processing logic can re...
View in text
Excerpt 8
quence | 173 Figure 6-4. Confusing Unaligned Fan-In example To mitigate this issue it is always better to check if the data orchestration tool pro‐ vides cus...
View in text
Tags
AI categories
data engineeringDataProgramming
ISBN: 1098165810
Publisher: O'Reilly Media
Publish Year: 2025
Language: English
Pages: 375
File Format: PDF
File Size: 7.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…