Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: James Densmore

Rating No ratings yet

Data pipelines are the foundation for success in data analytics. Moving data from numerous diverse sources and transforming it to provide context is the difference between having data and actually gaining value from it. This pocket reference defines data pipelines and explains how they work in today's modern data stack. You'll learn common considerations and key decision points when implementing pipelines, such as batch versus streaming data ingestion and build versus buy. This book addresses the most common decisions made by data professionals and discusses foundational concepts that apply to open source frameworks, commercial products, and homegrown solutions. You'll learn: • What a data pipeline is and how it works • How data is moved and processed on modern data infrastructure, including cloud platforms • Common tools and products used by data engineers to build pipelines • How pipelines support analytics and reporting needs • Considerations for pipeline maintenance, testing, and alerting

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A compact, decision-oriented field guide to building analytics data pipelines: what they are, how data moves from source to warehouse, and which trade-offs (batch vs. streaming, build vs. buy, transform-early vs. transform-late) you'll face along the way. Best for aspiring and practicing data engineers who want a practical map of the modern data stack rather than a deep dive into any single tool. 【Book Arc】 - **Opening (~0%–10%)**: Defines what a data pipeline is and why it matters, framing data as a resource that only gains value once refined and delivered. Introduces the book's scope: foundational concepts that apply across homegrown code, open source frameworks, and commercial products, with all examples in Python and SQL. - **Early (~10%–25%)**: Maps the modern data infrastructure — source interfaces (databases, REST APIs, Kafka streams, cloud storage, warehouses/lakes) and data structures (JSON, CSV, flat files, semi-structured logs) — then contrasts row-based storage for OLTP with columnar storage for analytics, and lays out the ELT pattern for analysts, data scientists, and ML/data products. - **Early–Middle (~25%–45%)**: Hands-on extraction. Setting up Python and cloud file storage, then pulling data from MySQL, PostgreSQL, MongoDB, and REST APIs. Covers full vs. incremental extraction, the pitfalls of incremental loads (missed deletes, unreliable timestamps), and change data capture via binlogs, Kafka, and Debezium. - **Middle (~45%–60%)**: Loading data into destinations — configuring Amazon Redshift and Snowflake warehouses, loading into them, and using file storage as a data lake. Weighs open source frameworks against commercial alternatives. - **Late (~60%–75%)**: Transformation and modeling. Distinguishes noncontextual transformations (cleaning, type casting) from contextual ones, asks when to transform (during or after ingestion), and introduces data modeling foundations. - **Ending (~75%–100%)**: Orchestration and operations with Apache Airflow — setup, building DAGs, additional pipeline tasks, and advanced configurations — plus the maintenance, testing, and alerting concerns that keep pipelines reliable over time. 【Key Takeaways】 - **Pipelines are the value chain for data** (Opening): Raw data sitting in source systems is inert; moving and transforming it into context is what turns it into something an organization can act on. This framing justifies every later design decision. - **Ingestion starts with two questions: interface and structure** (Early): Before writing code, identify how you'll reach the data (database, API, stream, file system) and what shape it arrives in (JSON, CSV, semi-structured logs). These determine your tooling and transformation strategy. - **Storage layout should match the workload** (Early): Row-based storage suits OLTP workloads that read and write single records frequently; analytics reads large volumes infrequently and often only a few columns, which favors columnar warehouses. Choosing the wrong layout is a performance tax you pay forever. - **Incremental extraction is faster but leaks** (Early–Middle): Pulling only changed rows via a `LastUpdated` timestamp is far more efficient than full extracts, but it misses deleted rows and depends on a reliably maintained timestamp column — a common source of silent data drift. - **CDC at scale means using a platform, not hand-rolling one** (Middle): Reading MySQL binlogs or Postgres WAL directly works for small cases, but production change data capture should lean on Kafka and Debezium rather than a custom-built replication layer. - **ELT serves different consumers differently** (Early): Analysts need modeled metrics and dashboards; data scientists need more granular, sometimes raw data; ML-powered data products need training and validation data. The extract and load steps stay similar — the transform step branches. - **Build vs. buy is a recurring, context-dependent decision** (Early–Middle): Cost, engineering culture, and legal/security concerns about external vendors all push teams toward custom code, but commercial ingestion tools (Fivetran, Stitch) and frameworks (Singer) exist precisely because ingestion is repetitive work. - **Pipelines are operated, not just built** (Ending): Orchestration with Airflow, plus monitoring, testing, and alerting, is what makes a pipeline trustworthy. Delivering data once is easy; delivering it reliably, securely, and on time is the actual job. 【Reading Tips】 - **Skim the code, read the reasoning.** The Python/SQL samples are deliberately simplified with minimal error handling — treat them as starting points, not production templates. The durable value is in the decision criteria around them. - **Deep-read the trade-off sections.** Batch vs. streaming, full vs. incremental, build vs. buy, transform-during vs. transform-after — these are the chapters you'll return to when designing a real pipeline. - **Use the table of contents as a checklist.** The book is structured as a reference; if you already know ingestion, jump to transformation, modeling, and orchestration. - **Pair it with tool documentation.** Debezium, Airflow, Redshift, and Snowflake each get overview-level treatment; the book tells you *when* and *why* to use them, not every configuration detail. - **Watch for the operational chapters.** Maintenance, testing, and alerting are easy to skip but are where most real pipelines fail. 【Coverage Limits】 This guide is synthesized from stratified excerpts covering the introduction, infrastructure, ingestion, loading, transformation, and orchestration chapters; specific code details, later operational chapters, and any appendices are only partially represented. Where the excerpts are thin (e.g., detailed data modeling and advanced Airflow configurations), the guide stays at the conceptual level.
Page 6
ostgreSQL Database 63 Extracting Data from MongoDB 67 Extracting Data from a REST API 74 Streaming Data Ingestions with Kafka and Debezium 79 Chapter 5: Data...
View in text
Excerpt 2
to data ingestion tools. Of particular interest is whether the value of a commercial solution is to make it easier for data engineers to build data ingestion...
View in text
Excerpt 3
tractions will bring back the latest version of the row. In the full extract, that’s true for all rows in the table as the extrac‐ tion retrieves a full copy...
View in text
Excerpt 4
this by calling the .find() function on mongo_collection to query the documents you’re looking for. In the following exam‐ ple, you’ll grab all documents wit...
View in text
Excerpt 5
that are used for tracking marketing and ad campaigns. They are common across most platforms and organizations. Parsing URLs is possible in both SQL and Pyth...
View in text
Excerpt 6
don’t have the full timestamp stored in the daily granular Data Modeling Foundations | 137 empty between steps 3 and 4 and can be queried right away. The dow...
View in text
Excerpt 7
.sensors.external_task_sensor \ import ExternalTaskSensor from datetime import timedelta from airflow.utils.dates import days_ago dag = DAG( 'sensor_test', d...
View in text
Excerpt 8
ing a test to check for a seasonality factor in growth of a row count in the Orders table. What if you want to just warn instead of halt the pipeline? You’ll...
View in text
Tags
AI categories
DataBig DataBackend
Publisher: O'Reilly Media
Publish Year: 2021
Language: English
Pages: 206
File Format: PDF
File Size: 7.6 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…