Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Rui Pedro Machado, Helder Russa

Rating No ratings yet

With the shift from data warehouses to data lakes, data now lands in repositories before it's been transformed, enabling engineers to model raw data into clean, well-defined datasets. dbt (data build tool) helps you take data further. This practical book shows data analysts, data engineers, BI developers, and data scientists how to create a true self-service transformation platform through the use of dynamic SQL. Authors Rui Machado from Monstarlab and Hélder Russa from Jumia show you how to quickly deliver new data products by focusing more on value delivery and less on architectural and engineering aspects. If you know your business well and have the technical skills to model raw data into clean, well-defined datasets, you'll learn how to design and deliver data models without any technical influence. With this book, you'll learn: What dbt is and how a dbt project is structured How dbt fits into the data engineering and analytics worlds How to collaborate on building data models The main tools and architectures for building useful, functional data models How to fit dbt into data warehousing and laking architecture How to build tests for data transformations

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical field guide to building trustworthy, scalable data models with SQL and dbt, aimed at analysts, data engineers, BI developers, and data scientists who want to move from ad-hoc scripts to a collaborative, testable transformation platform. 【Book Arc】 - **Opening (~0%–10%)**: Frames the shift from warehouses to lakes and introduces the analytics engineer role, dbt's place in the stack, and how the book is organized across six chapters. - **Early (~10%–30%)**: Traces the evolution of data management—from Inmon and Kimball through Hadoop, Redshift, and cloud platforms—and explains why SQL and stored procedures for ETL/ELT hit their limits. - **Early–Middle (~30%–50%)**: Covers data modeling fundamentals: normalization and its trade-offs, dimensional modeling with star and snowflake schemas, Data Vault concepts, and the pain of monolithic SQL scripts. - **Middle (~50%–70%)**: Moves into SQL practice for analytics, including common table expressions, window functions, distributed SQL processing, and hands-on data manipulation with DuckDB, Polars, and FugueSQL. - **Late (~70%–90%)**: Introduces dbt itself—its design philosophy, data flow, dbt Cloud setup with BigQuery and GitHub, the Cloud UI and IDE, and the structure of a dbt project using the Jaffle Shop example. - **Ending (~90%–100%)**: Focuses on modular data models, testing transformations, and generating documentation so that models are reusable, verifiable, and understandable by others. 【Key Takeaways】 - **The analytics engineer bridges data engineering and analytics** (Early): this role designs storage, builds pipelines, and collaborates with data scientists, making reliable insights possible without owning the entire platform. - **Normalization is for OLTP, not analytics** (Middle): transactional systems optimize for fast writes and integrity, while analytics systems need read-optimized, denormalized structures that support historical analysis. - **Star and snowflake schemas trade simplicity for integrity** (Middle): star schemas simplify queries; snowflake schemas reduce redundancy but require more joins—choose based on dataset size and relationship complexity. - **Monolithic SQL scripts don't scale as a practice** (Middle): without version control, dependency management, or idempotency, large SQL files become fragile and unreusable, which is the problem dbt is designed to solve. - **dbt brings software engineering discipline to SQL** (Late): project structure, YAML configuration, sources, models, and tests turn transformations into maintainable, collaborative assets. - **Testing and documentation are first-class concerns** (Ending): building tests for transformations and generating data documentation ensures models stay trustworthy as they evolve. - **SQL remains the core language of analytics engineering** (Middle): CTEs, window functions, and distributed SQL patterns are essential skills, with DuckDB, Polars, and FugueSQL as practical tools for local and scalable manipulation. 【Reading Tips】 - Skim the historical evolution in Chapter 1 if you already know the data landscape; slow down on the analytics engineer role and its responsibilities. - Deep-read the data modeling chapter—normalization, star/snowflake, and Data Vault concepts are the conceptual foundation for everything dbt does later. - Treat the SQL chapter as a working reference: run the DuckDB and Polars examples yourself rather than just reading them. - When you reach the dbt chapters, follow the Jaffle Shop setup end-to-end; the project structure, YAML files, and tests only click when you build them. - Don't skip the testing and documentation sections—they are what separate a hobby project from a production transformation platform. 【Coverage Limits】 The excerpts cover the book's structure, historical context, data modeling concepts, SQL tooling, and dbt project setup, but do not include detailed code from later dbt chapters or advanced deployment topics. Specific chapter-level depth beyond the table of contents and sampled sections is not fully represented.
Page 6
59 Debugging and Optimizing Data Models 60 Medallion Architecture Pattern 63 Summary 66 3. SQL for Analytics. . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 2
o their specific needs, whether in the cloud or on premises. In its cloud version, dbt integrates seamlessly with leading cloud platforms, including Microsof...
View in text
Excerpt 3
the lack of 18 | Chapter 1: Analytics Engineering CHAPTER 2 Data Modeling for Analytics In today’s data-driven world, organizations rely more and more on dat...
View in text
Excerpt 4
ical changes in specific attributes. Monolith Data Modeling Until recently, the prevailing approach to data modeling revolved around the creation of extensiv...
View in text
Excerpt 5
sistency between the two versions and minimizes the risk of future data discrepancies. This technique is particularly effective when working with a well-esta...
View in text
Excerpt 6
ING filter. An optional clause closely related to GROUP BY, the HAVING filter applies conditions to the grouped data. Compared with the WHERE clause, HAVING...
View in text
Excerpt 7
window functions is ranking results within a given window, which allows ranking per group or creating relative rankings based on specific crite‐ ria. In addi...
View in text
Excerpt 8
rn. It allows us to register data tables, apply SQL queries for data preprocessing, and create and train machine learning models within a SQL context. Howeve...
View in text
Tags
AI categories
DataDatabaseSQL
ISBN: 1098142381
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 324
File Format: PDF
File Size: 10.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…