Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Jim Dowling

Rating No ratings yet

Get up to speed on a new unified approach to building machine learning (ML) systems with a feature store. Using this practical book, data scientists and ML engineers will learn in detail how to develop and operate batch, real-time, and agentic ML systems. Author Jim Dowling introduces fundamental principles and practices for developing, testing, and operating ML and AI systems at scale.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Building Machine Learning Systems with a Feature Store ## 【One-Line Pitch】 A practical, hands-on guide for data scientists and ML engineers who want to master the unified approach of building batch, real-time, and LLM-powered ML systems using feature stores—with Hopsworks as the primary reference implementation. If you've struggled with offline/online skew, feature reuse, or production ML pipelines, this book shows you a coherent way forward. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces the core problem—ML systems need consistent, reusable features across training and inference—and lays out the book's structure: batch, real-time, and agentic (LLM) systems. The author positions Hopsworks as the reference feature store (he's a developer) while noting examples work with alternatives like Feast. - **Early (~10%–23%)**: Establishes the fundamental architecture: feature groups (mutable feature storage), feature views (query interfaces), and the taxonomy of transformations (reusable, model-specific, real-time). Walks through building a first end-to-end air quality forecasting system with XGBoost, a dashboard, and natural-language querying via LLM function calling. - **Early–Middle (~23%–39%)**: Dives deep into feature store design—point-in-time correctness, ASOF joins to prevent data leakage, slow-changing dimension (SCD) types, and data mesh patterns for organizing projects. Uses the recurring credit card fraud example to show how feature stores enrich raw transaction data with historical context. - **Middle (~39%–48%)**: Covers data transformation frameworks (Pandas, Polars, Spark, Flink) and feature engineering techniques—joins, categorical encoders, schema management, and mixed-mode UDFs that run as vectorized Pandas operations in training but low-latency Python functions in online inference. - **Late (~48%–end)**: Moves into production concerns: testing strategies (offline pytest during development vs. online tests in production), training pipeline design, model deployment with A/B testing, and operating streaming feature pipelines with windowed and rolling aggregations. ## 【Key Takeaways】 - **Feature stores solve the offline/online skew problem** (Early): By storing features in feature groups with point-in-time correctness, you ensure training and inference use identical feature definitions—eliminating the classic "works in training, breaks in production" failure mode. - **The three-type transformation taxonomy is the mental model** (Early): Reusable features (computed once, used everywhere), model-specific features (computed per model), and real-time features (computed at inference) each belong in different pipeline stages—knowing which is which prevents costly mistakes. - **ASOF joins are the backbone of time-series ML** (Early): Temporal joins with ASOF LEFT JOIN conditions prevent future data leakage while preserving label rows, ensuring your training data is point-in-time correct across multiple feature tables. - **Data mesh beats central data teams for feature ownership** (Early): Distributing feature groups across projects with read-only sharing (as in the credit card fraud example) lets domain experts own their data while enabling cross-team reuse—a pattern that scales better than a single monolithic project. - **Mixed-mode UDFs give you both performance and latency** (Middle): Writing transformation functions that execute as vectorized Pandas UDFs in training but as Python UDFs in online inference (like the `transaction_amount_deviation` example) means you don't sacrifice correctness for speed. - **Explicit schemas are non-negotiable in production** (Middle): Pandas and PySpark both infer types incorrectly (datetimes become strings, etc.); specifying schemas explicitly for feature groups prevents silent data corruption that's expensive to fix later. - **LLM function calling turns dashboards into conversational interfaces** (Early): By passing verbose, well-documented function declarations to an LLM, you can let users query forecasts in natural language ("What will air quality be like Tuesday?") without the model actually executing code—you parse its response and call the function yourself. ## 【Reading Tips】 - **Skim the Hopsworks-specific API details** (Chapters 4–5): The concepts (feature groups, feature views, ASOF joins) are portable; the exact method calls (`fs.get_feature_group()`, `fv.training_data()`) are implementation details you can look up later. - **Deep-read the transformation taxonomy and mixed-mode UDF sections** (Chapters 2, 7): These are the conceptual core that will change how you architect pipelines—worth rereading and taking notes on. - **Build the air quality project in Chapter 3**: It's the book's "hello world" that ties everything together—feature groups, training, batch inference, and LLM querying. Don't skip it even if you're not interested in weather data. - **Watch for the credit card fraud example as a recurring thread**: It appears across chapters to illustrate different concepts (data enrichment, data mesh, joins, UDFs). Following it through gives you a complete picture of a realistic production system. - **The testing chapter (Chapter 7) is more important than it looks**: The offline/online test distinction (pytest during development, not in production) is a practical insight that will save you debugging headaches. ## 【Coverage Limits】 Excerpts cover roughly the first half of the book (through Chapter 7 of 10+). The guide does not cover the later chapters on streaming feature pipelines (Chapter 9), training with unstructured data, or the full agentic/LLM system design sections—though the LLM function-calling pattern from Chapter 3 is included. ##
Page 15
escribes how to design and schedule batch feature pipelines. Chapter 9 describes how to design and operate streaming feature pipelines, including windowed ag...
View in text
Excerpt 2
registry A batch inference pipeline to download the model, make predictions on new feature data, and read from the feature store to produce air quality forec...
View in text
Excerpt 3
the data files. Instead, you can create a new feature group with a different name but with the same primary key and event time as the original feature group...
View in text
Excerpt 4
oding the encoder labels with a value target/label variable between 0 and n_classes-1 For features with a very large number of categories, feature hashing (t...
View in text
Excerpt 5
down to the online feature store that executes them as SQL. These provide insights into recent spikes or drops in activity (such as anomalous fraud activity...
View in text
Excerpt 6
acity. If it’s significantly lower, you may be bottlenecked by reading training data from object storage. One fix is to pre-copy training data from object st...
View in text
Excerpt 7
at clients can execute should also be deterministic, making them predictable and side-effect-free. MCP also supports resources, which are functions that retu...
View in text
Excerpt 8
at observability in agentic AI systems, where logging is a building block for error analysis and evals, both of which are key techniques in building reliable...
View in text
Tags
AI categories
AIBackendData
machine learning
Publish Year: 2025
Language: English
File Format: PDF
File Size: 13.6 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…