Learn how to build end-to-end scalable machine learning solutions with Apache Spark. With this practical guide, author Adi Polak introduces data and ML practitioners to creative solutions that supersede today's traditional methods. You'll learn a more holistic approach that takes you beyond specific requirements and organizational goals—allowing data and ML practitioners to collaborate and understand each other better.
Scaling Machine Learning with Spark examines several technologies for building end-to-end distributed ML workflows based on the Apache Spark ecosystem with Spark MLlib, MLflow, TensorFlow, and PyTorch. If you're a data scientist who works with machine learning, this book shows you when and why to use each technology.
You will:
• Explore machine learning, including distributed computing concepts and terminology
• Manage the ML lifecycle with MLflow
• Ingest data and perform basic preprocessing with Spark
• Explore feature engineering, and use Spark to extract features
• Train a model with MLlib and build a pipeline to reproduce it
• Build a data system to combine the power of Spark with deep learning
• Get a step-by-step example of working with distributed TensorFlow
• Use PyTorch to scale machine learning and its internal architecture
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical guide for data scientists and ML engineers who want to move beyond single-machine experiments and build scalable, end-to-end machine learning workflows using Apache Spark, MLlib, MLflow, TensorFlow, and PyTorch—covering everything from distributed computing fundamentals to production deployment patterns.
【Book Arc】
- **Opening (~0%–9%)**: Introduces the book's scope and roadmap—covering distributed ML concepts, the ML workflow stages, and the ecosystem of Spark, MLlib, MLflow, TensorFlow, and PyTorch. Sets expectations for hands-on tutorials and production-focused content.
- **Early (~9%–25%)**: Dives into distributed computing terminology and architecture—covering synchronous/asynchronous communication, barriers, allreduce/allgather/broadcast, and ensemble methods. Establishes the mental model needed to understand why distributed ML differs from local training.
- **Early–Middle (~25%–38%)**: Addresses key challenges in distributed ML—privacy concerns, portability issues with GPUs and cloud migration, and the trade-offs of different distributed topologies. Concludes with guidance on setting up a local learning environment for the book's code samples.
- **Middle (~38%–47%)**: Introduces Apache Spark and PySpark fundamentals—the distributed architecture, RDDs, DataFrame immutability, and functional programming paradigms (lambda functions). Explains why immutability and reproducibility are critical for scalable ML, and contrasts Spark DataFrames with pandas and scikit-learn workflows.
- **Middle (~47%–end of sample)**: Transitions to MLflow for managing the ML lifecycle—tracking experiments, logging parameters, organizing runs, and supporting individual users and data science teams. Sets up the foundation for later chapters on feature engineering, model training with MLlib, deep learning integration, and deployment patterns.
【Key Takeaways】
- **Distributed ML requires a new vocabulary** (Early): Understanding concepts like barriers, synchronous/asynchronous communication, and collective operations (allreduce, allgather) is essential before writing any distributed code—these determine how your cluster behaves under load.
- **Ensemble methods are a core distributed ML technique** (Early): Combining multiple algorithms reduces bias and variance, but requires clear model aggregation strategies at prediction time—a pattern that maps naturally to distributed computing.
- **Privacy and portability are unsolved distributed ML problems** (Early–Middle): There's no one-size-fits-all privacy solution, and hardware diversity (GPUs, cloud-native features) complicates workload portability—plan for these constraints early in your architecture.
- **Spark's immutability is a feature, not a limitation** (Middle): Every DataFrame operation is reproducible and trackable, which is critical for scaling ML experiments—but Spark prunes unused operations, so you must explicitly save references to preserve your work.
- **Functional programming underpins PySpark** (Middle): Lambda functions and stateless operations are the backbone of Spark's scalability—mastering these patterns makes your code more reliable and maintainable across distributed systems.
- **MLflow is the lifecycle manager for distributed ML** (Middle): It supports individual researchers and teams alike, enabling experiment tracking, parameter logging, and artifact sharing—use `log_params` for batch logging and tags for run organization.
- **The book's structure mirrors a real ML workflow** (Opening): From data ingestion and preprocessing → feature engineering → model training → deep learning → deployment, each chapter builds on the previous, so follow the sequence for best results.
【Reading Tips】
- **Skim Chapter 1's terminology if you're already familiar with distributed systems**—but don't skip the ensemble methods section, as it's referenced later in model training discussions.
- **Deep-read the Spark fundamentals chapter (Ch. 2)**—the immutability and functional programming sections are the most conceptually dense and will save you debugging time in later tutorials.
- **Set up your local environment twice as the author advises**—once for Chapters 2–6 (Spark/MLlib) and once for Chapters 7–10 (deep learning frameworks)—to avoid dependency conflicts.
- **Use the GitHub repository for code samples**—the book's tutorials are meant to be run, not just read; hands-on practice is essential for internalizing the concepts.
- **Pay special attention to the MLflow chapter**—it's the glue that ties together experiment tracking, model reproducibility, and deployment, and it's easy to underestimate its importance.
【Coverage Limits】
The excerpts cover the book's opening through the MLflow chapter (~47% of the book). Later chapters on feature engineering with Spark, MLlib pipeline training, distributed TensorFlow, PyTorch internals, and deployment patterns (batch, model-in-service, model-as-a-service) are mentioned in the table of contents but not detailed in this guide.
Page 6
ns are also available for most titles (https://oreilly.com). For more information, contact our corporate/institu‐ tional sales department: 800-998-9938 or co...
oday are driven by machine learning, using machine learning models to answer questions such as: How can my application automatically adapt itself to the cust...
gnificant sources of confusion when approaching distributed machine learning is the lack of a clear understanding of what precisely is distributed across the...
with anonymous functions. These are functions that are exe‐ cuted without state yet are not named. The idea originated in mathematics, with lambda calculus,...
high level, the two types of objects are as follows: Vector A vector object represents a numeric vector. You can think of it as an array, just like in Python...
hat to do: Do you fill in the missing values? Drop the col‐ umns with missing values altogether? Try using a different algorithm? The decisions you make can...
M errors. What is an out-of-memory exception? Good question. OOM errors are common in Java; they happen when an in-memory object requires more RAM than is av...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Scaling Machine Learning with Spark Distributed ML with MLlib, TensorFlow, and PyTorch (Adi Polak)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Scaling Machine Learning with Spark Distributed ML with MLlib, TensorFlow, and PyTorch (Adi Polak)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment