Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Gourav Gupta, Dr. Manish Gupta, Dr. Inder Singh Gupta

Rating No ratings yet

Who this book is for This book is valuable for data engineers, machine learning engineers, data scientists, data architects, business analysts, and technical consultants worldwide. It would be beneficial to have some familiarity with the fundamentals of Hadoop and Python. Table of Contents 1. Introduction to Machine Learning 2. Apache Spark Environment Setup and Configuration 3. Apache Spark 4. Apache Spark MLlib 5. Supervised Learning with Spark 6. Un-Supervised Learning with Apache Spark 7. Natural Language Processing with Apache Spark 8. Recommendation Engine with Distributed Framework 9. Deep Learning with Spark 10. Computer Vision with Apache Spark

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Practical Machine Learning with Spark ## 【One-Line Pitch】 A hands-on guide for data professionals who want to scale machine learning workflows using Apache Spark's distributed computing power, covering everything from environment setup to advanced topics like deep learning and computer vision. Ideal for data engineers, ML engineers, and data scientists who already know Hadoop and Python basics and want to move beyond single-machine ML. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces machine learning fundamentals—the evolution from AI to ML to deep learning, types of learning (supervised, unsupervised, semi-supervised, reinforcement), and real-world applications across healthcare, sentiment analysis, and video surveillance. Establishes why distributed processing matters for modern data volumes. - **Early (~10%–23%)**: Walks through Apache Spark environment setup in exhaustive detail—installing Cloudera VM, configuring AWS cloud instances, setting up HDP clusters with Ambari, and installing development tools like Sublime Text and Jupyter Notebook. Heavy on step-by-step configuration with screenshots. - **Early (~23%–32%)**: Covers Apache Spark core concepts—RDD vs. DataFrame vs. Dataset comparisons, transformations and actions, shared variables (accumulators and broadcast), job optimization techniques, storage levels, and workflow scheduling with Apache Oozie. Includes practical PySpark code for DataFrame operations like creating, renaming, and deduplicating data. - **Middle (~39%–48%)**: Dives into Spark MLlib—the distributed machine learning library. Covers ML pipeline components, data types (local/sparse/dense vectors, matrices), feature extraction (TF-IDF, Word2Vec, CountVectorizer), feature transformers (Tokenizer, N-Gram, Binarizer, StandardScaler), and feature selectors. Includes code examples for building unified ML pipelines. - **Middle (~48%–end)**: Moves into applied machine learning—supervised learning with regression and classification algorithms, unsupervised learning techniques, NLP with Spark, recommendation engines on distributed frameworks, deep learning integration, and computer vision applications. Includes model evaluation metrics (RMSE, MSE, MAE, R²) and practical implementation examples. ## 【Key Takeaways】 - **Distributed ML solves the standalone bottleneck** (Middle): When training data exceeds single-machine memory, processing loads can spike up to 95%, making training and testing painfully slow. Spark MLlib distributes the entire ML pipeline—preprocessing, feature extraction, model fitting, and evaluation—across clusters for scalable performance. - **RDD, DataFrame, and Dataset are not interchangeable** (Early): Each has distinct trade-offs—RDDs offer low-level control but slower aggregation; DataFrames provide auto schema discovery and faster exploratory analysis; Datasets combine type safety with performance. Choose based on whether you need compile-time safety or rapid iteration. - **Feature engineering is the heart of MLlib** (Middle): Spark provides a rich toolkit—TF-IDF for text weighting, N-Gram for sequence modeling, Binarizer for thresholding, StandardScaler for normalization, and PCA for dimensionality reduction. Mastering these transformers lets you build production-ready feature pipelines that scale. - **Shared variables optimize distributed computation** (Early): Accumulators and broadcast variables solve the cross-node communication problem in Spark. Broadcast variables cache large lookup tables on each worker, while accumulators provide fault-tolerant counters—essential for efficient distributed algorithms. - **Workflow orchestration matters for production ML** (Early): Apache Oozie binds scattered jobs into DAG-based workflows with coordinators that handle scheduling based on frequency and data availability. This turns ad-hoc ML scripts into repeatable, scheduled pipelines integrated with the Hadoop ecosystem. - **Model evaluation requires multiple metrics** (Middle): For regression tasks, RMSE, MSE, MAE, and R² each tell a different story about model quality. The book demonstrates computing all four in PySpark, emphasizing that no single metric captures model performance fully. - **MLlib bridges academic concepts and corporate practice** (Opening): The book explicitly aims to close the gap between academic ML knowledge and real-world industrial deployment, making it valuable for practitioners transitioning from theory to production systems. ## 【Reading Tips】 - **Skim the environment setup chapters** (Early, ~10%–23%): The Cloudera VM, AWS, and HDP installation steps are detailed but time-sensitive—versions and screenshots will age. Read for the conceptual flow, but don't memorize every click; use as reference when actually setting up your environment. - **Deep-read the MLlib chapter** (Middle, ~39%–48%): This is the book's core value—the feature extractors, transformers, and pipeline components are the building blocks you'll use daily. Pay special attention to TF-IDF, N-Gram, and the scaling/normalization transformers. - **Focus on the PySpark code patterns** (throughout): The book is code-heavy with practical examples. Rather than reading linearly, try running the DataFrame operations and ML pipeline examples in Jupyter Notebook or Google Colab to build muscle memory. - **Watch for the AI vs. ML conceptual foundation** (Opening, ~0%–10%): The early comparison tables and ML taxonomy provide useful mental models, but you can skim the application examples (healthcare, sentiment analysis) if you're already familiar with ML use cases. - **Skip the screenshots if you're experienced** (Early): Many figures show installation wizards and terminal outputs. If you've set up Spark before, these add little; focus instead on the architecture discussions and code snippets. ## 【Coverage Limits】 This guide covers the book's first half in depth (foundations, environment, Spark core, MLlib) and the supervised learning chapter's evaluation metrics. The later chapters on unsupervised learning, NLP, recommendation engines, deep learning, and computer vision are only briefly mentioned—the excerpts do not provide sufficient detail to summarize their specific algorithms and implementations. ##
Page 10
Gupta, for improving the standard and quality of this book. I agree that the content of this book will confound the reader with great interest. — Gourav Gupt...
View in text
Excerpt 2
-based Understand about the cloud instance setup using AWS. Install Apache Spark and Apache Hadoop on cloud using Amazon Elastic Compute Cloud (Amazon EC2)....
View in text
Excerpt 3
rmation on existing DataFrames using drop_duplicates(func): Figure 3.14: Program of drop duplicate function on existing Dataframe to remove the duplicity Fig...
View in text
Excerpt 4
lding value. Figure 4.26: Code and output of StandardScaler MinMaxScaler MinMaxScaler rescales each feature to a range varies between [0,1]. It transforms th...
View in text
Excerpt 5
ure, it is quite clear that the left side of data points of point l1 and right side of data points of point l2 are unable to find the classification with lin...
View in text
Excerpt 6
arkConf() import matplotlib.pyplot as plt from mpl_toolkits.mplot3d import Axes3D from pyspark.sql import SparkSession from pyspark.ml.feature import Standar...
View in text
Excerpt 7
to analyze the performance of a DL model. There are various categories of metrics in NNs; few categories are given as follows: Accuracy Metrics Probabilistic...
View in text
Excerpt 8
ted docker into the Amazon Web Services - Elastic MapReduce (AWS EMR) or Amazon Web Services - Elastic Compute Cluster (AWS EC2) where the DL model is config...
View in text
Tags
AI categories
Big DataArtificial IntelligenceBackend
ISBN: 9391392083
Publisher: BPB Publications
Publish Year: 2022
Language: English
Pages: 545
File Format: PDF
File Size: 18.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…