Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ankur A. Patel

Rating No ratings yet

converted pdf, Book description Many industry experts consider unsupervised learning the next frontier in artificial intelligence, one that may hold the key to general artificial intelligence. Since the majority of the world's data is unlabeled, conventional supervised learning cannot be applied. Unsupervised learning, on the other hand, can be applied to unlabeled datasets to discover meaningful patterns buried deep in the data, patterns that may be near impossible for humans to uncover. Author Ankur Patel shows you how to apply unsupervised learning using two simple, production-ready Python frameworks: Scikit-learn and TensorFlow using Keras. With code and hands-on examples, data scientists will identify difficult-to-find patterns in data and gain deeper business insight, detect anomalies, perform automatic feature engineering and selection, and generate synthetic datasets. All you need is programming and some machine learning experience to get started. * Compare the strengths and weaknesses of the different machine learning approaches: supervised, unsupervised, and reinforcement learning * Set up and manage machine learning projects end-to-end * Build an anomaly detection system to catch credit card fraud * Clusters users into distinct and homogeneous groups * Perform semisupervised learning * Develop movie recommender systems using restricted Boltzmann machines * Generate synthetic images using generative adversarial networks

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical bridge from classical machine learning into the unlabeled-data frontier: learn to find structure, flag anomalies, and generate synthetic data with Scikit-learn and TensorFlow/Keras. Best for working data scientists and engineers who already know supervised learning and want production-ready unsupervised techniques. 【Book Arc】 - **Opening (~0%–10%)**: Frames why unsupervised learning matters when most data is unlabeled, contrasts supervised, unsupervised, semisupervised, and reinforcement approaches, and introduces core terminology (features, outliers, data drift) using a spam-filter example. - **Early (~10%–30%)**: Walks through an end-to-end supervised project—environment setup with Git and Jupyter, data acquisition, standardization vs. normalization, handling NaNs and categorical values, k-fold cross-validation, and precision/recall trade-offs—using credit card fraud detection as the running case. - **Early–Middle (~30%–40%)**: Covers dimensionality reduction as both a goal and a pipeline step: PCA and its variants (incremental, sparse, kernel), SVD, manifold learning (Isomap), and independent component analysis, applied to the MNIST digits dataset. - **Middle (~40%–50%)**: Turns dimensionality reduction into anomaly detection—reconstruction-error scoring with PCA, sparse PCA, and sparse random projection to separate fraudulent transactions from normal ones. - **Late (~50%–75%)**: Moves into clustering and semisupervised learning, grouping users into homogeneous segments and blending labeled with unlabeled data. (Excerpts do not cover the specific algorithms in detail.) - **Ending (~75%–100%)**: Builds generative and recommender systems—restricted Boltzmann machines for movie recommendations and generative adversarial networks for synthetic images. (Excerpts do not cover these chapters in detail.) 【Key Takeaways】 - **Unlabeled data is the default, not the exception** (Opening): since most real-world data lacks labels, unsupervised methods unlock patterns supervised models simply cannot reach. - **Project hygiene precedes modeling** (Early): environment setup, version control, standardization, NaN handling, and stratified k-fold validation are treated as non-negotiable foundations before any algorithm runs. - **Dimensionality reduction serves two masters** (Early–Middle): it can be the end goal (anomaly detection) or a pipeline step that makes large-scale image, video, speech, and text problems tractable. - **PCA has a family, not a single form** (Middle): incremental, sparse, and kernel variants each trade off linearity, interpretability, and scalability—choosing well matters more than defaulting to vanilla PCA. - **Reconstruction error is a practical anomaly score** (Middle): projecting data down and back up, then measuring the residual, cleanly separates fraud from normal transactions on the credit card dataset. - **Precision and recall are business decisions** (Early): in fraud detection, high precision avoids antagonizing customers while high recall avoids losing money—the right balance depends on cost, not just metrics. - **Gradient boosting sets a strong supervised baseline** (Early): LightGBM and XGBoost outperform random forests and logistic regression in the fraud case, giving a benchmark against which unsupervised approaches can be judged. - **Generative models extend unsupervised learning into creation** (Ending): RBMs and GANs show how learned structure can power recommenders and synthetic image generation. 【Reading Tips】 - **Skim Chapter 2 if you're already fluent in supervised ML**: the environment setup and cross-validation material is foundational but not the book's core value. - **Deep-read the PCA variant comparisons and the anomaly detection chapter**: these are the most transferable techniques and the clearest worked examples. - **Run the code, don't just read it**: the book is explicitly hands-on, with Jupyter notebooks on GitHub; the fraud and MNIST cases only click when you execute them. - **Watch the preprocessing details**: standardization vs. normalization, NaN imputation, and categorical encoding are easy to gloss over but drive results. - **Treat the later generative chapters as a springboard**: use them to understand the concepts, then consult current literature since the tooling evolves quickly. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book; the clustering, semisupervised, RBM, and GAN chapters are referenced but not detailed in the source material.
Page 10
work-based system that can identify faces with 97% accuracy. This is near human-level performance and is a more than 27% improvement over previous systems. 2...
View in text
Excerpt 2
dized data, all the normalized data is on a positive scale. Identify nonnumerical values by feature Some machine learning algorithms cannot handle nonnumeric...
View in text
Excerpt 3
images, video, speech, and text. The MNIST Digits Database Before we introduce the dimensionality reduction algorithms, let’s explore the dataset that we wil...
View in text
Excerpt 4
st set. We will then use the Scikit-Learn inverse_transform function to recreate the original dimensions from the principal components matrix of the test set...
View in text
Excerpt 5
236.754581 9889.0 42536 85070.0 85077.0 298.587755 16786.0 42537 85058.0 85078.0 309.946867 16875.0 42538 85074.0 85079.0 375.698458 34870.0 42539 85065.0 85...
View in text
Excerpt 6
hape) X_test_AE_noisy = X_test_AE.copy() + noise_factor * \ np.random.normal(loc=0.0, scale=1.0, size=X_test_AE.shape) Denoising Autoencoder Compared to the...
View in text
Excerpt 7
l_validation) The MSE of this very naive prediction is 1.05. This is our baseline: Mean squared error using naive prediction: 1.055420084238528 Let’s see if...
View in text
Excerpt 8
et, test_set = pickle.load(f, encoding='latin1') f.close() X_train, y_train = train_set[0], train_set[1] X_validation, y_validation = validation_set[0], vali...
View in text
Tags
AI categories
Artificial IntelligencePythonData
ISBN: 1492035645
Publisher: O'Reilly Media
Publish Year: 2019
Language: English
Pages: 515
File Format: PDF
File Size: 6.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…