converted pdf, Book description
Many industry experts consider unsupervised learning the next frontier in artificial intelligence, one that may hold the key to general artificial intelligence. Since the majority of the world's data is unlabeled, conventional supervised learning cannot be applied. Unsupervised learning, on the other hand, can be applied to unlabeled datasets to discover meaningful patterns buried deep in the data, patterns that may be near impossible for humans to uncover.
Author Ankur Patel shows you how to apply unsupervised learning using two simple, production-ready Python frameworks: Scikit-learn and TensorFlow using Keras. With code and hands-on examples, data scientists will identify difficult-to-find patterns in data and gain deeper business insight, detect anomalies, perform automatic feature engineering and selection, and generate synthetic datasets. All you need is programming and some machine learning experience to get started.
* Compare the strengths and weaknesses of the different machine learning approaches: supervised, unsupervised, and reinforcement learning
* Set up and manage machine learning projects end-to-end
* Build an anomaly detection system to catch credit card fraud
* Clusters users into distinct and homogeneous groups
* Perform semisupervised learning
* Develop movie recommender systems using restricted Boltzmann machines
* Generate synthetic images using generative adversarial networks
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical bridge from classical machine learning into the unlabeled-data frontier: learn to find structure, flag anomalies, and generate synthetic data with Scikit-learn and TensorFlow/Keras. Best for working data scientists and engineers who already know supervised learning and want production-ready unsupervised techniques.
【Book Arc】
- **Opening (~0%–10%)**: Frames why unsupervised learning matters when most data is unlabeled, contrasts supervised, unsupervised, semisupervised, and reinforcement approaches, and introduces core terminology (features, outliers, data drift) using a spam-filter example.
- **Early (~10%–30%)**: Walks through an end-to-end supervised project—environment setup with Git and Jupyter, data acquisition, standardization vs. normalization, handling NaNs and categorical values, k-fold cross-validation, and precision/recall trade-offs—using credit card fraud detection as the running case.
- **Early–Middle (~30%–40%)**: Covers dimensionality reduction as both a goal and a pipeline step: PCA and its variants (incremental, sparse, kernel), SVD, manifold learning (Isomap), and independent component analysis, applied to the MNIST digits dataset.
- **Middle (~40%–50%)**: Turns dimensionality reduction into anomaly detection—reconstruction-error scoring with PCA, sparse PCA, and sparse random projection to separate fraudulent transactions from normal ones.
- **Late (~50%–75%)**: Moves into clustering and semisupervised learning, grouping users into homogeneous segments and blending labeled with unlabeled data. (Excerpts do not cover the specific algorithms in detail.)
- **Ending (~75%–100%)**: Builds generative and recommender systems—restricted Boltzmann machines for movie recommendations and generative adversarial networks for synthetic images. (Excerpts do not cover these chapters in detail.)
【Key Takeaways】
- **Unlabeled data is the default, not the exception** (Opening): since most real-world data lacks labels, unsupervised methods unlock patterns supervised models simply cannot reach.
- **Project hygiene precedes modeling** (Early): environment setup, version control, standardization, NaN handling, and stratified k-fold validation are treated as non-negotiable foundations before any algorithm runs.
- **Dimensionality reduction serves two masters** (Early–Middle): it can be the end goal (anomaly detection) or a pipeline step that makes large-scale image, video, speech, and text problems tractable.
- **PCA has a family, not a single form** (Middle): incremental, sparse, and kernel variants each trade off linearity, interpretability, and scalability—choosing well matters more than defaulting to vanilla PCA.
- **Reconstruction error is a practical anomaly score** (Middle): projecting data down and back up, then measuring the residual, cleanly separates fraud from normal transactions on the credit card dataset.
- **Precision and recall are business decisions** (Early): in fraud detection, high precision avoids antagonizing customers while high recall avoids losing money—the right balance depends on cost, not just metrics.
- **Gradient boosting sets a strong supervised baseline** (Early): LightGBM and XGBoost outperform random forests and logistic regression in the fraud case, giving a benchmark against which unsupervised approaches can be judged.
- **Generative models extend unsupervised learning into creation** (Ending): RBMs and GANs show how learned structure can power recommenders and synthetic image generation.
【Reading Tips】
- **Skim Chapter 2 if you're already fluent in supervised ML**: the environment setup and cross-validation material is foundational but not the book's core value.
- **Deep-read the PCA variant comparisons and the anomaly detection chapter**: these are the most transferable techniques and the clearest worked examples.
- **Run the code, don't just read it**: the book is explicitly hands-on, with Jupyter notebooks on GitHub; the fraud and MNIST cases only click when you execute them.
- **Watch the preprocessing details**: standardization vs. normalization, NaN imputation, and categorical encoding are easy to gloss over but drive results.
- **Treat the later generative chapters as a springboard**: use them to understand the concepts, then consult current literature since the tooling evolves quickly.
【Coverage Limits】
This guide is based on stratified excerpts covering roughly the first half of the book; the clustering, semisupervised, RBM, and GAN chapters are referenced but not detailed in the source material.
Page 10
work-based system that can identify faces with 97% accuracy. This is near human-level performance and is a more than 27% improvement over previous systems. 2...
dized data, all the normalized data is on a positive scale. Identify nonnumerical values by feature Some machine learning algorithms cannot handle nonnumeric...
images, video, speech, and text. The MNIST Digits Database Before we introduce the dimensionality reduction algorithms, let’s explore the dataset that we wil...
st set. We will then use the Scikit-Learn inverse_transform function to recreate the original dimensions from the principal components matrix of the test set...
l_validation) The MSE of this very naive prediction is 1.05. This is our baseline: Mean squared error using naive prediction: 1.055420084238528 Let’s see if...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Hands-On Unsupervised Learning Using Python How to Build Applied Machine Learning Solutions from Unlabeled Data (Ankur A. Patel)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Hands-On Unsupervised Learning Using Python How to Build Applied Machine Learning Solutions from Unlabeled Data (Ankur A. Patel)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment