Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Norman Matloff

Rating No ratings yet

Machine learning without advanced math! This book presents a serious, practical look at machine learning, preparing you for valuable insights on your own data. The Art of Machine Learning is packed with real dataset examples and sophisticated advice on how to make full use of powerful machine learning methods. Readers will need only an intuitive grasp of charts, graphs, and the slope of a line, as well as familiarity with the R programming language. You’ll become skilled in a range of machine learning methods, starting with the simple k-Nearest Neighbors method (k-NN), then on to random forests, gradient boosting, linear/logistic models, support vector machines, the LASSO, and neural networks.Final chapters introduce text and image classification, as well as time series. You’ll learn not only how to use machine learning methods, but also why these methods work, providing the strong foundational background you’ll need in practice. Additional features How to avoid common problems, such as dealing with “dirty” data and factor variables with large numbers of levels A look at typical misconceptions, such as dealing with unbalanced data Exploration of the famous Bias-Variance Tradeoff, central to machine learning, and how it plays out in practice for each machine learning method Dozens of illustrative examples involving real datasets of varying size and field of application Standard R packages are used throughout, with a simple wrapper interface to provide convenient access. After finishing this book, you will be well equipped to start applying machine learning techniques to your own datasets.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# The Art of Machine Learning: A Hands-On Guide to Machine Learning with R ## 【One-Line Pitch】 A practical, math-light introduction to machine learning using R, teaching you not just how to apply methods like k-NN, random forests, and neural networks, but also why they work—perfect for analysts and programmers who want real-world ML skills without a statistics degree. ## 【Book Arc】 - **Opening (~0%–9%)**: Sets the "anti-cookbook" philosophy—understanding the why behind ML methods—and introduces the first method, k-NN, using a bike ridership prediction example. Covers regression functions, dummy variables, and the qeML/regtools package setup. - **Early (~9%–25%)**: Deepens k-NN with real datasets (MLB player weights), explains the bias-variance tradeoff through polling analogies, and introduces hyperparameter tuning, feature scaling (mmscale vs. scale), and the kNN() function's arguments. - **Early (~25%–34%)**: Shifts to classification models with the Telco Churn dataset, clarifying the confusing terminology (numeric-Y vs. classification), and demonstrates multiclass problems with vertebral column data and medical no-show prediction. - **Middle (~34%–44%)**: Covers model evaluation with ROC curves and AUC (using qeROC()), then formalizes the bias-variance tradeoff, overfitting/underfitting, and why too many features (like 42,000 ZIP codes) inflate variance. - **Middle (~44%–47%)**: Introduces dimension reduction with the Million Song Dataset (515,000 rows, 90 features), showing PCA via prcomp() and qePCA() to reduce feature sets, avoid overfitting, and speed up computation. ## 【Key Takeaways】 - **ML is regression in disguise** (Early): All ML methods estimate regression functions—classification is just regression with categorical Y. Understanding this unifies seemingly different techniques and clarifies confusing terminology. - **k-NN is the intuitive foundation** (Opening): Predicting by averaging over similar cases (neighbors) is both simple and powerful. The hyperparameter k is a "Goldilocks" choice—too small overfits, too large underfits. - **Bias-variance tradeoff is central** (Early): Reducing one increases the other. Small k or too many features inflate variance; omitting key features (like height in weight prediction) introduces bias. This tradeoff drives all hyperparameter and feature decisions. - **Feature scaling matters** (Early): Unscaled features give undue influence to variables with large values. The mmscale() function maps to [0,1] for bounded, unitless variables, unlike scale() which produces unbounded values. - **Dirty data and p-hacking are real threats** (Opening): Remove IDs and near-unique identifiers (like patient IDs in medical data) to avoid overfitting. Be aware of how feature engineering can lead to spurious findings. - **ROC curves evaluate classification honestly** (Middle): qeROC() wraps pROC to plot ROC and compute AUC on holdout sets—essential for assessing churn prediction models beyond simple accuracy. - **Dimension reduction fights the curse of dimensionality** (Middle): With 90 features in the Million Song Dataset, PCA creates new features as combinations of originals, reducing candidate sets from 2^90 to just 90—saving computation and preventing overfitting. ## 【Reading Tips】 - **Skim the code-heavy sections** (Early ~25%): The kNN() function arguments and R output dumps are reference material—grasp the concept, then return when you need syntax. - **Deep-read the bias-variance explanations** (Early ~16%–19%): The polling analogy (landline bias, margin of error) is the clearest explanation of variance and bias you'll find—master this and the rest of the book clicks. - **Work through the Telco Churn example** (Early ~28%): This is the book's flagship classification case. Follow the full pipeline: data loading, cleaning, k-NN modeling, then ROC evaluation. - **Don't skip the "Pitfall" sections**: They're marked recurring themes for a reason—they contain hard-won practical advice like removing IDs and handling factor variables with many levels. - **Install packages before starting** (Opening): You need R, qeML, and regtools (version 1.7+) installed and loaded. The book assumes you've done this, so set up early to avoid friction. ## 【Coverage Limits】 This guide covers the book's first half (through dimension reduction with PCA). Excerpts do not cover later chapters on random forests, gradient boosting, logistic models, SVMs, LASSO, neural networks, text/image classification, or time series—though the blurb indicates these follow the same practical, why-it-works approach. ##
Page 7
hapter 3: Bias, Variance, Overfitting, and Cross-Validation Chapter 4: Dealing with Large Numbers of Features PART II: TREE-BASED METHODS Chapter 5: A Step B...
View in text
Excerpt 2
n estimate of the true population regression function r(). Forming a prediction from just the closest k = 5 neighbors works from a very small sample. Imagine...
View in text
Excerpt 3
2 39.06 10.06 25.02 29.00 114.41  4.56 DH 3 68.83 22.22 50.09 46.61 105.99 -3.53 DH 4 69.30 24.65 44.31 44.64 101.87 11.21 DH 5 49.71  9.65 28.32 40.06 108.1...
View in text
Excerpt 4
  V75      V76      V77       V78      V79       V80 1 -41.1245 -8.40816 7.19877 -8.60176 -5.90857 -12.32437 14.68734 -54.32125        V81     V82       V83 ...
View in text
Excerpt 5
in the entire tree, as can be seen by typing dtout$nNodes.) Second, we see in part how that reduction was accomplished: DT was able to form its own groups of...
View in text
Excerpt 6
not take the ordering of results in a grid search literally. The first few “best” results may actually be similar. Moreover, the apparent “best” may actually...
View in text
Excerpt 7
e fact that the formula attempts to correct for overfitting. The qeLin() reports this too, and again a large discrepancy between this value and the first R2...
View in text
Excerpt 8
384401 PART IV METHODS BASED ON SEPARATING LINES AND PLANES The methods we’ve looked at so far were developed by statisticians or, in the case of boosting, b...
View in text
Tags
AI categories
Programmingmachine learning
ISBN: 1098168755
Publish Year: 2023
Language: English
Pages: 272
File Format: PDF
File Size: 6.4 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…