Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Carl McBride Ellis

Rating No ratings yet

Master the essential tools and techniques for supervised machine learning on tabular data with this practical guide to regression and classification. Through clear explanations, code snippets, and hands-on notebooks, you'll learn how to use Python and leading machine learning libraries, including pandas, scikit-learn, CatBoost, LightGBM, XGBoost, TabPFN, and TabICL, to build predictive models for real-world datasets. The book covers the complete workflow, from data exploration and cleaning to model development, evaluation, and optimization. You'll learn how to perform regression analysis for accurate point predictions and estimate uncertainty using conformal prediction intervals. For classification tasks, you'll explore probabilistic predictions and calibration techniques to improve model reliability. You'll also discover practical approaches to feature engineering, feature selection, and hyperparameter optimization to enhance model performance. In addition, the book introduces tabular foundation models and in-context learning techniques, providing insight into the latest advances in machine learning for structured data. By the end of the book, you'll have the skills and confidence to develop, evaluate, and deploy supervised machine learning models for a wide range of tabular data applications.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to supervised machine learning on tabular data, walking you from data cleaning and cross-validation through regression, classification, and modern gradient-boosting and foundation models. Best for practitioners who already know some Python and want a rigorous, workflow-oriented reference for building and evaluating predictive models. 【Book Arc】 - **Opening (~0%–15%)**: Frames core vocabulary and statistical foundations — estimators vs. approximators, errors vs. residuals, aleatoric vs. epistemic uncertainty, and descriptive statistics/visualization (histograms, box/violin plots, correlation coefficients). - **Early (~15%–30%)**: Covers the unglamorous but decisive preprocessing work: missing-value handling (MCAR/MAR/MNAR, imputation strategies), outlier detection, encoding, scaling, and cross-validation design including data leakage and covariate shift. - **Middle (~30%–55%)**: Builds the two core supervised tasks — regression (linear models, trees, regularization, bias-variance, quantile and conformal prediction intervals) and classification (logistic regression, decision trees, probabilistic metrics, calibration, multiclass and imbalanced data). - **Late (~55%–75%)**: Extends into generalized linear/additive models, ensemble estimators (Random Forest, boosting, stacking), and hyperparameter optimization, including the Optimizer's Curse and automated search routines. - **Ending (~75%–100%)**: Moves into feature engineering and selection, then introduces tabular foundation models and in-context learning (TabPFN, TabICL) as the frontier of structured-data prediction. 【Key Takeaways】 - **Supervised tabular ML is a workflow, not a single model choice** (Opening): the book sequences exploration → cleaning → validation → modeling → evaluation → optimization, so you learn where each decision fits. - **Uncertainty deserves explicit treatment** (Early–Middle): the text distinguishes aleatoric from epistemic uncertainty and teaches conformal prediction intervals and quantile regression for honest point-prediction ranges. - **Classification is fundamentally probabilistic** (Middle): calibration, reliability diagrams, strictly proper scoring rules, and Venn-ABERS calibration are treated as first-class concerns, not afterthoughts. - **Cross-validation design prevents self-deception** (Early): nested CV, fold-count bias, permutation invariance, and data leakage are covered to keep out-of-sample estimates trustworthy. - **Hyperparameter optimization has its own failure mode** (Late): the Optimizer's Curse and optimistic bias are named explicitly, with guidance on early stopping and automated search. - **Ensembles and boosting are the practical workhorses** (Late): Random Forest, gradient-boosted decision trees, stacking, and convex combinations of predictions are presented as the default strong baselines for tabular data. - **Feature engineering and selection remain decisive** (Late): interaction terms, selection methods, and preprocessing choices are shown to materially affect performance before any exotic model is tried. - **Tabular foundation models are emerging** (Ending): TabPFN and TabICL with in-context learning are introduced as a newer alternative paradigm for structured data. 【Reading Tips】 - **Deep-read the cross-validation and data-cleaning chapters** (Early): these underpin every later result; skimming them will make the regression and classification chapters harder to trust. - **Skim the statistics refresher if you're already comfortable** (Opening): use it as a vocabulary check rather than a full study pass. - **Treat the regression and classification chapters as parallel tracks** (Middle): read one fully, then the other, comparing how uncertainty and evaluation are handled differently. - **Don't skip the hyperparameter optimization and feature chapters** (Late): they contain the practical judgment calls that separate a working model from a well-tuned one. - **Run the accompanying notebooks** (throughout): the book is code-driven, and the GitHub repository is the intended companion for hands-on practice. 【Coverage Limits】 This guide is synthesized from stratified excerpts covering the table of contents and early chapters; detailed content of later chapters (ensembles, HPO, feature engineering, foundation models) is inferred from section headings rather than full text. Specific code examples, datasets, and numerical results are not covered here.
Page 2
rrors, or conceptual mistakes are entirely of my own making. I would be delighted if you could bring such issues to my attention either using the dedicated D...
View in text
Page 4
tion of NaN with missingno . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 4.1.2 Types of missing: MCAR, MAR, and MNAR . . . . ....
View in text
Page 6
Linear regression learning curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108 7.12 Decision tree regressor . . . . . . . . . . ....
View in text
Page 3
. . . . . . . 159 8.12.7 Negative rates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 8.12.8 Matthe...
View in text
Page 9
. 216 13.2 Why no traditional ANN? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219 13.3 Foundation models . . . . . . ....
View in text
Page 12
k “The Nature of Statistical Learning Theory” by Vladimir N. Vapnik). 1.4 Prediction vs forecast The etymology of the verb to predict is to say before, but b...
View in text
Page 15
us Facure “Causal Inference in Python”, O’Reilly Media, Inc. (2023) • Aleksander Molak “Causal Inference and Discovery in Python”, Packt Publishing Limited (...
View in text
Page 19
Dispersion: range, variance, MAD, and quartiles In Figure 2.3 we have two distributions that have exactly the same mean value. Figure 2.3: Two distributions...
View in text
Tags
AI categories
Artificial IntelligenceDataPython
ISBN: 1808081315
Publish Year: 2026
Language: English
Pages: 238
File Format: PDF
File Size: 14.8 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…