Page
1
A I-A ssisted Sta tistics for D ata Scientists AI-Assisted Statistics for Data Scientists 50+ Essential Concepts Using R and Python Peter Bruce, Andrew Bruce & Peter Gedeck 3rd Edition
Page
2
9 7 9 8 3 4 1 6 6 6 2 8 3 5 7 9 9 9 ISBN: 979-8-341-66628-3 US $79.99 CAN $99.99 DATA SCIENCE Statistical methods are a key part of AI and data science, yet few data scientists have formal statistical training. The third edition of this popular guide expands its practical foundations in R and Python into the modern AI toolkit, with new chapters on neural networks, deep learning, and large language models. Generative AI is integrated throughout, showing how tools such as ChatGPT, Claude, and Gemini work, and how they can support real-world statistical workflows. This book highlights concepts that matter most when working with data, building predictive models, and deploying AI responsibly. If you’re comfortable with R or Python and have had some exposure to basic statistics, this concise reference will boost your statistical literacy, your understanding of how AI works, and your confidence in real-world data science and AI projects. • Conduct exploratory data analysis to improve models • Apply sampling and experimental design to reduce bias and answer questions with clarity • Use regression to understand data-generating processes and detect anomalies • Build predictive models using classification, clustering, and unsupervised learning with unbalanced data • Understand how generative AI works, and use tools like ChatGPT, Claude, and Gemini to help you analyze data and build machine learning models Peter Bruce is the founder of the Institute for Statistics Education at Statistics.com, which offers about 80 courses, roughly half of which are aimed at data scientists. Andrew Bruce, principal research scientist at Amazon, has over 30 years of experience in statistics and data science in academia, government, and business. Peter Gedeck is a senior data scientist at Collaborative Drug Discovery and lecturer at UVA who specializes in algorithms to predict biological and physicochemical properties of drug candidates. AI-Assisted Statistics for Data Scientists “This book makes the connection between useful statistical terms and principles and today’s data mining lingo and practices, with clear explanations and plenty of examples.” Galit Shmueli, lead author of Machine Learning for Business Analytics and chair professor, National Tsing Hua University, Taiwan “Brilliantly balances the power of modern AI with a critical eye on its statistical limitations and risks.” Ari Joury, founder and CEO at Wangari Global
Page
3
Peter Bruce, Andrew Bruce, and Peter Gedeck AI-Assisted Statistics for Data Scientists 50+ Essential Concepts Using R and Python THIRD EDITION
Page
4
979-8-341-66628-3 [LSI] AI-Assisted Statistics for Data Scientists by Peter Bruce, Andrew Bruce, and Peter Gedeck Copyright © 2026 Peter Bruce, Andrew Bruce, and Peter Gedeck. All rights reserved. Published by O’Reilly Media, Inc., 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (https://oreilly.com). For more information, contact our corporate/institu‐ tional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Michelle Smith Development Editor: Corbin Collins Production Editor: Ashley Stussy Copyeditor: Emily Wydeven Proofreader: Krsta Technology Solutions Indexers: Peter Bruce, Andrew Bruce, and Peter Gedeck Cover Designer: Susan Brown Cover Illustrator: Karen Montgomery Interior Designer: David Futato Interior Illustrator: Kate Dullea May 2017: First Edition May 2020: Second Edition June 2026: Third Edition Revision History for the First Edition 2026-06-16: First Release See https://oreilly.com/catalog/errata.csp?isbn=9798341666283 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. AI-Assisted Statistics for Data Scientists, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Page
5
Peter Bruce and Andrew Bruce would like to dedicate this book to the memories of our parents, Victor G. Bruce and Nancy C. Bruce, who cultivated a passion for math and science; and to our early mentors John W. Tukey and Julian Simon and our lifelong friend Geoff Watson, who helped inspire us to pursue a career in statistics. Peter Gedeck would like to dedicate this book to Tim Clark and Christian Kramer, with deep thanks for their scientific collaboration and friendship.
Page
6
(This page has no text content)
Page
7
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xv 1. Exploratory Data Analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 Elements of Structured Data 2 Further Reading 4 Data Structures 4 Rectangular Data 5 Data Frames and Indexes 6 Nonrectangular Data Structures 7 Further Reading 8 Data Dictionaries and Catalogs 8 Further Reading 12 Estimates of Location 13 Mean 14 Median and Robust Estimates 15 Example: Location Estimates of Population and Murder Rates 17 Further Reading 18 Estimates of Variability 19 Standard Deviation and Related Estimates 20 Estimates Based on Percentiles 22 Example: Variability Estimates of State Population 23 Further Reading 24 Exploring the Data Distribution 24 Percentiles and Boxplots 25 Frequency Tables and Histograms 27 Density Plots and Estimates 30 Further Reading 32 Exploring Binary and Categorical Data 32 v
Page
8
Mode 34 Expected Value 34 Probability 35 Further Reading 35 Correlation 35 Scatterplots 39 Further Reading 41 Exploring Two or More Variables 41 Hexagonal Binning and Contours (Plotting Numeric-Versus-Numeric Data) 42 Two Categorical Variables 44 Categorical and Numeric Data 46 Visualizing Multiple Variables 48 Further Reading 50 Interpreting Visualization Results with AI 50 Summary 52 Exploration with AI 52 2. Data and Sampling Distributions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 Random Sampling and Sample Bias 56 Bias 58 Random Selection 59 Size Versus Quality: When Does Size Matter? 60 Sample Mean Versus Population Mean 61 Further Reading 61 Selection Bias 62 Regression to the Mean 63 Further Reading 65 Sampling Distribution of a Statistic 65 Central Limit Theorem 68 Standard Error 68 Further Reading 69 The Bootstrap 69 Resampling Versus Bootstrapping 73 Further Reading 73 Confidence Intervals 73 Further Reading 76 Normal Distribution 76 Standard Normal and QQ-Plots 78 Long-Tailed Distributions 80 Further Reading 82 Student’s t-Distribution 82 Further Reading 85 vi | Table of Contents
Page
9
Binomial Distribution 85 Further Reading 87 Chi-Square Distribution 88 Further Reading 89 F-Distribution 89 Further Reading 90 Poisson and Related Distributions 90 Poisson Distributions 90 Exponential Distribution 91 Estimating the Failure Rate 92 Weibull Distribution 92 Further Reading 93 Summary 93 Exploration with AI 93 3. Statistical Experiments and Significance Testing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 A/B Testing 96 Why Have a Control Group? 98 Why Just A/B? Why Not C, D,…? 100 Further Reading 100 Hypothesis Tests 101 The Null Hypothesis 102 Alternative Hypothesis 102 One-Way Versus Two-Way Hypothesis Tests 103 Further Reading 103 Resampling 104 Permutation Test 104 Example: Web Stickiness 105 Exhaustive and Bootstrap Permutation Tests 109 Permutation Tests: The Bottom Line for Data Science 110 Further Reading 111 Statistical Significance and p-Values 111 p-Value 113 Alpha 114 Type 1 and Type 2 Errors 116 Data Science and p-Values 116 Further Reading 117 t-Tests 117 Further Reading 119 Multiple Testing 119 Further Reading 122 Degrees of Freedom 123 Table of Contents | vii
Page
10
Further Reading 124 ANOVA 124 F-Statistic 127 Two-Way ANOVA 129 Further Reading 130 Chi-Square Test 130 Chi-Square Test: A Resampling Approach 131 Chi-Square Test: Statistical Theory 133 Fisher’s Exact Test 134 Relevance for Data Science 136 Further Reading 137 Multi-Arm Bandit Algorithm 137 Further Reading 140 Power and Sample Size 141 Sample Size 142 Further Reading 144 Summary 145 Exploration with AI 145 4. Regression and Prediction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 Simple Linear Regression 147 The Regression Equation 149 Fitted Values and Residuals 151 Least Squares 153 Prediction Versus Explanation (Profiling) 154 Further Reading 154 Multiple Linear Regression 155 Example: King County Housing Data 156 Assessing the Model 157 Cross-Validation 160 Model Selection and Stepwise Regression 160 Weighted Regression 164 Further Reading 165 Prediction Using Regression 165 The Dangers of Extrapolation 166 Confidence and Prediction Intervals 166 Factor Variables in Regression 168 Dummy Variables Representation 169 Factor Variables with Many Levels 171 Ordered Factor Variables 173 Interpreting the Regression Equation 174 Correlated Predictors 175 viii | Table of Contents
Page
11
Multicollinearity 177 Confounding Variables 177 Interactions and Main Effects 179 Regression Diagnostics 181 Outliers 182 Influential Values 184 Heteroskedasticity, Nonnormality, and Correlated Errors 187 Partial Residual Plots and Nonlinearity 190 Polynomial and Spline Regression 192 Polynomial 193 Splines 194 Generalized Additive Models 197 Further Reading 198 Analyzing the Results of a Regression Analysis with AI 199 Summary 200 Exploration with AI 201 5. Classification. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203 Naive Bayes 204 Why Exact Bayesian Classification Is Impractical 205 The Naive Solution 206 Numeric Predictor Variables 209 Further Reading 212 Discriminant Analysis 212 Covariance Matrix 213 Fisher’s Linear Discriminant 214 A Simple Example 215 Extensions of Discriminant Analysis 218 Further Reading 218 Logistic Regression 219 Logistic Response Function and Logit 219 Logistic Regression and the GLM 221 Generalized Linear Models 223 Predicted Values from Logistic Regression 223 Interpreting the Coefficients and Odds Ratios 224 Linear and Logistic Regression: Similarities and Differences 225 Assessing the Model 227 Further Reading 230 Evaluating Classification Models 231 Confusion Matrix 232 The Rare Class Problem 235 Precision, Recall, and Specificity 235 Table of Contents | ix
Page
12
ROC Curve 236 AUC 238 Lift 240 Further Reading 241 Strategies for Imbalanced Data 241 Undersampling 242 Oversampling and Up/Down Weighting 244 Data Generation 245 Cost-Based Classification 246 Exploring the Predictions 246 Further Reading 247 Summary 248 Exploration with AI 248 6. Statistical Machine Learning. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249 K-Nearest Neighbors 250 A Small Example: Predicting Loan Default 251 Distance Metrics 254 One-Hot Encoder 255 Standardization (Normalization, z-Scores) 256 Choosing K 259 KNN as a Feature Engine 260 Tree Models 262 A Simple Example 263 The Recursive Partitioning Algorithm 266 Measuring Homogeneity or Impurity 268 Stopping the Tree from Growing 269 Predicting a Continuous Value 271 How Trees Are Used 271 Further Reading 272 Bagging and the Random Forest 272 Bagging 274 Random Forest 274 Variable Importance 279 Hyperparameters 282 Boosting 284 The Boosting Algorithm 285 XGBoost 286 Regularization: Avoiding Overfitting 288 Hyperparameters and Cross-Validation 292 Summary 296 Exploration with AI 297 x | Table of Contents
Page
13
7. Unsupervised Learning. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299 Principal Components Analysis 300 A Simple Example 301 Computing the Principal Components 304 Interpreting Principal Components 305 Correspondence Analysis 308 Analyzing the Results of Principal Components and Correspondence Analysis with AI 310 Further Reading 311 K-Means Clustering 311 A Simple Example 312 K-Means Algorithm 315 Interpreting the Clusters 316 Analyzing the Clusters with AI 318 Selecting the Number of Clusters 319 Hierarchical Clustering 322 A Simple Example 322 The Dendrogram 323 The Agglomerative Algorithm 325 Measures of Dissimilarity 325 Model-Based Clustering 327 Multivariate Normal Distribution 327 Mixtures of Normals 329 Selecting the Number of Clusters 331 Further Reading 334 Scaling and Categorical Variables 334 Scaling the Variables 335 Dominant Variables 337 Categorical Data and Gower’s Distance 339 Problems with Clustering Mixed Data 342 Returning to AI 343 Summary 344 Exploration with AI 345 8. Neural Networks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347 The Basic Structure of a Neural Net 347 Further Reading 349 Fitting a Network to Data 349 Weights and Biases 351 Activation Functions in the Hidden Layers 351 The Final Layer 352 The Training Process 354 Table of Contents | xi
Page
14
A Simple Model 354 Further Reading 358 Backpropagation and Gradient Descent 358 Overfitting 360 Regularization 362 Interpretation and Variable Importance 365 Hyperparameter Optimization 368 Further Reading 369 Loss (Cost) Functions 369 Loss Functions for Regression 370 Loss Functions for Classification 370 Performance 371 Summary 373 Exploration with AI 373 9. Deep Learning. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375 Advantages of Depth 375 RNNs 377 Text Data 379 Memory Issues 380 Further Reading 383 CNNs 383 Intuition Behind CNNs 385 CNN Architecture 387 Image Data and Transfer Learning 389 Putting It All Together: An Example 391 Further Reading 396 Transformers 396 Embeddings 398 Attention Mechanisms 399 Architecture Overview 403 Revisiting the MNIST Image Classification Example 405 Further Reading 409 Reinforcement Learning 409 Multi-Arm Bandits 410 Markov Decision Processes 411 Deep Reinforcement Learning 414 Further Reading 415 Summary 415 Exploration with AI 415 xii | Table of Contents
Page
15
10. Generative AI and Large Language Models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 417 Generative AI and LLMs 418 Predicting the Next Word 418 Creativity and Temperature 419 Learning Language and Generating Images: Attention 420 Further Reading 422 Reinforcement Learning from Human Feedback 422 Alignment 423 Reward Models 423 Further Reading 424 Image Generation 424 Integrating Text Input 426 Further Reading 427 Scale and Power 427 Parallelization and GPUs 428 Mixture of Experts Models 429 Emergent Capabilities 429 Artificial General Intelligence 430 Further Reading 431 Training LLMs and Adapting LLMs for Specific Purposes 431 Fine-Tuning 432 Retrieval-Augmented Generation (RAG) 432 Web Searches 433 Prompt Engineering 433 Pattern Matching and In-Context Learning 435 Interactive Conversation 435 Further Reading 436 Deploying and Using Generative AI 436 Augmenting AI Training Data 438 Integrating Generative AI into Applications 438 Agentic AI 439 Task Decomposition and Chain-of-Thought Reasoning 439 The Generative AI Stack 439 Further Reading 441 Summary 441 Exploration with AI 441 11. Caveats and Concerns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 443 Accuracy and Reliability 443 Model Collapse 445 Value of Proprietary Data 445 Vibe Coding 447 Table of Contents | xiii
Page
16
Using AI for Statistical Analysis 449 Misinterpreting Visual Information 450 Inconsistency and Sloppiness 451 Arithmetic and Problem Solving 452 Self-Criticism 453 Best Practices 454 AI, Society, and the Individual 455 Personal Companions 457 Big Brother 458 Artificial General Intelligence 459 World Models 460 Data Curation and Governance 461 Data Curation Pipeline 461 Data Ownership and Control Protection 463 Summary 465 Exploration with AI 466 Bibliography. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 467 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 469 xiv | Table of Contents
Page
17
Preface This book, which has been updated with extensive new chapters and material on AI (neural networks, deep learning, and generative AI), is aimed at the data scientist and the machine learning engineer who has some prior (perhaps spotty or ephemeral) exposure to statistics. Illustrations are provided using the R and/or Python program‐ ming languages. Some familiarity with those languages is useful, but, with the vibe coding capabilities of modern AI tools, even those with no programming background can benefit from this book. Two of the authors came to the world of data science from the world of statistics, and have some appreciation of the contribution that statistics can make to the art of data science. At the same time, we are well aware of the limitations of traditional statistics instruction: statistics as a discipline is a century and a half old, and most statistics textbooks and courses are laden with the momentum and inertia of an ocean liner. The methods in the first portion of this book have some connection—historical or methodological—to the discipline of statistics. Neural nets—the underpinning of modern AI—evolved mainly out of computer science and are covered in the latter part of the book. Data scientists, machine learning practitioners, and software engineers have come to know AI as a powerful coding tool. Vibe coding refers to the ability of AI models to produce code based on specifications (a prompt) written in English or some other natural language. A related concept, agentic coding, is broader, covering autonomous task execution by AI agents within a structured workflow that includes goal and task specification, testing, validation, and oversight. In this book, examples in R and Python are shown, and the initial chapters (the ones that focus on statistics) discuss the use of vibe coding, primarily at the conclusion of the chapters. The latter chapters discuss AI in more depth, and go into more detail on its use for coding. Why learn about statistics, coding, and the models that underpin AI if large language models (LLMs) can do the work for you? If you can use AI as a quasi-magical tool to do your work, why do you really need to understand the underlying concepts? xv
Page
18
• AI can make mistakes, mistakes which may not be readily apparent. • If used properly in statistics and data science, AI will ask you questions and offer choices that require some knowledge of the underlying statistical concepts to answer. • If you have some understanding of neural nets and how—based on a statistical and machine learning foundation—they enable deep learning and generative AI, you are in a better position to understand their strengths and weaknesses. • While AI is very effective at providing “textbook” solutions to well-defined statis‐ tical problems, it does not yet have the ability to do the kind of critical thinking and problem solving that is required to work through an ambiguous data science project from end-to-end. With respect to statistics, this book seeks to: • Lay out, in digestible, navigable, and easily referenced form, key concepts from statistics that are relevant to data science. • Explain which concepts are important and useful from a data science perspective, which are less so, and why. With respect to AI, this book seeks to: • Provide a conceptual overview, at a high level in nontechnical terms, the statis‐ tical and machine learning foundations and the current algorithms of neural networks, deep learning, and generative AI. • Illustrate the use of generative AI for statistical analysis. Note that this book is not intended as a practitioners guide to deep learning and generative AI. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, keywords, packages, and libraries. xvi | Preface
Page
19
Key Terms Data science is a fusion of multiple disciplines, including statistics, computer science, information technology, and domain-specific fields. As a result, several different terms could be used to reference a given concept. Key terms and their synonyms will be highlighted throughout the book in a sidebar such as this. This element signifies a tip or suggestion. This element signifies a general note. This element indicates a warning or caution. Using Code Examples In all cases, this book gives code examples first in R and then in Python. In order to avoid unnecessary repetition, we generally show only output and plots created by the R code. You can find the complete code in R and Python as well as the data sets for download at https://github.com/gedeck/ai-assisted-statistics-for-data-scientists. We encourage you to use vibe coding to duplicate or extend the examples presented in the book. We do include some examples of prompting and the AI responses to illustrate overall strategies. We have not included comprehensive detailed instructions on using AI throughout the book, as there are different tools and they are evolving rapidly. Moreover, it is vital that you gain an understanding of the statistical and machine learning methods presented here. In the future, analytics and engineering jobs will go increasingly to those who understand enough of the methods and con‐ cepts to be able to properly structure the validation, testing, and oversight of the AI tools that will increasingly be used. It will not be enough to simply submit a prompt to an AI and then put the response into production. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You Preface | xvii
Page
20
do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “AI-Assisted Statistics for Data Scien‐ tists by Peter Bruce, Andrew Bruce, and Peter Gedeck (O’Reilly). Copyright 2026 Peter Bruce, Andrew Bruce, and Peter Gedeck, 979-8-341-66628-3.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit http://oreilly.com. How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 141 Stony Circle, Suite 195 Santa Rosa, CA 95401 800-889-8969 (in the United States or Canada) 707-827-7019 (international or local) 707-829-0104 (fax) support@oreilly.com https://oreilly.com/about/contact.html xviii | Preface