Digital Library

Data Mining Practical Machine Learning Tools and Techniques, Fourth Edition (Ian H. Witten, Eibe Frank, Mark A. Hall etc.) (z-library.sk, 1lib.sk, z-lib.sk)

Ian H. Witten, Eibe Frank, Mark A. Hall, Christopher J. Pal

Data Mining Practical Machine Learning Tools and Techniques, Fourth Edition (Ian H. Witten, Eibe Frank, Mark A. Hall etc.) (z-library.sk, 1lib.sk, z-lib.sk)

Author Ian H. Witten, Eibe Frank, Mark A. Hall, Christopher J. Pal

data
Language English

Data Mining: Practical Machine Learning Tools and Techniques, Fourth Edition, offers a thorough grounding in machine learning concepts, along with practical advice on applying these tools and techniques in real-world data mining situations. This highly anticipated fourth edition of the most acclaimed work on data mining and machine learning teaches readers everything they need to know to get going, from preparing inputs, interpreting outputs, evaluating results, to the algorithmic methods at the heart of successful data mining approaches. Extensive updates reflect the technical changes and modernizations that have taken place in the field since the last edition, including substantial new chapters on probabilistic methods and on deep learning. Accompanying the book is a new version of the popular WEKA machine learning software from the University of Waikato. Authors Witten, Frank, Hall, and Pal include today's techniques coupled with the methods at the leading edge of contemporary research. Please visit the book companion website at contains Powerpoint slides for Chapters 1-12. This is a very comprehensive teaching resource, with many PPT slides covering each chapter of the book Online Appendix on the Weka workbench; again a very comprehensive learning aid for the open source software that goes with the book Table of contents, highlighting the many new sections in the 4th edition, along with reviews of the 1st edition, errata, etc. Provides a thorough grounding in machine learning concepts, as well as practical advice on applying the tools and techniques to data mining projects Presents concrete tips and techniques for performance improvement that work by transforming the input or output in machine learning methods Includes a downloadable WEKA software toolkit, a comprehensive collection of machine learning algorithms for data mining tasks-in an easy-to-use interactive interface Includes open-access online courses that introduce practical applications of the mate

Format PDF
Size 4.8 MB
8
Views
0
Downloads
0.00
Total Donations
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
Data Mining
Page 3
This page intentionally left blank
Page 4
Data Mining Practical Machine Learning Tools and Techniques Fourth Edition Ian H. Witten University of Waikato, Hamilton, New Zealand Eibe Frank University of Waikato, Hamilton, New Zealand Mark A. Hall University of Waikato, Hamilton, New Zealand Christopher J. Pal Polytechnique Montréal, and the Université de Montréal, Montreal, QC, Canada AMSTERDAM • BOSTON • HEIDELBERG • LONDON NEW YORK • OXFORD • PARIS • SAN DIEGO SAN FRANCISCO • SINGAPORE • SYDNEY • TOKYO Morgan Kaufmann is an imprint of Elsevier
Page 5
Morgan Kaufmann is an imprint of Elsevier 50 Hampshire Street, 5th Floor, Cambridge, MA 02139, United States Copyright © 2017, 2011, 2005, 2000 Elsevier Inc. All rights reserved. No part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or any information storage and retrieval system, without permission in writing from the publisher. Details on how to seek permission, further information about the Publisher’s permissions policies and our arrangements with organizations such as the Copyright Clearance Center and the Copyright Licensing Agency, can be found at our website: www.elsevier.com/permissions. This book and the individual contributions contained in it are protected under copyright by the Publisher (other than as may be noted herein). Notices Knowledge and best practice in this field are constantly changing. As new research and experience broaden our understanding, changes in research methods, professional practices, or medical treatment may become necessary. Practitioners and researchers must always rely on their own experience and knowledge in evaluating and using any information, methods, compounds, or experiments described herein. In using such information or methods they should be mindful of their own safety and the safety of others, including parties for whom they have a professional responsibility. To the fullest extent of the law, neither the Publisher nor the authors, contributors, or editors, assume any liability for any injury and/or damage to persons or property as a matter of products liability, negligence or otherwise, or from any use or operation of any methods, products, instructions, or ideas contained in the material herein. British Library Cataloguing-in-Publication Data A catalogue record for this book is available from the British Library Library of Congress Cataloging-in-Publication Data A catalog record for this book is available from the Library of Congress ISBN: 978-0-12-804291-5 For Information on all Morgan Kaufmann publications visit our website at https://www.elsevier.com Publisher: Todd Green Acquisition Editor: Tim Pitts Editorial Project Manager: Charlotte Kent Production Project Manager: Nicky Carter Designer: Matthew Limbert Typeset by MPS Limited, Chennai, India
Page 6
Contents List of Figures..........................................................................................................xv List of Tables..........................................................................................................xxi Preface ................................................................................................................. xxiii PART I INTRODUCTION TO DATA MINING CHAPTER 1 What’s it all about? ..........................................................3 1.1 Data Mining and Machine Learning..............................................4 Describing Structural Patterns ....................................................... 6 Machine Learning.......................................................................... 7 Data Mining ................................................................................... 9 1.2 Simple Examples: The Weather Problem and Others...................9 The Weather Problem.................................................................. 10 Contact Lenses: An Idealized Problem....................................... 12 Irises: A Classic Numeric Dataset .............................................. 14 CPU Performance: Introducing Numeric Prediction .................. 16 Labor Negotiations: A More Realistic Example......................... 16 Soybean Classification: A Classic Machine Learning Success ...................................................................................19 1.3 Fielded Applications ....................................................................21 Web Mining ................................................................................. 21 Decisions Involving Judgment .................................................... 22 Screening Images......................................................................... 23 Load Forecasting ......................................................................... 24 Diagnosis...................................................................................... 25 Marketing and Sales .................................................................... 26 Other Applications....................................................................... 27 1.4 The Data Mining Process.............................................................28 1.5 Machine Learning and Statistics..................................................30 1.6 Generalization as Search..............................................................31 Enumerating the Concept Space ................................................. 32 Bias .............................................................................................. 33 1.7 Data Mining and Ethics ...............................................................35 Reidentification............................................................................ 36 Using Personal Information......................................................... 37 Wider Issues................................................................................. 38 1.8 Further Reading and Bibliographic Notes ...................................38 v
Page 7
CHAPTER 2 Input: concepts, instances, attributes......................... 43 2.1 What’s a Concept? .......................................................................44 2.2 What’s in an Example?................................................................46 Relations ...................................................................................... 47 Other Example Types .................................................................. 51 2.3 What’s in an Attribute?................................................................53 2.4 Preparing the Input.......................................................................56 Gathering the Data Together ....................................................... 56 ARFF Format ............................................................................... 57 Sparse Data .................................................................................. 60 Attribute Types ............................................................................ 61 Missing Values ............................................................................ 62 Inaccurate Values......................................................................... 63 Unbalanced Data.......................................................................... 64 Getting to Know Your Data ........................................................ 65 2.5 Further Reading and Bibliographic Notes ...................................65 CHAPTER 3 Output: knowledge representation ............................... 67 3.1 Tables ...........................................................................................68 3.2 Linear Models ..............................................................................68 3.3 Trees .............................................................................................70 3.4 Rules .............................................................................................75 Classification Rules ..................................................................... 75 Association Rules ........................................................................ 79 Rules With Exceptions ................................................................ 80 More Expressive Rules................................................................ 82 3.5 Instance-Based Representation ....................................................84 3.6 Clusters .........................................................................................87 3.7 Further Reading and Bibliographic Notes ...................................88 CHAPTER 4 Algorithms: the basic methods ..................................... 91 4.1 Inferring Rudimentary Rules .......................................................93 Missing Values and Numeric Attributes ..................................... 94 4.2 Simple Probabilistic Modeling ....................................................96 Missing Values and Numeric Attributes ................................... 100 Naı̈ve Bayes for Document Classification ................................ 103 Remarks ..................................................................................... 105 4.3 Divide-and-Conquer: Constructing Decision Trees ..................105 Calculating Information............................................................. 108 Highly Branching Attributes ..................................................... 110 vi Contents
Page 8
4.4 Covering Algorithms: Constructing Rules ..............................113 Rules Versus Trees .................................................................. 114 A Simple Covering Algorithm ................................................ 115 Rules Versus Decision Lists.................................................... 119 4.5 Mining Association Rules........................................................120 Item Sets .................................................................................. 120 Association Rules .................................................................... 122 Generating Rules Efficiently ................................................... 124 4.6 Linear Models ..........................................................................128 Numeric Prediction: Linear Regression .................................. 128 Linear Classification: Logistic Regression ............................. 129 Linear Classification Using the Perceptron ............................ 131 Linear Classification Using Winnow ...................................... 133 4.7 Instance-Based Learning..........................................................135 The Distance Function............................................................. 135 Finding Nearest Neighbors Efficiently ................................... 136 Remarks ................................................................................... 141 4.8 Clustering .................................................................................141 Iterative Distance-Based Clustering ........................................ 142 Faster Distance Calculations ................................................... 144 Choosing the Number of Clusters ........................................... 146 Hierarchical Clustering............................................................ 147 Example of Hierarchical Clustering........................................ 148 Incremental Clustering............................................................. 150 Category Utility ....................................................................... 154 Remarks ................................................................................... 156 4.9 Multi-instance Learning...........................................................156 Aggregating the Input.............................................................. 157 Aggregating the Output ........................................................... 157 4.10 Further Reading and Bibliographic Notes...............................158 4.11 WEKA Implementations..........................................................160 CHAPTER 5 Credibility: evaluating what’s been learned............. 161 5.1 Training and Testing ..................................................................163 5.2 Predicting Performance..............................................................165 5.3 Cross-Validation.........................................................................167 5.4 Other Estimates ..........................................................................169 Leave-One-Out .......................................................................... 169 The Bootstrap............................................................................. 169 5.5 Hyperparameter Selection..........................................................171 viiContents
Page 9
5.6 Comparing Data Mining Schemes...........................................172 5.7 Predicting Probabilities............................................................176 Quadratic Loss Function.......................................................... 177 Informational Loss Function ................................................... 178 Remarks ................................................................................... 179 5.8 Counting the Cost ....................................................................179 Cost-Sensitive Classification ................................................... 182 Cost-Sensitive Learning........................................................... 183 Lift Charts ................................................................................ 183 ROC Curves ............................................................................. 186 Recall-Precision Curves........................................................... 190 Remarks ................................................................................... 190 Cost Curves.............................................................................. 192 5.9 Evaluating Numeric Prediction ...............................................194 5.10 The MDL Principle..................................................................197 5.11 Applying the MDL Principle to Clustering.............................200 5.12 Using a Validation Set for Model Selection ...........................201 5.13 Further Reading and Bibliographic Notes...............................202 PART II MORE ADVANCED MACHINE LEARNING SCHEMES CHAPTER 6 Trees and rules ............................................................. 209 6.1 Decision Trees............................................................................210 Numeric Attributes .................................................................... 210 Missing Values .......................................................................... 212 Pruning ....................................................................................... 213 Estimating Error Rates .............................................................. 215 Complexity of Decision Tree Induction.................................... 217 From Trees to Rules .................................................................. 219 C4.5: Choices and Options........................................................ 219 Cost-Complexity Pruning .......................................................... 220 Discussion .................................................................................. 221 6.2 Classification Rules....................................................................221 Criteria for Choosing Tests ....................................................... 222 Missing Values, Numeric Attributes ......................................... 223 Generating Good Rules ............................................................. 224 Using Global Optimization........................................................ 226 Obtaining Rules From Partial Decision Trees .......................... 227 Rules With Exceptions .............................................................. 231 Discussion .................................................................................. 233 viii Contents
Page 10
6.3 Association Rules.......................................................................234 Building a Frequent Pattern Tree .............................................. 235 Finding Large Item Sets ............................................................ 240 Discussion .................................................................................. 241 6.4 WEKA Implementations............................................................242 CHAPTER 7 Extending instance-based and linear models .......... 243 7.1 Instance-Based Learning............................................................244 Reducing the Number of Exemplars ......................................... 245 Pruning Noisy Exemplars.......................................................... 245 Weighting Attributes ................................................................. 246 Generalizing Exemplars............................................................. 247 Distance Functions for Generalized Exemplars........................ 248 Generalized Distance Functions ................................................ 250 Discussion .................................................................................. 250 7.2 Extending Linear Models...........................................................252 The Maximum Margin Hyperplane........................................... 253 Nonlinear Class Boundaries ...................................................... 254 Support Vector Regression........................................................ 256 Kernel Ridge Regression ........................................................... 258 The Kernel Perceptron............................................................... 260 Multilayer Perceptrons............................................................... 261 Radial Basis Function Networks ............................................... 270 Stochastic Gradient Descent...................................................... 270 Discussion .................................................................................. 272 7.3 Numeric Prediction With Local Linear Models........................273 Model Trees ............................................................................... 274 Building the Tree ....................................................................... 275 Pruning the Tree ........................................................................ 275 Nominal Attributes .................................................................... 276 Missing Values .......................................................................... 276 Pseudocode for Model Tree Induction...................................... 277 Rules From Model Trees........................................................... 281 Locally Weighted Linear Regression........................................ 281 Discussion .................................................................................. 283 7.4 WEKA Implementations............................................................284 CHAPTER 8 Data transformations .................................................... 285 8.1 Attribute Selection .....................................................................288 Scheme-Independent Selection.................................................. 289 Searching the Attribute Space ................................................... 292 Scheme-Specific Selection ........................................................ 293 ixContents
Page 11
8.2 Discretizing Numeric Attributes ................................................296 Unsupervised Discretization...................................................... 297 Entropy-Based Discretization.................................................... 298 Other Discretization Methods.................................................... 301 Entropy-Based Versus Error-Based Discretization................... 302 Converting Discrete to Numeric Attributes .............................. 303 8.3 Projections ..................................................................................304 Principal Component Analysis .................................................. 305 Random Projections................................................................... 307 Partial Least Squares Regression .............................................. 307 Independent Component Analysis............................................. 309 Linear Discriminant Analysis.................................................... 310 Quadratic Discriminant Analysis .............................................. 310 Fisher’s Linear Discriminant Analysis...................................... 311 Text to Attribute Vectors........................................................... 313 Time Series ................................................................................ 314 8.4 Sampling.....................................................................................315 Reservoir Sampling ................................................................... 315 8.5 Cleansing ....................................................................................316 Improving Decision Trees ......................................................... 316 Robust Regression ..................................................................... 317 Detecting Anomalies ................................................................. 318 One-Class Learning ................................................................... 319 Outlier Detection ....................................................................... 320 Generating Artificial Data ......................................................... 321 8.6 Transforming Multiple Classes to Binary Ones ........................322 Simple Methods ......................................................................... 323 Error-Correcting Output Codes ................................................. 324 Ensembles of Nested Dichotomies............................................ 326 8.7 Calibrating Class Probabilities...................................................328 8.8 Further Reading and Bibliographic Notes .................................331 8.9 WEKA Implementations............................................................334 CHAPTER 9 Probabilistic methods .................................................. 335 9.1 Foundations ................................................................................336 Maximum Likelihood Estimation ............................................. 338 Maximum a Posteriori Parameter Estimation ........................... 339 9.2 Bayesian Networks.....................................................................339 Making Predictions .................................................................... 340 x Contents
Page 12
Learning Bayesian Networks .................................................... 344 Specific Algorithms ................................................................... 347 Data Structures for Fast Learning ............................................. 349 9.3 Clustering and Probability Density Estimation .........................352 The Expectation Maximization Algorithm for a Mixture of Gaussians .........................................................................353 Extending the Mixture Model ................................................... 356 Clustering Using Prior Distributions......................................... 358 Clustering With Correlated Attributes ...................................... 359 Kernel Density Estimation ........................................................ 361 Comparing Parametric, Semiparametric and Nonparametric Density Models for Classification .......................................362 9.4 Hidden Variable Models ............................................................363 Expected Log-Likelihoods and Expected Gradients................. 364 The Expectation Maximization Algorithm ............................... 365 Applying the Expectation Maximization Algorithm to Bayesian Networks ..........................................................366 9.5 Bayesian Estimation and Prediction ..........................................367 Probabilistic Inference Methods................................................ 368 9.6 Graphical Models and Factor Graphs........................................370 Graphical Models and Plate Notation ....................................... 371 Probabilistic Principal Component Analysis............................. 372 Latent Semantic Analysis .......................................................... 376 Using Principal Component Analysis for Dimensionality Reduction .............................................................................377 Probabilistic LSA....................................................................... 378 Latent Dirichlet Allocation........................................................ 379 Factor Graphs............................................................................. 382 Markov Random Fields ............................................................. 385 Computing Using the Sum-Product and Max-Product Algorithms ...........................................................................386 9.7 Conditional Probability Models.................................................392 Linear and Polynomial Regression as Probability Models..................................................................................392 Using Priors on Parameters ....................................................... 393 Multiclass Logistic Regression.................................................. 396 Gradient Descent and Second-Order Methods.......................... 400 Generalized Linear Models ....................................................... 400 Making Predictions for Ordered Classes................................... 402 Conditional Probabilistic Models Using Kernels...................... 402 xiContents
Page 13
9.8 Sequential and Temporal Models ............................................403 Markov Models and N-gram Methods .................................... 403 Hidden Markov Models........................................................... 404 Conditional Random Fields ..................................................... 406 9.9 Further Reading and Bibliographic Notes...............................410 Software Packages and Implementations................................ 414 9.10 WEKA Implementations..........................................................416 CHAPTER 10 Deep learning .................................................. 417 10.1 Deep Feedforward Networks ...................................................420 The MNIST Evaluation ........................................................... 421 Losses and Regularization....................................................... 422 Deep Layered Network Architecture ...................................... 423 Activation Functions................................................................ 424 Backpropagation Revisited...................................................... 426 Computation Graphs and Complex Network Structures ........ 429 Checking Backpropagation Implementations ......................... 430 10.2 Training and Evaluating Deep Networks ................................431 Early Stopping ......................................................................... 431 Validation, Cross-Validation, and Hyperparameter Tuning ... 432 Mini-Batch-Based Stochastic Gradient Descent ..................... 433 Pseudocode for Mini-Batch Based Stochastic Gradient Descent.................................................................................434 Learning Rates and Schedules................................................. 434 Regularization With Priors on Parameters.............................. 435 Dropout .................................................................................... 436 Batch Normalization................................................................ 436 Parameter Initialization............................................................ 436 Unsupervised Pretraining......................................................... 437 Data Augmentation and Synthetic Transformations............... 437 10.3 Convolutional Neural Networks ..............................................437 The ImageNet Evaluation and Very Deep Convolutional Networks ..............................................................................438 From Image Filtering to Learnable Convolutional Layers..... 439 Convolutional Layers and Gradients....................................... 443 Pooling and Subsampling Layers and Gradients .................... 444 Implementation ........................................................................ 445 10.4 Autoencoders............................................................................445 Pretraining Deep Autoencoders With RBMs.......................... 448 Denoising Autoencoders and Layerwise Training.................. 448 Combining Reconstructive and Discriminative Learning....... 449 xii Contents
Page 14
10.5 Stochastic Deep Networks .......................................................449 Boltzmann Machines ............................................................... 449 Restricted Boltzmann Machines.............................................. 451 Contrastive Divergence ........................................................... 452 Categorical and Continuous Variables.................................... 452 Deep Boltzmann Machines...................................................... 453 Deep Belief Networks ............................................................. 455 10.6 Recurrent Neural Networks .....................................................456 Exploding and Vanishing Gradients ....................................... 457 Other Recurrent Network Architectures ................................. 459 10.7 Further Reading and Bibliographic Notes...............................461 10.8 Deep Learning Software and Network Implementations........464 Theano...................................................................................... 464 Tensor Flow ............................................................................. 464 Torch ........................................................................................ 465 Computational Network Toolkit.............................................. 465 Caffe......................................................................................... 465 Deeplearning4j......................................................................... 465 Other Packages: Lasagne, Keras, and cuDNN........................ 465 10.9 WEKA Implementations..........................................................466 CHAPTER 11 Beyond supervised and unsupervised learning ..... 467 11.1 Semisupervised Learning.........................................................468 Clustering for Classification.................................................... 468 Cotraining ................................................................................ 470 EM and Cotraining .................................................................. 471 Neural Network Approaches ................................................... 471 11.2 Multi-instance Learning...........................................................472 Converting to Single-Instance Learning ................................. 472 Upgrading Learning Algorithms ............................................. 475 Dedicated Multi-instance Methods.......................................... 475 11.3 Further Reading and Bibliographic Notes...............................477 11.4 WEKA Implementations..........................................................478 CHAPTER 12 Ensemble learning...................................................... 479 12.1 Combining Multiple Models....................................................480 12.2 Bagging ....................................................................................481 BiasVariance Decomposition ............................................... 482 Bagging With Costs................................................................. 483 12.3 Randomization .........................................................................484 Randomization Versus Bagging .............................................. 485 Rotation Forests ....................................................................... 486 xiiiContents
Page 15
12.4 Boosting ...................................................................................486 AdaBoost.................................................................................. 487 The Power of Boosting............................................................ 489 12.5 Additive Regression.................................................................490 Numeric Prediction .................................................................. 491 Additive Logistic Regression .................................................. 492 12.6 Interpretable Ensembles...........................................................493 Option Trees ............................................................................ 494 Logistic Model Trees............................................................... 496 12.7 Stacking....................................................................................497 12.8 Further Reading and Bibliographic Notes...............................499 12.9 WEKA Implementations..........................................................501 CHAPTER 13 Moving on: applications and beyond..................... 503 13.1 Applying Machine Learning..................................................504 13.2 Learning From Massive Datasets ..........................................506 13.3 Data Stream Learning............................................................509 13.4 Incorporating Domain Knowledge ........................................512 13.5 Text Mining ...........................................................................515 Document Classification and Clustering............................... 516 Information Extraction........................................................... 517 Natural Language Processing ................................................ 518 13.6 Web Mining ...........................................................................519 Wrapper Induction ................................................................. 519 Page Rank .............................................................................. 520 13.7 Images and Speech ................................................................522 Images .................................................................................... 523 Speech .................................................................................... 524 13.8 Adversarial Situations............................................................524 13.9 Ubiquitous Data Mining ........................................................527 13.10 Further Reading and Bibliographic Notes ............................529 13.11 WEKA Implementations .......................................................532 Appendix A: Theoretical foundations...................................................................533 Appendix B: The WEKA workbench ...................................................................553 References..............................................................................................................573 Index ......................................................................................................................601 xiv Contents
Page 16
List of Figures Figure 1.1 Rules for the contact lens data. 13 Figure 1.2 Decision tree for the contact lens data. 14 Figure 1.3 Decision trees for the labor negotiations data. 18 Figure 1.4 Life cycle of a data mining project. 29 Figure 2.1 A family tree and two ways of expressing the sister-of relation. 48 Figure 2.2 ARFF file for the weather data. 58 Figure 2.3 Multi-instance ARFF file for the weather data. 60 Figure 3.1 A linear regression function for the CPU performance data. 69 Figure 3.2 A linear decision boundary separating Iris setosas from Iris versicolors. 70 Figure 3.3 Constructing a decision tree interactively: (A) creating a rectangular test involving petallength and petalwidth; (B) the resulting (unfinished) decision tree. 73 Figure 3.4 Models for the CPU performance data: (A) linear regression; (B) regression tree; (C) model tree. 74 Figure 3.5 Decision tree for a simple disjunction. 76 Figure 3.6 The exclusive-or problem. 77 Figure 3.7 Decision tree with a replicated subtree. 77 Figure 3.8 Rules for the iris data. 81 Figure 3.9 The shapes problem. 82 Figure 3.10 Different ways of partitioning the instance space. 86 Figure 3.11 Different ways of representing clusters. 88 Figure 4.1 Pseudocode for 1R. 93 Figure 4.2 Tree stumps for the weather data. 106 Figure 4.3 Expanded tree stumps for the weather data. 108 Figure 4.4 Decision tree for the weather data. 109 Figure 4.5 Tree stump for the ID code attribute. 111 Figure 4.6 Covering algorithm: (A) covering the instances; (B) decision tree for the same problem. 113 Figure 4.7 The instance space during operation of a covering algorithm. 115 Figure 4.8 Pseudocode for a basic rule learner. 118 xv
Page 17
Figure 4.9 (A) Finding all item sets with sufficient coverage; (B) finding all sufficiently accurate association rules for a k-item set. 127 Figure 4.10 Logistic regression: (A) the logit transform; (B) example logistic regression function. 130 Figure 4.11 The perceptron: (A) learning rule; (B) representation as a neural network. 132 Figure 4.12 The Winnow algorithm: (A) unbalanced version; (B) balanced version. 134 Figure 4.13 A kD-tree for four training instances: (A) the tree; (B) instances and splits. 137 Figure 4.14 Using a kD-tree to find the nearest neighbor of the star. 137 Figure 4.15 Ball tree for 16 training instances: (A) instances and balls; (B) the tree. 139 Figure 4.16 Ruling out an entire ball (gray) based on a target point (star) and its current nearest neighbor. 140 Figure 4.17 Iterative distance-based clustering. 143 Figure 4.18 A ball tree: (A) two cluster centers and their dividing line; (B) corresponding tree. 145 Figure 4.19 Hierarchical clustering displays. 149 Figure 4.20 Clustering the weather data. 151 Figure 4.21 Hierarchical clusterings of the iris data. 153 Figure 5.1 A hypothetical lift chart. 185 Figure 5.2 Analyzing the expected benefit of a mailing campaign when the cost of mailing is (A) $0.50 and (B) $0.80. 187 Figure 5.3 A sample ROC curve. 188 Figure 5.4 ROC curves for two learning schemes. 189 Figure 5.5 Effect of varying the probability threshold: (A) error curve; (B) cost curve. 193 Figure 6.1 Example of subtree raising, where node C is “raised” to subsume node B. 214 Figure 6.2 Pruning the labor negotiations decision tree. 216 Figure 6.3 Algorithm for forming rules by incremental reduced- error pruning. 226 Figure 6.4 RIPPER: (A) algorithm for rule learning; (B) meaning of symbols. 228 Figure 6.5 Algorithm for expanding examples into a partial tree. 229 Figure 6.6 Example of building a partial tree. 230 Figure 6.7 Rules with exceptions for the iris data. 232 xvi List of Figures
Page 18
Figure 6.8 Extended prefix trees for the weather data: (A) the full data; (B) the data conditional on temperature5mild; (C) the data conditional on humidity5 normal. 238 Figure 7.1 A boundary between two rectangular classes. 249 Figure 7.2 A maximum margin hyperplane. 253 Figure 7.3 Support vector regression: (A) ε5 1; (B) ε5 2; (C) ε5 0.5. 257 Figure 7.4 Example data sets and corresponding perceptrons. 262 Figure 7.5 Step vs sigmoid: (A) step function; (B) sigmoid function. 264 Figure 7.6 Gradient descent using the error function w21 1. 265 Figure 7.7 Multilayer perceptron with a hidden layer (omitting bias inputs). 267 Figure 7.8 Hinge, squared and 01 loss functions. 271 Figure 7.9 Pseudocode for model tree induction. 278 Figure 7.10 Model tree for a data set with nominal attributes. 279 Figure 8.1 Attribute space for the weather dataset. 292 Figure 8.2 Discretizing the temperature attribute using the entropy method. 299 Figure 8.3 The result of discretizing the temperature attribute. 299 Figure 8.4 Class distribution for a two-class, two-attribute problem. 302 Figure 8.5 Principal component transform of a dataset: (A) variance of each component; (B) variance plot. 306 Figure 8.6 Comparing principal component analysis and Fisher’s linear discriminant analysis. 312 Figure 8.7 Number of international phone calls from Belgium, 19501973. 318 Figure 8.8 Overoptimistic probability estimation for a two-class problem. 329 Figure 9.1 A simple Bayesian network for the weather data. 341 Figure 9.2 Another Bayesian network for the weather data. 342 Figure 9.3 The Markov blanket for variable x6 in a 10-variable Bayesian network. 348 Figure 9.4 The weather data: (A) reduced version; (B) corresponding AD tree. 350 Figure 9.5 A two-class mixture model. 354 Figure 9.6 DensiTree showing possible hierarchical clusterings of a given data set. 360 Figure 9.7 Probability contours for three types of model, all based on Gaussians. 362 xviiList of Figures
Page 19
Figure 9.8 (A) Bayesian network for a mixture model; (B) multiple copies of the Bayesian network, one for each observation; (C) plate notation version of (B). 371 Figure 9.9 (A) Bayesian network for probabilistic PCA; (B) equal- probability contour for a Gaussian distribution along with its covariance matrix’s principal eigenvector. 372 Figure 9.10 The singular value decomposition of a t by d matrix. 377 Figure 9.11 Graphical models for (A) pLSA, (B) LDAb, and (C) smoothed LDAb. 379 Figure 9.12 (A) Bayesian network and (B) corresponding factor graph. 382 Figure 9.13 The Markov blanket for variable x6 in a 10-variable factor graph. 383 Figure 9.14 (A) and (B) Bayesian network and corresponding factor graph; (C) and (D) Naı̈ve Bayes model and corresponding factor graph. 384 Figure 9.15 (A) Bayesian network representing the joint distribution of y and its parents; (B) factor graph for a logistic regression for the conditional distribution of y given its parents. 384 Figure 9.16 (A) Undirected graph representing a Markov random field structure; (B) corresponding factor graph. 385 Figure 9.17 Message sequence in an example factor graph. 389 Figure 9.18 (A) and (B) First- and second-order Markov models for a sequence of variables; (C) Hidden Markov model; (D) Markov random field. 404 Figure 9.19 Mining emails for meeting details. 406 Figure 9.20 (A) Dynamic Bayesian network representation of a hidden Markov model; (B) similarly structured Markov random field; (C) factor graph for (A); and (D) factor graph for a linear chain conditional random field. 407 Figure 10.1 A feedforward neural network. 424 Figure 10.2 Computation graph showing forward propagation in a deep network. 426 Figure 10.3 Backpropagation in a deep network (the forward computation is shown with gray arrows). 429 Figure 10.4 Parameter updates that follow the forward and backward propagation steps (shown with gray arrows). 430 Figure 10.5 Typical learning curves for the training and validation sets. 431 xviii List of Figures
Page 20
Figure 10.6 Pseudocode for mini-batch based stochastic gradient descent. 435 Figure 10.7 Typical convolutional neural network architecture. 439 Figure 10.8 Original image; filtered with the two Sobel operators; magnitude of the result. 441 Figure 10.9 Examples of what random neurons detect in different layers of a convolutional neural network using the visualization approach of Zeiler and Fergus (2013). Underlying imagery kindly provided by Matthew Zeiler. 442 Figure 10.10 Example of the convolution, pooling, and decimation operations used in convolutional neural networks. 443 Figure 10.11 A simple autoencoder. 445 Figure 10.12 A deep autoencoder with multiple layers of transformation. 447 Figure 10.13 Low-dimensional principal component space (left) compared with one learned by a deep autoencoder (right). 447 Figure 10.14 Boltzmann machines: (A) fully connected; (B) restricted; (C) more general form of (B). 450 Figure 10.15 (A) Deep Boltzmann machine and (B) deep belief network. 453 Figure 10.16 (A) Feedforward network transformed into a recurrent network; (B) hidden Markov model; and (C) recurrent network obtained by unwrapping (A). 456 Figure 10.17 Structure of a “long short-term memory” unit. 459 Figure 10.18 Recurrent neural networks: (A) bidirectional, (B) encoder-decoder. 460 Figure 10.19 A deep encoder-decoder recurrent network. 460 Figure 12.1 Algorithm for bagging. 483 Figure 12.2 Algorithm for boosting. 488 Figure 12.3 Algorithm for additive logistic regression. 493 Figure 12.4 Simple option tree for the weather data. 494 Figure 12.5 Alternating decision tree for the weather data. 495 Figure 13.1 A tangled web. 521 xixList of Figures
The above is a preview of the first 20 pages. Register to read the complete e-book.

Support Author

0.00
Total Amount (¥)
0
Donation Count
Please enter an amount Minimum ¥1

You will be redirected to Alipay to complete payment, then return here.

Recommended for You

Loading recommended books...
Failed to load, please try again later
Back to List