Page
1
(This page has no text content)
Page
2
BIRMINGHAM—MUMBAI Machine Learning for Streaming Data with Python Copyright © 2022 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. Publishing Product Manager: Dinesh Chaudhary Content Development Editor: Joseph Sunil Technical Editor: Rahul Limbachiya Copy Editor: Safis Editing Project Coordinator: Farheen Fathima Proofreader: Safis Editing Indexer: Sejal Dsilva Production Designer: Shankar Kalbhor Marketing Coordinator: Shifa Ansari and Abeer Riyaz Dawe
Page
3
First published: July 2022 Production reference: 1240622 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-80324-836-3 www.packt.com Contributors About the author Joos Korstanje, with his master's degrees in both environmental sciences and data science, has been working on statistics and data science for almost 10 years. Through his work in different companies including Disney, AXA, and others, he has closely followed developments in data science and related fields. This experience in the business world has allowed him to write about data science from an applied point of view (through his books, Medium, Towards Data Science, LinkedIn, and more). About the reviewer Olivia Petris is a big data engineer working as an IT consultant in a technology and advisory services firm based in Paris. On her professional journey, she's always looking for challenging and interesting assignments. Since her engineering diploma in computer science, she has chosen to be in the data science and big data field. Therefore, she continues to improve her skills and keep up to date with new IT and technology developments. In her free time, she enjoys traveling, practicing karate, and hanging out with her family and friends.
Page
4
Table of Contents Preface
Page
5
Part 1: Introduction and Core Concepts of Streaming Data
Page
6
Chapter 1 : An Introduction to Streaming Data Technical requirements Setting up a Python environment A short history of data science Working with streaming data Streaming data versus batch data Advantages of streaming data Examples of successful implementation of streaming analytics Challenges of streaming data How to get started with streaming data Common use cases for streaming data Streaming versus big data Real-time data formats and importing an example dataset in Python Summary Further reading
Page
7
Chapter 2 : Architectures for Streaming and Real- Time Machine Learning Technical requirements Python environment Defining your analytics as a function Understanding microservices architecture Communicating between services through APIs Demystifying the HTTP protocol The GET request The POST request JSON format for communication between systems RESTful APIs Building a simple API on AWS API Gateway in AWS Lambda in AWS Data-generating process on a local machine Implementing the example More architectural considerations Other AWS services and other services in general that have the same functionality Big data tools for real time streaming Calling a big data environment in real time Summary Further reading
Page
8
Chapter 3 : Data Analysis on Streaming Data Technical requirements Python environment Descriptive statistics on streaming data Why are descriptive statistics different on streaming data? Introduction to sampling theory Comparing population and sample Population parameters and sample statistics Sampling distribution Sample size calculations and confidence level Rolling descriptive statistics from streaming Exponential weight Tracking convergence as an additional KPI Overview of the main descriptive statistics The mean The median The mode Standard deviation Variance Quartiles and interquartile range Correlations Real-time visualizations Opening the dashboard Comparing Plotly's Dash and other real-time visualization tools
Page
9
Building basic alerting systems Alerting systems on extreme values Alerting systems on process stability (mean and median) Alerting systems on constant variability (std and variance) Basic alerting systems using statistical process control Summary Further reading
Page
10
Part 2: Exploring Use Cases for Data Streaming
Page
11
Chapter 4 : Online Learning with River Technical requirements Python environment What is online machine learning? How is online learning different from regular learning? Advantages of online learning Challenges of online learning Types of online learning Using River for online learning Training an online model with River Improving the model evaluation Building a multiclass classifier using one-vs-rest Summary Further reading
Page
12
Chapter 5 : Online Anomaly Detection Technical requirements Python environment Defining anomaly detection Are outliers a problem? Exploring use cases of anomaly detection Fraud detection in financial institutions Anomaly detection on your log data Fault detection in manufacturing and production lines Hacking detection in computer networks (cyber security) Medical risks in health data Predictive maintenance and sensor data Comparing anomaly detection and imbalanced classification The problem of imbalanced data The F1 score SMOTE oversampling Anomaly detection versus classification Algorithms for detecting anomalies in River The use of thresholders in River anomaly detection Anomaly detection algorithm 1 – One-Class SVM Anomaly detection algorithm 2 – Half-Space-Trees Going further with anomaly detection Summary Further reading
Page
13
Chapter 6 : Online Classif ication Technical requirements Python environment Defining classification Identifying use cases of classification Use case 1 – email spam classification Use case 2 – face detection in phone camera Use case 3 – online marketing ad selection Overview of classification algorithms in River Classification algorithm 1 – LogisticRegression Classification algorithm 2 – Perceptron Classification algorithm 3 – AdaptiveRandomForestClassifier Classification algorithm 4 – ALMAClassifier Classification algorithm 5 – PAClassifier Evaluating benchmark results Summary Further reading
Page
14
Chapter 7 : Online Regression Technical requirements Python environment Defining regression Use cases of regression Use case 1 – Forecasting Use case 2 – Predicting the number of faulty products in manufacturing Overview of regression algorithms in River Regression algorithm 1 – LinearRegression Regression algorithm 2 – HoeffdingAdaptiveTreeRegressor Regression algorithm 3 – SGTRegressor Regression algorithm 4 – SRPRegressor Summary Further reading
Page
15
Chapter 8 : Reinforcement Learning Technical requirements Python environment Defining reinforcement learning Comparing online and offline reinforcement learning A more detailed overview of feedback loops in reinforcement learning The main steps of a reinforcement learning model Making the decisions Updating the decision rules Exploring Q-learning The goal of Q-learning Parameters of the Q-learning algorithm Deep Q-learning Using reinforcement learning for streaming data Use cases of reinforcement learning Use case one – trading system Use case two – social network ranking system Use case three – a self-driving car Use case four – chatbots Use case five – learning games Implementing reinforcement learning in Python Summary Further reading
Page
16
Part 3: Advanced Concepts and Best Practices around Streaming Data
Page
17
Chapter 9 : Drif t and Drif t Detection Technical requirements Python environment Defining drift Three types of drift Introducing model explicability Measuring drift Measuring data drift Measuring concept drift Measuring drift in Python A basic intuitive approach to measuring drift Measuring drift with robust tools Counteracting drift Offline learning with retraining strategies against drift Online learning against drift Summary Further reading
Page
18
Chapter 10 : Feature Transformation and Scaling Technical requirements Python environment Challenges of data preparation with streaming data Scaling data for streaming Introducing scaling Adapting scaling to a streaming context Transforming features in a streaming context Introducing PCA Mathematical definition of PCA Regular PCA in Python Incremental PCA for streaming Summary Further reading
Page
19
Chapter 11 : Catastrophic Forgett ing Technical requirements Python environment Introducing catastrophic forgetting Catastrophic forgetting in online models Detecting catastrophic forgetting Using Python to detect catastrophic forgetting Model explicability versus catastrophic forgetting Explaining models using linear coefficients Explaining models using dendrograms Explaining models using variable importance Summary Further reading
Page
20
Chapter 12 : Conclusion and Best Practices Going further Summary Other Books You May Enjoy