Digital Library

Practical Recommender Systems (Kim Falk)(Z-Library)

Kim Falk

Practical Recommender Systems (Kim Falk)(Z-Library)

Author Kim Falk

technology
Language English

Recommender systems are practically a necessity for keeping a site's content current, useful, and interesting to visitors. Recommender systems are everywhere, helping you find everything from movies to jobs, restaurants to hospitals, even romance. Practical Recommender Systems goes behind the curtain to show readers how recommender systems work and, more importantly, how to create and apply them for their site. This hands-on guide covers scaling problems and other issues they may encounter as their site grows.

Format PDF
Size 14.8 MB
19
Views
0
Downloads
0.00
Total Donations
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
M A N N I N G Kim Falk
Page 2
Practical Recommender Systems
Page 3
(This page has no text content)
Page 4
Practical Recommender Systems KIM FALK M A N N I N G SHELTER ISLAND
Page 5
For online information and ordering of this and other Manning books, please visit www.manning.com. The publisher offers discounts on this book when ordered in quantity. For more information, please contact Special Sales Department Manning Publications Co. 20 Baldwin Road PO Box 761 Shelter Island, NY 11964 Email: orders@manning.com © 2019 by Manning Publications Co. All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted, in any form or by means electronic, mechanical, photocopying, or otherwise, without prior written permission of the publisher. Many of the designations used by manufacturers and sellers to distinguish their products are claimed as trademarks. Where those designations appear in the book, and Manning Publications was aware of a trademark claim, the designations have been printed in initial caps or all caps. Recognizing the importance of preserving what has been written, it is Manning’s policy to have the books we publish printed on acid-free paper, and we exert our best efforts to that end. Recognizing also our responsibility to conserve the resources of our planet, Manning books are printed on paper that is at least 15 percent recycled and processed without the use of elemental chlorine. Manning Publications Co. Development editor: Helen Stergius 20 Baldwin Road Production editor: Janet Vail PO Box 761 Copy editors: Katie Petito and Frances Buran Shelter Island, NY 11964 Proofreader: Elizabeth Martin Technical proofreaders: Valentin Crettaz and Furkan Kamaci Typesetter: Dottie Marsico Cover designer: Marija Tudor ISBN 9781617292705 Printed in the United States of America 1 2 3 4 5 6 7 8 9 10 – SP – 24 23 22 21 20 19
Page 6
To the loves of my life: my wife, Sara, and my son, Peter, the small Superhero
Page 7
(This page has no text content)
Page 8
brief contents PART 1 GETTING READY FOR RECOMMENDER SYSTEMS................... 1 1 ■ What is a recommender? 3 2 ■ User behavior and how to collect it 30 3 ■ Monitoring the system 57 4 ■ Ratings and how to calculate them 77 5 ■ Non-personalized recommendations 102 6 ■ The user (and content) who came in from the cold 128 PART 2 RECOMMENDER ALGORITHMS........................................ 149 7 ■ Finding similarities among users and among content 151 8 ■ Collaborative filtering in the neighborhood 181 9 ■ Evaluating and testing your recommender 211 10 ■ Content-based filtering 248 11 ■ Finding hidden genres with matrix factorization 284 12 ■ Taking the best of all algorithms: implementing hybrid recommenders 329 13 ■ Ranking and learning to rank 357 14 ■ Future of recommender systems 384vii
Page 9
(This page has no text content)
Page 10
contents preface xvii acknowledgments xix about this book xx about the author xxiii about the cover illustration xxiv PART 1 GETTING READY FOR RECOMMENDER SYSTEMS....1 1 What is a recommender? 3 1.1 Real-life recommendations 3 Recommender systems are at home on the internet 5 ■ The long tail 5 ■ The Netflix recommender system 6 ■ Recommender system definition 12 1.2 Taxonomy of recommender systems 15 Domain 15 ■ Purpose 16 ■ Context 16 ■ Personalization level 17 ■ Whose opinions 18 ■ Privacy and trustworthiness 18 ■ Interface 19 ■ Algorithms 22 1.3 Machine learning and the Netflix Prize 23 1.4 The MovieGEEKs website 24 Design and specification 26 ■ Architecture 26 1.5 Building a recommender system 28ix
Page 11
CONTENTSx2 User behavior and how to collect it 30 2.1 How (I think) Netflix gathers evidence while you browse 31 The evidence Netflix collects 33 2.2 Finding useful user behavior 35 Capturing visitor impressions 35 ■ What you can learn from a shop browser 36 ■ Act of buying 40 ■ Consuming products 41 Visitor ratings 42 ■ Getting to know your customers the (old) Netflix way 45 2.3 Identifying users 46 2.4 Getting visitor data from other sources 46 2.5 The collector 47 Building the project files 48 ■ The data model 48 The snitch: Client-side evidence collector 49 ■ Integrating the collector into MovieGEEKs 50 2.6 What users in the system are and how to model them 52 3 Monitoring the system 57 3.1 Why adding a dashboard is a good idea 58 Answering “How are we doing?” 58 3.2 Doing the analytics 60 Web analytics 60 ■ The basic statistics 60 ■ Conversions 61 Analyzing the path up to conversion 64 ■ Conversion path 66 3.3 Personas 68 3.4 MovieGEEKs dashboard 71 Auto-generating data to your log 71 ■ Specification and design of the analytics dashboard 72 ■ Analytics dashboard wireframe 72 ■ Architecture 73 4 Ratings and how to calculate them 77 4.1 User-item preferences 78 Definition of ratings 78 ■ User-item matrix 79 4.2 Explicit or implicit ratings 81 How we use trusted sources for recommendations 82 4.3 Revisiting explicit ratings 83 4.4 What are implicit ratings? 83 People suggestions 85 ■ Considerations of calculating ratings 85
Page 12
CONTENTS xi4.5 Calculating implicit ratings 88 Looking at the behavioral data 89 ■ This could be considered a machine learning problem 93 4.6 How to implement implicit ratings 93 Adding the time aspect 97 4.7 Less frequent items provide more value 99 5 Non-personalized recommendations 102 5.1 What’s a non-personalized recommendation? 103 What’s a commercial? 103 ■ What does a recommendation do? 105 5.2 How to make recommendations when you have no data 105 Top 10: A chart of items 107 5.3 Implementing the chart and the groundwork for the recommender system component 108 The recommender system component 108 ■ MovieGEEKs code from GitHub 110 ■ A recommender system 110 ■ Adding a chart to MovieGEEKs 110 ■ Making the content look more attractive 111 5.4 Seeded recommendations 113 Frequently bought items similar to the one you’re viewing 114 Association rules 115 ■ Implementing association rules 120 Saving the association rules in the database 123 ■ Running the association rules calculator 124 ■ Using different events to create the association rules 126 6 The user (and content) who came in from the cold 128 6.1 What’s a cold start? 128 Cold products 130 ■ A cold visitor 130 ■ Gray sheep 132 Let’s look at real-life examples 132 ■ What can you do about cold starts? 133 6.2 Keeping track of visitors 134 Persisting anonymous users 134 6.3 Addressing cold-start problems with algorithms 134 Using association rules to create recs for cold users 135 Using domain knowledge and business rules 136 ■ Using segments 137 ■ Using categories to get around the gray sheep problem and how to introduce cold product 139
Page 13
CONTENTSxii6.4 Those who doesn’t ask, won’t know 140 When the visitor is no longer new 141 6.5 Using association rules to start recommending things fast 142 Find the collected items 143 ■ Retrieve association rules and order them according to confidence 143 ■ Displaying the recs 144 ■ Implementation evaluation 147 PART 2 RECOMMENDER ALGORITHMS.........................149 7 Finding similarities among users and among content 151 7.1 Why similarity? 152 What’s a similarity function? 153 7.2 Essential similarity functions 153 Jaccard distance 155 ■ Measuring distance with Lp-norms 156 Cosine similarity 159 ■ Finding similarity with Pearson’s correlation coefficient 162 ■ Test running a Pearson similarity 163 ■ Pearson correlation is similar to cosine 165 7.3 k-means clustering 165 The k-means clustering algorithm 166 ■ Translating k-means clustering into Python 168 7.4 Implementi5ng similarities 172 Implementing the similarity in the MovieGEEKs site 174 Implementing the clustering in the MovieGEEKs site 177 8 Collaborative filtering in the neighborhood 181 8.1 Collaborative filtering: A history lesson 183 When information became collaboratively filtered 183 Helping each other 183 ■ The rating matrix 185 ■ The collaborative filtering pipeline 186 ■ Should you use user-user or item-item collaborative filtering? 186 ■ Data requirements 187 8.2 Calculating recommendations 188 8.3 Calculating similarities 188 8.4 Amazon’s algorithm to precalculate item similarity 189 8.5 Ways to select the neighborhood 194 8.6 Finding the right neighborhood 195 8.7 Ways to calculate predicted ratings 196 8.8 Prediction with item-based filtering 197 Computing item predictions 198
Page 14
CONTENTS xiii8.9 Cold-start problems 199 8.10 A few words on machine learning terms 199 8.11 Collaborative filtering on the MovieGEEKs site 200 Item-based filtering 202 8.12 What’s the difference between association rule recs and collaborative recs? 207 8.13 Levers to fiddle with for collaborative filtering 207 8.14 Pros and cons of collaborative filtering 209 9 Evaluating and testing your recommender 211 9.1 Business wants lift, cross-sales, up-sales, and conversions 212 9.2 Why is it important to evaluate? 213 9.3 How to interpret user behavior 214 9.4 What to measure 214 Understanding my taste: Minimizing prediction error 216 Diversity 216 ■ Coverage 217 ■ Serendipity 219 9.5 Before implementing the recommender… 219 Verify the algorithm 220 ■ Regression testing 221 9.6 Types of evaluation 222 9.7 Offline evaluation 222 What to do when the algorithm doesn’t produce any recommendations 223 9.8 Offline experiments 223 Preparing the data for the experiment 229 9.9 Implementing the experiment in MovieGEEKs 235 The to-do list 235 9.10 Evaluating the test set 239 Starting out with the baseline predictor 239 ■ Finding the right parameters 242 9.11 Online evaluation 243 Controlled experiments 243 ■ A/B testing 244 9.12 Continuous testing with exploit/explore 245 Feedback loops 246
Page 15
CONTENTSxiv10 Content-based filtering 248 10.1 Descriptive example 249 10.2 Content-based filtering 251 10.3 Content analyzer 253 Feature extraction for the item profile 253 ■ Categorical data with small numbers 255 ■ Converting the year to a comparable feature 255 10.4 Extracting metadata from descriptions 256 Preparing descriptions 256 10.5 Finding important words with TF-IDF 260 10.6 Topic modeling using the LDA 261 What knobs can you turn to tweak the LDA? 268 10.7 Finding similar content 271 10.8 Creating the user profile 272 Creating the user profile with LDA 272 ■ Creating the user profile with TF-IDF 272 10.9 Content-based recommendations in MovieGEEKs 274 Loading data 274 ■ Training the model 277 ■ Creating item profiles 278 ■ Creating user profiles 278 ■ Showing recommendations 280 10.10 Evaluation of the content-based recommender 281 10.11 Pros and cons of content-based filtering 282 11 Finding hidden genres with matrix factorization 284 11.1 Sometimes it’s good to reduce the amount of data 285 11.2 Example of what you want to solve 287 11.3 A whiff of linear algebra 290 Matrix 290 ■ What’s factorization? 292 11.4 Constructing the factorization using SVD 293 Adding a new user by folding in 299 ■ How to do recommendations with SVD 301 ■ Baseline predictors 302 Temporal dynamic 304 11.5 Constructing the factorization using Funk SVD 305 Root Mean Squared Error 305 ■ Gradient descent 306 Stochastic gradient descent 309 ■ And finally, to the factorization 309 ■ Adding biases 310 ■ How to start and when to stop 311
Page 16
CONTENTS xv11.6 Doing recommendations with Funk SVD 315 11.7 Funk SVD implementation in MovieGEEKs 318 What to do with outliers 322 ■ Keeping the model up to date 324 ■ Faster implementation 324 11.8 Explicit vs. implicit data 324 11.9 Evaluation 324 11.10 Levers to fiddle with for Funk SVD 326 12 Taking the best of all algorithms: Implementing hybrid recommenders 329 12.1 The confused world of hybrids 330 12.2 The monolithic 331 Mixing content-based features with behavioral data to improve collaborative filtering recommenders 332 12.3 Mixed hybrid recommender 333 12.4 The ensemble 334 Switched ensemble recommender 335 ■ Weighted ensemble recommender 336 ■ Linear regression 337 12.5 Feature-weighted linear stacking (FWLS) 338 Meta features: Weights as functions 339 ■ The algorithm 341 12.6 Implementation 348 13 Ranking and learning to rank 357 13.1 Learning to rank an example at Foursquare 358 13.2 Re-ranking 362 13.3 What’s learning to rank again? 363 The three types of LTR algorithms 363 13.4 Bayesian Personalized Ranking 365 Ranking with BPR 368 ■ Math magic (advanced wizardry) 369 ■ The BPR algorithm 372 ■ BPR with matrix factorization 373 13.5 Implementation of BPR 373 Doing the recommendations 378 13.6 Evaluation 380 13.7 Levers to fiddle with for BPR 382
Page 17
CONTENTSxvi14 Future of recommender systems 384 14.1 This book in a few sentences 385 14.2 Topics to study next 388 Further reading 388 ■ Algorithms 389 ■ Context 389 Human-computer interactions 390 ■ Choosing a good architecture 390 14.3 What’s the future of recommender systems? 391 14.4 Final thoughts 395 index 397
Page 18
preface When I finished university in 2003, it was with the threat that no computer scientists would be needed in Europe because everything would be developed in countries where salaries were much lower. That never materialized, thank goodness, for many reasons. I’d venture that one of the larger issues was that companies underestimated the problem of developers not understanding the culture where their software was going to run. Software requests were implemented, but the functionality was different from what customers expected. Today, there’s a similar menace for people interested in machine learning and data science. But now the threat is not low salaries, but software as a service (SaaS), where you upload data and then the system does the work for you. I’m as concerned as anyone else that machines don’t understand domains and people. Machines aren’t intelligent enough yet that you can take humans out of the equation. Things are moving quickly, but I venture that anyone who is reading this book will be able to work with recommenders until the end of their career. Where did I drop into the mix? I was working as a software engineer in Italy and was moving to England and needed a job that required more thought than doing CRUD operations on a database. Luckily, I was contacted by a great recruiter from RedRock Consulting Ltd. They matched me with a recommender system provider, where I worked on the engine. And that was it; I was lost in machine learning (“lost” in the sense of being really interested and engaged). In addition to working on rec- ommender systems, I also started trawling for knowledge on the internet and read myriad books on the subject and related topics. Today you can’t throw a stick without having at least 10 people try to teach you something about machine learning. I find it amusing when I see one-page orxvii
Page 19
PREFACExviiione-hour tutorials that claim to teach you all you need to know about machine learn- ing. I can create a similarly effective tutorial on how to be a fighter pilot: You take off and you fly using the stick. If you need to shoot, you press a button. Then, you land before you run out of gas. A fighter pilot tutorial like that will probably be great to get you started—it’s where I started. But don’t fool yourself: understanding machine learning is complex. Add to that the human factor, which always makes things a bit wobblier. To get back to my story, I worked with recommenders and was happy about it, and then I changed jobs. In my new position, I was supposed to continue working on rec- ommender systems, but that project was delayed. At that point I was nervous I wouldn’t be working with recommenders anymore, but that was when Manning offered me the opportunity to write a book about recommender systems. What could I do, other than jump at the task? Immediately after I signed the contract, the recom- mender project started after all. Writing this book has been a great learning experi- ence, and I hope you’ll benefit from and enjoy it. The goal of the book is to introduce you to recommender systems—not only the algorithms, but also the recommender system ecosystem. The algorithms aren’t too complex, but to understand and run them requires understanding the users who are to receive the recommendations. The book’s contents have evolved during writing, because I’ve tried to fit more and more in. I hope reading this book will provide every- thing you need to know to get started on recommenders and give you a solid founda- tion to build on as you learn more.
Page 20
acknowledgments I want to mention and acknowledge two groups of people here: those who actively worked on the book and those who suffered and supported my constantly distracted presence for the last three years while this book project has been under way. It might be my name on the cover of Practical Recommender Systems, but this book couldn’t have come into existence without the great work of the people at Manning. I want especially to thank Helen Stergius for her relentless help and guidance as my development editor. She and all the others have translated my slightly dyslexic writing into something that teaches people to implement recommenders. I also want to thank Furkan Kamaci and Valentin Crettaz, my technical proofread- ers, and all the reviewers who took the time to read the early versions and helped the manuscript become more connected. They include Adhir Ramjiavan, Alexander Myltsev, Alvin Raj, Amit Lamba, Andrew Collier, Fazel Keshtkar, Jared Duncan, Jaromir Nemec, Martin Beer, Mayur Patil, Mike Dalrymple, Noreen Dertinger, Olivier Ducatteeuw, Peter Hampton, Simeon Leyzerzon, Søren Lind Kristiansen, Steven Parr, Tobias Bürger, Tobias Getrost, and Vipul Gupta. Many libraries, systems, and packages have been used to write this story, and I’m very grateful to the communities that helped me. I’m also thankful for the tools the open source communities have provided so everything didn’t need to be imple- mented from the ground up. Most important, I want to thank my wife, my son, my mother-in-law, the rest of my family, and my close friends for their support, love, and, most of all, patience. It hasn’t been easy for them to have a family member and friend who’s always sneaking away to write while moving to a new house and seeing our homes in Italy shaken to pieces by earthquakes. Not to mention that said writer started not one, but two, new jobs in the process. Thank you, and I promise no new projects for at least a couple of years. Love to all of you!xix
The above is a preview of the first 20 pages. Register to read the complete e-book.

Support Author

0.00
Total Amount (¥)
0
Donation Count
Please enter an amount Minimum ¥1

You will be redirected to Alipay to complete payment, then return here.

Recommended for You

Loading recommended books...
Failed to load, please try again later
Back to List