Share E-Book

Practical Python Data Wrangling and Data Quality Getting Started with Reading, Cleaning, and Analyzing Data (Susan E. McGregor)(Z-Library)

Author Susan E. McGregor

Python
Language English

No Description

Format PDF
Size 13.4 MB
89
Views
0
Downloads
0.00
Total Donations
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
(This page has no text content)
Page 3
Susan E. McGregor Practical Python Data Wrangling and Data Quality Boston Farnham Sebastopol TokyoBeijing
Page 4
978-1-492-09150-9 [LSI] Practical Python Data Wrangling and Data Quality by Susan E. McGregor Copyright © 2022 Susan E. McGregor. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Jessica Haberman Development Editor: Jeff Bleiel Production Editor: Daniel Elfanbaum Copyeditor: Sonia Saruba Proofreader: Piper Editorial Consulting, LLC Indexer: nSight, Inc. Interior Designer: David Futato Cover Designer: Jose Marzan Jr. Illustrator: Kate Dullea December 2021: First Edition Revision History for the First Edition 2021-12-02: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781492091509 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Practical Python Data Wrangling and Data Quality, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the author, and do not represent the publisher’s views. While the publisher and the author have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the author disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Page 5
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix 1. Introduction to Data Wrangling and Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 What Is “Data Wrangling”? 2 What Is “Data Quality”? 3 Data Integrity 4 Data “Fit” 5 Why Python? 6 Versatility 6 Accessibility 7 Readability 7 Community 7 Python Alternatives 8 Writing and “Running” Python 8 Working with Python on Your Own Device 11 Getting Started with the Command Line 11 Installing Python, Jupyter Notebook, and a Code Editor 14 Working with Python Online 19 Hello World! 20 Using Atom to Create a Standalone Python File 20 Using Jupyter to Create a New Python Notebook 21 Using Google Colab to Create a New Python Notebook 22 Adding the Code 23 In a Standalone File 23 In a Notebook 23 Running the Code 23 In a Standalone File 23 In a Notebook 24 iii
Page 6
Documenting, Saving, and Versioning Your Work 24 Documenting 24 Saving 25 Versioning 26 Conclusion 35 2. Introduction to Python. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 The Programming “Parts of Speech” 38 Nouns ≈ Variables 39 Verbs ≈ Functions 42 Cooking with Custom Functions 46 Libraries: Borrowing Custom Functions from Other Coders 47 Taking Control: Loops and Conditionals 47 In the Loop 48 One Condition… 51 Understanding Errors 55 Syntax Snafus 56 Runtime Runaround 58 Logic Loss 60 Hitting the Road with Citi Bike Data 62 Starting with Pseudocode 63 Seeking Scale 68 Conclusion 70 3. Understanding Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 Assessing Data Fit 73 Validity 74 Reliability 76 Representativeness 77 Assessing Data Integrity 79 Necessary, but Not Sufficient 81 Important 82 Achievable 85 Improving Data Quality 88 Data Cleaning 88 Data Augmentation 89 Conclusion 90 4. Working with File-Based and Feed-Based Data in Python. . . . . . . . . . . . . . . . . . . . . . . . 91 Structured Versus Unstructured Data 93 Working with Structured Data 97 File-Based, Table-Type Data—Take It to Delimit 97 iv | Table of Contents
Page 7
Wrangling Table-Type Data with Python 99 Real-World Data Wrangling: Understanding Unemployment 105 XLSX, ODS, and All the Rest 107 Finally, Fixed-Width 114 Feed-Based Data—Web-Driven Live Updates 118 Wrangling Feed-Type Data with Python 120 Working with Unstructured Data 133 Image-Based Text: Accessing Data in PDFs 134 Wrangling PDFs with Python 135 Accessing PDF Tables with Tabula 139 Conclusion 140 5. Accessing Web-Based Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 Accessing Online XML and JSON 143 Introducing APIs 145 Basic APIs: A Search Engine Example 146 Specialized APIs: Adding Basic Authentication 148 Getting a FRED API Key 149 Using Your API key to Request Data 150 Reading API Documentation 151 Protecting Your API Key When Using Python 153 Creating Your “Credentials” File 155 Using Your Credentials in a Separate Script 155 Getting Started with .gitignore 157 Specialized APIs: Working With OAuth 159 Applying for a Twitter Developer Account 160 Creating Your Twitter “App” and Credentials 162 Encoding Your API Key and Secret 167 Requesting an Access Token and Data from the Twitter API 168 API Ethics 172 Web Scraping: The Data Source of Last Resort 173 Carefully Scraping the MTA 176 Using Browser Inspection Tools 178 The Python Web Scraping Solution: Beautiful Soup 180 Conclusion 184 6. Assessing Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 The Pandemic and the PPP 187 Assessing Data Integrity 187 Is It of Known Pedigree? 188 Is It Timely? 189 Is It Complete? 189 Table of Contents | v
Page 8
Is It Well-Annotated? 201 Is It High Volume? 206 Is It Consistent? 208 Is It Multivariate? 211 Is It Atomic? 213 Is It Clear? 213 Is It Dimensionally Structured? 215 Assessing Data Fit 215 Validity 216 Reliability 219 Representativeness 220 Conclusion 222 7. Cleaning, Transforming, and Augmenting Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 225 Selecting a Subset of Citi Bike Data 226 A Simple Split 227 Regular Expressions: Supercharged String Matching 229 Making a Date 233 De-crufting Data Files 235 Decrypting Excel Dates 239 Generating True CSVs from Fixed-Width Data 241 Correcting for Spelling Inconsistencies 244 The Circuitous Path to “Simple” Solutions 250 Gotchas That Will Get Ya! 252 Augmenting Your Data 253 Conclusion 256 8. Structuring and Refactoring Your Code. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 257 Revisiting Custom Functions 258 Will You Use It More Than Once? 258 Is It Ugly and Confusing? 258 Do You Just Really Hate the Default Functionality? 259 Understanding Scope 259 Defining the Parameters for Function “Ingredients” 262 What Are Your Options? 263 Getting Into Arguments? 263 Return Values 264 Climbing the “Stack” 265 Refactoring for Fun and Profit 267 A Function for Identifying Weekdays 267 Metadata Without the Mess 270 Documenting Your Custom Scripts and Functions with pydoc 277 vi | Table of Contents
Page 9
The Case for Command-Line Arguments 281 Where Scripts and Notebooks Diverge 284 Conclusion 285 9. Introduction to Data Analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 287 Context Is Everything 288 Same but Different 289 What’s Typical? Evaluating Central Tendency 290 What’s That Mean? 290 Embrace the Median 291 Think Different: Identifying Outliers 292 Visualization for Data Analysis 292 What’s Our Data’s Shape? Understanding Histograms 296 The Significance of Symmetry 297 Counting “Clusters” 305 The $2 Million Question 306 Proportional Response 317 Conclusion 321 10. Presenting Your Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323 Foundations for Visual Eloquence 324 Making Your Data Statement 326 Charts, Graphs, and Maps: Oh My! 327 Pie Charts 327 Bar and Column Charts 330 Line Charts 335 Scatter Charts 339 Maps 342 Elements of Eloquent Visuals 345 The “Finicky” Details Really Do Make a Difference 345 Trust Your Eyes (and the Experts) 345 Selecting Scales 347 Choosing Colors 347 Above All, Annotate! 348 From Basic to Beautiful: Customizing a Visualization with seaborn and matplotlib 349 Beyond the Basics 354 Conclusion 355 11. Beyond Python. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357 Additional Tools for Data Review 358 Spreadsheet Programs 358 Table of Contents | vii
Page 10
OpenRefine 359 Additional Tools for Sharing and Presenting Data 361 Image Editing for JPGs, PNGs, and GIFs 361 Software for Editing SVGs and Other Vector Formats 362 Reflecting on Ethics 363 Conclusion 364 A. More Python Programming Resources. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365 B. A Bit More About Git. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369 C. Finding Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375 D. Resources for Visualization and Information Design. . . . . . . . . . . . . . . . . . . . . . . . . . . . 381 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383 viii | Table of Contents
Page 11
Preface Welcome! If you’ve picked up this book, you’re likely one of the many millions of people intrigued by the processes and possibilities surrounding “data”—that incred‐ ible, elusive new “currency” that’s transforming the way we live, work, and even connect with one another. Most of us, for example, are vaguely aware of the fact that data—collected by our electronic devices and other activities—is being used to shape what advertisements we see, what media is recommended to us, and which search results populate first when we look for something online. What many people may not appreciate is that the tools and skills for accessing, transforming, and generating insight from data are readily available to them. This book aims to help those people— you, if you like—do just that. Data is not something that is only available or useful to big companies or governmen‐ tal number crunchers. Being able to access, understand, and gather insight from data is a valuable skill whether you’re a data scientist or a day care worker. The tools needed to use data effectively are more accessible than ever before. Not only can you do significant data work using only free software and programming languages, you don’t even need an expensive computer. All of the exercises in this book, for example, were designed and run on a Chromebook that cost less than $500. You can even just use free online platforms through the internet connection at your local library. The goal of this book is to provide the guidance and confidence that data novices need to begin exploring the world of data—first by accessing it, then evaluating its quality. With those foundations in place, we’ll move on to some of the basic methods of analyzing and presenting data to generate meaningful insight. While these latter sections will be far from comprehensive (both data analysis and visualization are robust fields unto themselves), they will give you the core skills needed to gener‐ ate accurate, informative analyses and visualizations using your newly cleaned and acquired data. ix
Page 12
1 For a long time, installing the tools was also a huge obstacle. Now all you need is an internet connection! Who Should Read This Book? This book is intended for true beginners; all you need are a basic understanding of how to use computers (e.g., how to download a file, open a program, copy and paste, etc.), an open mind, and a willingness to experiment. I especially encourage you to take a chance on this book if you are someone who feels intimidated by data or programming, if you’re “bad at math,” or imagine that working with data or learning to program is too hard for you. I have spent nearly a decade teaching hundreds of people who didn’t think of themselves as “technical” the exact skills contained in this book, and I have never once had a student who was genuinely unable to get through this material. In my experience, the most challenging part of programming and working with data is not the difficulty of the material but the quality of the instruction.1 I am grateful both to the many students over the years whose questions have helped me immeasurably in finding ways to convey this material better, and for the opportunity to share what I learned from them with so many others through this book. While a book cannot truly replace the kind of support provided by a human teacher, I hope it will at least give the tools you need to master the basics—and perhaps the inspiration to take those skills to the next level. Folks who have some experience with data wrangling but have reached the limits of spreadsheet tools or want to expand the range of data formats they can easily access and manipulate will also find this book useful, as will those with frontend programming skills (in JavaScript or PHP, for example) who are looking for a way to get started with Python. Where Would You like to Go? In the preface to media theorist Douglas Rushkoff ’s book Program or Be Programmed (OR Books), he compares the act of programming to that of driving a car. Unless you learn to program, Rushkoff writes, you are a perpetual passenger in the digital world, one who “is getting driven from place to place. Only the car has no windows and if the driver tells you there is only one supermarket in the county, you have to believe him.” “You can relegate your programming to others,” Rushkoff continues, “but then you have to trust them that their programs are really doing what you’re asking, and in a way that is in your best interests.” More and more these days, the latter assertion is being thrown into question. Over the years, I’ve asked several hundred students if they believe anyone can learn to drive, and the answer has always been yes. At the same time, I have met few people, apart from myself, who truly believe that anyone can program. Yet driving a motor x | Preface
Page 13
vehicle is, in reality, vastly more complex than programming a computer. Why, then, do so many of us imagine that programming will be “too hard” for us? For me, this is where the real strength of Rushkoff ’s analogy shows, because his “windowless car” doesn’t just hide the outside world from the passenger—it also hides the “driver” from passersby. It’s easy to believe that anyone can drive a car because we actually see all kinds of people driving cars, every day. When it comes to programming, though, we rarely get to see who is “behind the wheel,” which means our ideas about who can and should program are largely defined by media that portray programmers as typically white and overwhelmingly male. As a result, those characteristics have come to dominate who does program—but there’s no reason why it should. Because if you can drive a car—or even punctuate a sentence—I promise you can program a computer, too. Who Shouldn’t Read This Book? As noted previously, this book is intended for beginners. So while you may find some sections useful if you are new to data analysis or visualization, this volume is not designed to serve those with prior experience in Python or another data-focused programming language (like R). Fortunately, O’Reilly has many specialized volumes that deal with advanced Python topics and libraries, such as Wes McKinney’s Python for Data Analysis (O’Reilly) or the Python Data Science Handbook by Jake VanderPlas (O’Reilly). What to Expect from This Volume The content of this book is designed to be followed in the order presented, as the con‐ cepts and exercises in each chapter build on those explored previously. Throughout, however, you will find that exercises are presented in two ways: as code “notebooks” and as “standalone” programming files. The purpose of this is twofold. First, it allows you, the reader, to use whichever approach you prefer or find more accessible; sec‐ ond, it provides a way to compare these two methods of interacting with data-driven Python code. In my experience, Python “notebooks” are extremely useful for getting up and running quickly but can become tedious if you develop a reliable piece of code that you wish to run repeatedly. Since the code from one format often cannot simply be copied and pasted to the other, both are provided in the accompanying GitHub repo. Data files, too, are available via Google Drive. As you follow along with the exercises, you will be able to use the format you prefer and will also have the option of seeing the differences in the code for each format firsthand. Although Python is the primary tool used in this book, effective data wrangling and analysis are made easier through the smart use of a range of tools, from text editors (the programs in which you will actually write your code) to spreadsheet programs. Preface | xi
Page 14
Because of this, there are occasional exercises in this book that rely on other free and/or open source tools besides Python. Wherever these are introduced, I will offer some context as to why that tool has been chosen, along with sufficient instructions to complete the example task. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Monospaced Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords. Monospaced bold Shows commands or other text that should be typed literally by the user. Monospaced italic Shows text that should be replaced with user-supplied values or by values deter‐ mined by context. This element signifies a tip or suggestion. This element signifies a general note. This element indicates a warning or caution. xii | Preface
Page 15
Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download at https://github.com/PracticalPythonDataWranglingAndQuality. If you have a technical question or a problem using the code examples, please send email to bookquestions@oreilly.com. The code in this book is here to help you develop your skills. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Practical Python Data Wrangling and Data Quality by Susan E. McGregor (O’Reilly). Copyright 2022 Susan E. McGregor, 978-1-492-09150-9.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit http://oreilly.com. Preface | xiii
Page 16
How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-998-9938 (in the United States or Canada) 707-829-0515 (international or local) 707-829-0104 (fax) We have a web page for this book, where we list errata, examples, and any addi‐ tional information. You can access this page at https://www.oreilly.com/library/view/ practical-python-data/9781492091493. Email bookquestions@oreilly.com to comment or ask technical questions about this book. For news and information about our books and courses, visit http://oreilly.com. Find us on Facebook: http://facebook.com/oreilly Follow us on Twitter: http://twitter.com/oreillymedia Watch us on YouTube: http://www.youtube.com/oreillymedia Acknowledgments As I mentioned previously, this book owes much to my many students over the years who were brave enough to try something new and ask sincere questions along the way. The process of writing this book (to say nothing of the text itself) was made immeasurably better by my editor, Jeff Bleiel, whose pleasantness, flexibility, and light touch tempered my excesses while making space for my personal style. I am also grateful for the thoughtful and generous comments of my reviewers: Joanna S. Kao, Anne Bonner, and Randy Au. I would also like to thank Jess Haberman, who offered me the chance to make this material my own, as well as Jacqueline Kazil and Katharine Jarmul, who helped put me in her way. I’d also like to thank Jeannette Wing and Cliff Stein and the staff at Columbia University’s Data Science Institute, whose interest in this work has already helped it generate exciting new opportunities. And of course, I want to thank my friends and relations for their interest and support, even—and especially—when they had no idea what I was talking about. xiv | Preface
Page 17
Finally, I’d want to thank my family (including the children too young to read this) for staying supportive even when the Sad SpongeBob days set in. You make the work worth doing. Preface | xv
Page 18
(This page has no text content)
Page 19
CHAPTER 1 Introduction to Data Wrangling and Data Quality These days it seems like data is the answer to everything: we use the data in product and restaurant reviews to decide what to buy and where to eat; companies use the data about what we read, click, and watch to decide what content to produce and which advertisements to show; recruiters use data to decide which applicants get job interviews; the government uses data to decide everything from how to allocate highway funding to where your child goes to school. Data—whether it’s a basic table of numbers or the foundation of an “artificial intelligence” system—permeates our lives. The pervasive impact that data has on our experiences and opportunities every day is precisely why data wrangling is—and will continue to be—an essential skill for anyone interested in understanding and influencing how data-driven systems operate. Likewise, the ability to assess—and even improve—data quality is indispen‐ sable for anyone interested in making these sometimes (deeply) flawed systems work better. Yet because both the terms data wrangling and data quality will mean different things to different people, we’ll begin this chapter with a brief overview of the three main topics addressed in this book: data wrangling, data quality, and the Python program‐ ming language. The goal of this overview is to give you a sense of my approach to these topics, partly so you can determine if this book is right for you. After that, we’ll spend some time on the necessary logistics of how to access and configure the software tools and other resources you’ll need to follow along with and complete the exercises in this book. Though all of the resources that this book will reference are free to use, many programming books and tutorials take for granted that readers will be coding on (often quite expensive) computers that they own. Since I really believe that anyone who wants to can learn to wrangle data with Python, however, I wanted to make sure that the material in this book can work for you even if you don’t 1
Page 20
have access to a full-featured computer of your own. To help ensure this, all of the solutions you’ll find here and in the following chapters were written and tested on a Chromebook; they can also be run using free, online-only tools using either your own device or a shared computer, for example, at school or a public library. I hope that by illustrating how accessible not just the knowledge but also the tools of data wrangling can be will encourage you to explore this exciting and empowering practice. What Is “Data Wrangling”? Data wrangling is the process of taking “raw” or “found” data, and transforming it into something that can be used to generate insight and meaning. Driving every substantive data wrangling effort is a question: something about the world you want to investigate or learn more about. Of course, if you came to this book because you’re really excited about learning to program, then data wrangling can be a great way to get started, but let me urge you now not to try to skip straight to the programming without engaging the data quality processes in the chapters ahead. Because as much as data wrangling may benefit from programming skills, it is about much more than simply learning how to access and manipulate data; it’s about making judgments, inferences, and selections. As this book will illustrate, most data that is readily available is not especially good quality, so there’s no way to do data wrangling without making choices that will influence the substance of the resulting data. To attempt data wrangling without considering data quality is like trying drive a car without steering: you may get somewhere—and fast!—but it’s probably nowhere you want to be. If you’re going to spend time wrangling and analyzing data, you want to try to make sure it’s at least likely to be worth the effort. Just as importantly, though, there’s no better way to learn a new skill than to connect it to something you genuinely want to get “right,” because that personal interest is what will carry you through the inevitable moments of frustration. This doesn’t mean that question you choose has to be something of global importance. It can be a question about your favorite video games, bands, or types of tea. It can be a question about your school, your neighborhood, or your social media life. It can be a question about economics, politics, faith, or money. It just has to be something that you genuinely care about. Once you have your question in hand, you’re ready to begin the data wrangling pro‐ cess. While the specific steps may need adjusting (or repeating) depending on your particular project, in principle data wrangling involves some or all of the following steps: 1. Locating or collecting data 2. Reviewing the data 3. “Cleaning,” standardizing, transforming, and/or augmenting the data 2 | Chapter 1: Introduction to Data Wrangling and Data Quality
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
Back to List