Practical Python Data Wrangling and Data Quality Getting Started with Reading, Cleaning, and Analyzing Data (Susan E. McGregor)(Z-Library)
Python
No Description
89
Views
0
Downloads
0.00
Total Donations
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Page
1
(This page has no text content)
Page
2
(This page has no text content)
Page
3
Susan E. McGregor Practical Python Data Wrangling and Data Quality Boston Farnham Sebastopol TokyoBeijing
Page
4
978-1-492-09150-9 [LSI] Practical Python Data Wrangling and Data Quality by Susan E. McGregor Copyright © 2022 Susan E. McGregor. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Jessica Haberman Development Editor: Jeff Bleiel Production Editor: Daniel Elfanbaum Copyeditor: Sonia Saruba Proofreader: Piper Editorial Consulting, LLC Indexer: nSight, Inc. Interior Designer: David Futato Cover Designer: Jose Marzan Jr. Illustrator: Kate Dullea December 2021: First Edition Revision History for the First Edition 2021-12-02: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781492091509 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Practical Python Data Wrangling and Data Quality, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the author, and do not represent the publisher’s views. While the publisher and the author have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the author disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Page
5
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix 1. Introduction to Data Wrangling and Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 What Is “Data Wrangling”? 2 What Is “Data Quality”? 3 Data Integrity 4 Data “Fit” 5 Why Python? 6 Versatility 6 Accessibility 7 Readability 7 Community 7 Python Alternatives 8 Writing and “Running” Python 8 Working with Python on Your Own Device 11 Getting Started with the Command Line 11 Installing Python, Jupyter Notebook, and a Code Editor 14 Working with Python Online 19 Hello World! 20 Using Atom to Create a Standalone Python File 20 Using Jupyter to Create a New Python Notebook 21 Using Google Colab to Create a New Python Notebook 22 Adding the Code 23 In a Standalone File 23 In a Notebook 23 Running the Code 23 In a Standalone File 23 In a Notebook 24 iii
Page
6
Documenting, Saving, and Versioning Your Work 24 Documenting 24 Saving 25 Versioning 26 Conclusion 35 2. Introduction to Python. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 The Programming “Parts of Speech” 38 Nouns ≈ Variables 39 Verbs ≈ Functions 42 Cooking with Custom Functions 46 Libraries: Borrowing Custom Functions from Other Coders 47 Taking Control: Loops and Conditionals 47 In the Loop 48 One Condition… 51 Understanding Errors 55 Syntax Snafus 56 Runtime Runaround 58 Logic Loss 60 Hitting the Road with Citi Bike Data 62 Starting with Pseudocode 63 Seeking Scale 68 Conclusion 70 3. Understanding Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 Assessing Data Fit 73 Validity 74 Reliability 76 Representativeness 77 Assessing Data Integrity 79 Necessary, but Not Sufficient 81 Important 82 Achievable 85 Improving Data Quality 88 Data Cleaning 88 Data Augmentation 89 Conclusion 90 4. Working with File-Based and Feed-Based Data in Python. . . . . . . . . . . . . . . . . . . . . . . . 91 Structured Versus Unstructured Data 93 Working with Structured Data 97 File-Based, Table-Type Data—Take It to Delimit 97 iv | Table of Contents
Page
7
Wrangling Table-Type Data with Python 99 Real-World Data Wrangling: Understanding Unemployment 105 XLSX, ODS, and All the Rest 107 Finally, Fixed-Width 114 Feed-Based Data—Web-Driven Live Updates 118 Wrangling Feed-Type Data with Python 120 Working with Unstructured Data 133 Image-Based Text: Accessing Data in PDFs 134 Wrangling PDFs with Python 135 Accessing PDF Tables with Tabula 139 Conclusion 140 5. Accessing Web-Based Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 Accessing Online XML and JSON 143 Introducing APIs 145 Basic APIs: A Search Engine Example 146 Specialized APIs: Adding Basic Authentication 148 Getting a FRED API Key 149 Using Your API key to Request Data 150 Reading API Documentation 151 Protecting Your API Key When Using Python 153 Creating Your “Credentials” File 155 Using Your Credentials in a Separate Script 155 Getting Started with .gitignore 157 Specialized APIs: Working With OAuth 159 Applying for a Twitter Developer Account 160 Creating Your Twitter “App” and Credentials 162 Encoding Your API Key and Secret 167 Requesting an Access Token and Data from the Twitter API 168 API Ethics 172 Web Scraping: The Data Source of Last Resort 173 Carefully Scraping the MTA 176 Using Browser Inspection Tools 178 The Python Web Scraping Solution: Beautiful Soup 180 Conclusion 184 6. Assessing Data Quality. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 The Pandemic and the PPP 187 Assessing Data Integrity 187 Is It of Known Pedigree? 188 Is It Timely? 189 Is It Complete? 189 Table of Contents | v
Page
8
Is It Well-Annotated? 201 Is It High Volume? 206 Is It Consistent? 208 Is It Multivariate? 211 Is It Atomic? 213 Is It Clear? 213 Is It Dimensionally Structured? 215 Assessing Data Fit 215 Validity 216 Reliability 219 Representativeness 220 Conclusion 222 7. Cleaning, Transforming, and Augmenting Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 225 Selecting a Subset of Citi Bike Data 226 A Simple Split 227 Regular Expressions: Supercharged String Matching 229 Making a Date 233 De-crufting Data Files 235 Decrypting Excel Dates 239 Generating True CSVs from Fixed-Width Data 241 Correcting for Spelling Inconsistencies 244 The Circuitous Path to “Simple” Solutions 250 Gotchas That Will Get Ya! 252 Augmenting Your Data 253 Conclusion 256 8. Structuring and Refactoring Your Code. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 257 Revisiting Custom Functions 258 Will You Use It More Than Once? 258 Is It Ugly and Confusing? 258 Do You Just Really Hate the Default Functionality? 259 Understanding Scope 259 Defining the Parameters for Function “Ingredients” 262 What Are Your Options? 263 Getting Into Arguments? 263 Return Values 264 Climbing the “Stack” 265 Refactoring for Fun and Profit 267 A Function for Identifying Weekdays 267 Metadata Without the Mess 270 Documenting Your Custom Scripts and Functions with pydoc 277 vi | Table of Contents
Page
9
The Case for Command-Line Arguments 281 Where Scripts and Notebooks Diverge 284 Conclusion 285 9. Introduction to Data Analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 287 Context Is Everything 288 Same but Different 289 What’s Typical? Evaluating Central Tendency 290 What’s That Mean? 290 Embrace the Median 291 Think Different: Identifying Outliers 292 Visualization for Data Analysis 292 What’s Our Data’s Shape? Understanding Histograms 296 The Significance of Symmetry 297 Counting “Clusters” 305 The $2 Million Question 306 Proportional Response 317 Conclusion 321 10. Presenting Your Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323 Foundations for Visual Eloquence 324 Making Your Data Statement 326 Charts, Graphs, and Maps: Oh My! 327 Pie Charts 327 Bar and Column Charts 330 Line Charts 335 Scatter Charts 339 Maps 342 Elements of Eloquent Visuals 345 The “Finicky” Details Really Do Make a Difference 345 Trust Your Eyes (and the Experts) 345 Selecting Scales 347 Choosing Colors 347 Above All, Annotate! 348 From Basic to Beautiful: Customizing a Visualization with seaborn and matplotlib 349 Beyond the Basics 354 Conclusion 355 11. Beyond Python. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357 Additional Tools for Data Review 358 Spreadsheet Programs 358 Table of Contents | vii
Page
10
OpenRefine 359 Additional Tools for Sharing and Presenting Data 361 Image Editing for JPGs, PNGs, and GIFs 361 Software for Editing SVGs and Other Vector Formats 362 Reflecting on Ethics 363 Conclusion 364 A. More Python Programming Resources. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365 B. A Bit More About Git. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369 C. Finding Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375 D. Resources for Visualization and Information Design. . . . . . . . . . . . . . . . . . . . . . . . . . . . 381 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383 viii | Table of Contents
Page
11
Preface Welcome! If you’ve picked up this book, you’re likely one of the many millions of people intrigued by the processes and possibilities surrounding “data”—that incred‐ ible, elusive new “currency” that’s transforming the way we live, work, and even connect with one another. Most of us, for example, are vaguely aware of the fact that data—collected by our electronic devices and other activities—is being used to shape what advertisements we see, what media is recommended to us, and which search results populate first when we look for something online. What many people may not appreciate is that the tools and skills for accessing, transforming, and generating insight from data are readily available to them. This book aims to help those people— you, if you like—do just that. Data is not something that is only available or useful to big companies or governmen‐ tal number crunchers. Being able to access, understand, and gather insight from data is a valuable skill whether you’re a data scientist or a day care worker. The tools needed to use data effectively are more accessible than ever before. Not only can you do significant data work using only free software and programming languages, you don’t even need an expensive computer. All of the exercises in this book, for example, were designed and run on a Chromebook that cost less than $500. You can even just use free online platforms through the internet connection at your local library. The goal of this book is to provide the guidance and confidence that data novices need to begin exploring the world of data—first by accessing it, then evaluating its quality. With those foundations in place, we’ll move on to some of the basic methods of analyzing and presenting data to generate meaningful insight. While these latter sections will be far from comprehensive (both data analysis and visualization are robust fields unto themselves), they will give you the core skills needed to gener‐ ate accurate, informative analyses and visualizations using your newly cleaned and acquired data. ix
Page
12
1 For a long time, installing the tools was also a huge obstacle. Now all you need is an internet connection! Who Should Read This Book? This book is intended for true beginners; all you need are a basic understanding of how to use computers (e.g., how to download a file, open a program, copy and paste, etc.), an open mind, and a willingness to experiment. I especially encourage you to take a chance on this book if you are someone who feels intimidated by data or programming, if you’re “bad at math,” or imagine that working with data or learning to program is too hard for you. I have spent nearly a decade teaching hundreds of people who didn’t think of themselves as “technical” the exact skills contained in this book, and I have never once had a student who was genuinely unable to get through this material. In my experience, the most challenging part of programming and working with data is not the difficulty of the material but the quality of the instruction.1 I am grateful both to the many students over the years whose questions have helped me immeasurably in finding ways to convey this material better, and for the opportunity to share what I learned from them with so many others through this book. While a book cannot truly replace the kind of support provided by a human teacher, I hope it will at least give the tools you need to master the basics—and perhaps the inspiration to take those skills to the next level. Folks who have some experience with data wrangling but have reached the limits of spreadsheet tools or want to expand the range of data formats they can easily access and manipulate will also find this book useful, as will those with frontend programming skills (in JavaScript or PHP, for example) who are looking for a way to get started with Python. Where Would You like to Go? In the preface to media theorist Douglas Rushkoff ’s book Program or Be Programmed (OR Books), he compares the act of programming to that of driving a car. Unless you learn to program, Rushkoff writes, you are a perpetual passenger in the digital world, one who “is getting driven from place to place. Only the car has no windows and if the driver tells you there is only one supermarket in the county, you have to believe him.” “You can relegate your programming to others,” Rushkoff continues, “but then you have to trust them that their programs are really doing what you’re asking, and in a way that is in your best interests.” More and more these days, the latter assertion is being thrown into question. Over the years, I’ve asked several hundred students if they believe anyone can learn to drive, and the answer has always been yes. At the same time, I have met few people, apart from myself, who truly believe that anyone can program. Yet driving a motor x | Preface
Page
13
vehicle is, in reality, vastly more complex than programming a computer. Why, then, do so many of us imagine that programming will be “too hard” for us? For me, this is where the real strength of Rushkoff ’s analogy shows, because his “windowless car” doesn’t just hide the outside world from the passenger—it also hides the “driver” from passersby. It’s easy to believe that anyone can drive a car because we actually see all kinds of people driving cars, every day. When it comes to programming, though, we rarely get to see who is “behind the wheel,” which means our ideas about who can and should program are largely defined by media that portray programmers as typically white and overwhelmingly male. As a result, those characteristics have come to dominate who does program—but there’s no reason why it should. Because if you can drive a car—or even punctuate a sentence—I promise you can program a computer, too. Who Shouldn’t Read This Book? As noted previously, this book is intended for beginners. So while you may find some sections useful if you are new to data analysis or visualization, this volume is not designed to serve those with prior experience in Python or another data-focused programming language (like R). Fortunately, O’Reilly has many specialized volumes that deal with advanced Python topics and libraries, such as Wes McKinney’s Python for Data Analysis (O’Reilly) or the Python Data Science Handbook by Jake VanderPlas (O’Reilly). What to Expect from This Volume The content of this book is designed to be followed in the order presented, as the con‐ cepts and exercises in each chapter build on those explored previously. Throughout, however, you will find that exercises are presented in two ways: as code “notebooks” and as “standalone” programming files. The purpose of this is twofold. First, it allows you, the reader, to use whichever approach you prefer or find more accessible; sec‐ ond, it provides a way to compare these two methods of interacting with data-driven Python code. In my experience, Python “notebooks” are extremely useful for getting up and running quickly but can become tedious if you develop a reliable piece of code that you wish to run repeatedly. Since the code from one format often cannot simply be copied and pasted to the other, both are provided in the accompanying GitHub repo. Data files, too, are available via Google Drive. As you follow along with the exercises, you will be able to use the format you prefer and will also have the option of seeing the differences in the code for each format firsthand. Although Python is the primary tool used in this book, effective data wrangling and analysis are made easier through the smart use of a range of tools, from text editors (the programs in which you will actually write your code) to spreadsheet programs. Preface | xi
Page
14
Because of this, there are occasional exercises in this book that rely on other free and/or open source tools besides Python. Wherever these are introduced, I will offer some context as to why that tool has been chosen, along with sufficient instructions to complete the example task. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Monospaced Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords. Monospaced bold Shows commands or other text that should be typed literally by the user. Monospaced italic Shows text that should be replaced with user-supplied values or by values deter‐ mined by context. This element signifies a tip or suggestion. This element signifies a general note. This element indicates a warning or caution. xii | Preface
Page
15
Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download at https://github.com/PracticalPythonDataWranglingAndQuality. If you have a technical question or a problem using the code examples, please send email to bookquestions@oreilly.com. The code in this book is here to help you develop your skills. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Practical Python Data Wrangling and Data Quality by Susan E. McGregor (O’Reilly). Copyright 2022 Susan E. McGregor, 978-1-492-09150-9.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit http://oreilly.com. Preface | xiii
Page
16
How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-998-9938 (in the United States or Canada) 707-829-0515 (international or local) 707-829-0104 (fax) We have a web page for this book, where we list errata, examples, and any addi‐ tional information. You can access this page at https://www.oreilly.com/library/view/ practical-python-data/9781492091493. Email bookquestions@oreilly.com to comment or ask technical questions about this book. For news and information about our books and courses, visit http://oreilly.com. Find us on Facebook: http://facebook.com/oreilly Follow us on Twitter: http://twitter.com/oreillymedia Watch us on YouTube: http://www.youtube.com/oreillymedia Acknowledgments As I mentioned previously, this book owes much to my many students over the years who were brave enough to try something new and ask sincere questions along the way. The process of writing this book (to say nothing of the text itself) was made immeasurably better by my editor, Jeff Bleiel, whose pleasantness, flexibility, and light touch tempered my excesses while making space for my personal style. I am also grateful for the thoughtful and generous comments of my reviewers: Joanna S. Kao, Anne Bonner, and Randy Au. I would also like to thank Jess Haberman, who offered me the chance to make this material my own, as well as Jacqueline Kazil and Katharine Jarmul, who helped put me in her way. I’d also like to thank Jeannette Wing and Cliff Stein and the staff at Columbia University’s Data Science Institute, whose interest in this work has already helped it generate exciting new opportunities. And of course, I want to thank my friends and relations for their interest and support, even—and especially—when they had no idea what I was talking about. xiv | Preface
Page
17
Finally, I’d want to thank my family (including the children too young to read this) for staying supportive even when the Sad SpongeBob days set in. You make the work worth doing. Preface | xv
Page
18
(This page has no text content)
Page
19
CHAPTER 1 Introduction to Data Wrangling and Data Quality These days it seems like data is the answer to everything: we use the data in product and restaurant reviews to decide what to buy and where to eat; companies use the data about what we read, click, and watch to decide what content to produce and which advertisements to show; recruiters use data to decide which applicants get job interviews; the government uses data to decide everything from how to allocate highway funding to where your child goes to school. Data—whether it’s a basic table of numbers or the foundation of an “artificial intelligence” system—permeates our lives. The pervasive impact that data has on our experiences and opportunities every day is precisely why data wrangling is—and will continue to be—an essential skill for anyone interested in understanding and influencing how data-driven systems operate. Likewise, the ability to assess—and even improve—data quality is indispen‐ sable for anyone interested in making these sometimes (deeply) flawed systems work better. Yet because both the terms data wrangling and data quality will mean different things to different people, we’ll begin this chapter with a brief overview of the three main topics addressed in this book: data wrangling, data quality, and the Python program‐ ming language. The goal of this overview is to give you a sense of my approach to these topics, partly so you can determine if this book is right for you. After that, we’ll spend some time on the necessary logistics of how to access and configure the software tools and other resources you’ll need to follow along with and complete the exercises in this book. Though all of the resources that this book will reference are free to use, many programming books and tutorials take for granted that readers will be coding on (often quite expensive) computers that they own. Since I really believe that anyone who wants to can learn to wrangle data with Python, however, I wanted to make sure that the material in this book can work for you even if you don’t 1
Page
20
have access to a full-featured computer of your own. To help ensure this, all of the solutions you’ll find here and in the following chapters were written and tested on a Chromebook; they can also be run using free, online-only tools using either your own device or a shared computer, for example, at school or a public library. I hope that by illustrating how accessible not just the knowledge but also the tools of data wrangling can be will encourage you to explore this exciting and empowering practice. What Is “Data Wrangling”? Data wrangling is the process of taking “raw” or “found” data, and transforming it into something that can be used to generate insight and meaning. Driving every substantive data wrangling effort is a question: something about the world you want to investigate or learn more about. Of course, if you came to this book because you’re really excited about learning to program, then data wrangling can be a great way to get started, but let me urge you now not to try to skip straight to the programming without engaging the data quality processes in the chapters ahead. Because as much as data wrangling may benefit from programming skills, it is about much more than simply learning how to access and manipulate data; it’s about making judgments, inferences, and selections. As this book will illustrate, most data that is readily available is not especially good quality, so there’s no way to do data wrangling without making choices that will influence the substance of the resulting data. To attempt data wrangling without considering data quality is like trying drive a car without steering: you may get somewhere—and fast!—but it’s probably nowhere you want to be. If you’re going to spend time wrangling and analyzing data, you want to try to make sure it’s at least likely to be worth the effort. Just as importantly, though, there’s no better way to learn a new skill than to connect it to something you genuinely want to get “right,” because that personal interest is what will carry you through the inevitable moments of frustration. This doesn’t mean that question you choose has to be something of global importance. It can be a question about your favorite video games, bands, or types of tea. It can be a question about your school, your neighborhood, or your social media life. It can be a question about economics, politics, faith, or money. It just has to be something that you genuinely care about. Once you have your question in hand, you’re ready to begin the data wrangling pro‐ cess. While the specific steps may need adjusting (or repeating) depending on your particular project, in principle data wrangling involves some or all of the following steps: 1. Locating or collecting data 2. Reviewing the data 3. “Cleaning,” standardizing, transforming, and/or augmenting the data 2 | Chapter 1: Introduction to Data Wrangling and Data Quality
The above is a preview of the first 20 pages. Register to read the complete e-book.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
# Practical Python Data Wrangling and Data Quality
## 【One-Line Pitch】
A hands-on guide for journalists, analysts, and researchers who need to turn messy, real-world data into clean, trustworthy datasets using Python—covering everything from data quality assessment to code refactoring. If you've ever felt that your data is "dirty" and you're not sure where to start, this book is your practical roadmap.
## 【Book Arc】
- **Opening (~0%–10%)**: Introduces the core philosophy that data is inherently human-made and therefore flawed, establishing the two key axes of data quality: integrity (how well the data is structured and documented) and fit (how appropriate it is for your specific question). Sets up the book's practical, project-based approach.
- **Early (~10%–25%)**: Walks through setting up a complete Python environment (Miniconda, Jupyter Notebook, Atom editor) and establishing a "data wrangling diary" workflow—emphasizing the importance of writing your research question as a single sentence and using version control (Git/GitHub) to track all changes.
- **Early (~25%–35%)**: Covers Python fundamentals through the lens of data wrangling, introducing the five core data types (numbers, strings, lists, dictionaries, booleans) and the crucial concept that "nouns ≈ variables"—using syntax highlighting and punctuation to reliably identify data types.
- **Middle (~35%–50%)**: Dives into programming structures—custom functions (described as "recipes"), loops, conditionals, and nesting—with a strong emphasis on Python's whitespace-dependent indentation. Introduces exception handling but argues that for data quality work, you often *want* errors to surface rather than silently pass over bad data.
- **Middle (~50%–70%)**: Applies these skills to real data wrangling tasks, using the Citi Bike dataset as a running example. Covers selecting data subsets, regular expressions for string matching, handling dates, cleaning fixed-width files, correcting spelling inconsistencies, and augmenting data with additional sources.
- **Late (~70%–100%)**: Focuses on structuring and refactoring code for maintainability, revisiting custom functions with criteria like "will you use it more than once?" and "is it ugly and confusing?"—plus a comprehensive framework for assessing data quality across dimensions like annotation, volume, consistency, atomicity, clarity, and dimensionality.
## 【Key Takeaways】
- **Data quality has two distinct axes: integrity and fit** (Early): Integrity concerns the data's internal characteristics (completeness, atomicity, annotation), while fit asks whether the data actually answers your question. Both matter, and neither is guaranteed by sophisticated tools.
- **Computers amplify human judgment—they don't replace it** (Early): Even "intelligent" systems are just pattern-matching on human-selected data. The responsibility for data quality rests with the humans who collect, clean, and analyze it.
- **Write your research question as a single sentence** (Early): This simple discipline prevents you from losing track of your goal when data wrangling inevitably leads you down "rabbit holes." Your question becomes the anchor for all subsequent decisions.
- **Python's five data types are identifiable by punctuation alone** (Early): Numbers are just digits, strings have matching quotes, lists use square brackets, dictionaries use braces, and booleans are True/False. This makes type identification reliable and beginner-friendly.
- **Indentation is not optional in Python—it's structural** (Middle): Unlike curly-brace languages, Python uses whitespace to define code blocks. Understanding nesting (one tab per level) is essential for writing loops, conditionals, and functions that work correctly.
- **Exceptions are your friends in data quality work** (Middle): While you *can* write code to handle errors gracefully, the book argues you often *shouldn't*—because silent error handling can hide data quality problems you need to know about.
- **Comments are for your future self, not the computer** (Middle): The book's examples deliberately include more comment lines than code, because you *will* forget what your code does and why. Detailed comments are a sign of good practice, not verbosity.
- **Data cleaning is rarely a straight line** (Middle): The "circuitous path to simple solutions" is normal—real data has spelling inconsistencies, weird date formats, and fixed-width files that need conversion. Expect detours and build them into your workflow.
## 【Reading Tips】
- **Skim the setup chapters (10%–25%) if you already have Python installed**: The Miniconda/Jupyter/Atom installation walkthrough is thorough but platform-specific; focus instead on the "data diary" concept and Git workflow, which are transferable habits.
- **Deep-read the Citi Bike examples (Middle section)**: The running example of counting Subscriber vs. Customer user types is deceptively simple—it demonstrates the full loop of reading data, transforming it, and checking for unexpected values. This is where the book's philosophy becomes concrete.
- **Pay special attention to the "Gotchas" sections**: These are the hard-won lessons from real data wrangling—things like empty values not converting to zero, or Excel dates needing decryption. These will save you hours of debugging later.
- **Don't skip the code comments**: The book deliberately includes extensive inline comments in its examples. Reading them is like having the author explain her thinking process, which is more valuable than just copying the code.
- **Use the data quality framework as a checklist**: The late chapters on assessing data fit (validity, reliability, representativeness) and structural quality (atomicity, clarity, dimensionality) are worth returning to as a reference whenever you start a new dataset.
## 【Coverage Limits】
This guide covers the book's core philosophy, setup, Python fundamentals, and data cleaning approach through the middle sections. The excerpts do not cover the later chapters on advanced data augmentation techniques, code refactoring in depth, or the final data quality assessment frameworks in detail—those sections are summarized from the table of contents rather than excerpted content.
##
Passage locations
Excerpt 1
226 A Simple Split 227 Regular Expressions: Supercharged String Matching 229 Making a Date 233 De-crufting Data Files 235 Decrypting Excel Dates 239 Generati...
View in text
Excerpt 2
it possible to do a wider range of more conclusive analyses. In most cases, however, you’ll find that a given dataset is lacking on any number of data integr...
View in text
Excerpt 3
uction to Data Wrangling and Data Quality Nouns ≈ Variables In the English language, nouns are often described as any word that refers to a “person, place, o...
View in text
Excerpt 4
se for us to deal with runtime errors as they arise, rather than trying to plan for (and handle) all possible errors in advance.6 Of course, it is possible t...
View in text
Recommended for You
{{#thumbnailUrl}}
{{/thumbnailUrl}}
{{^thumbnailUrl}}
{{/thumbnailUrl}}
Loading recommended books...
Failed to load, please try again later
Tip the Site
Scan the WeChat Pay or Alipay code to tip. No login required.
WeChat Pay
Alipay