Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Reuven M. Lerner

Rating No ratings yet

Practice makes perfect pandas! Work out your pandas skills against dozens of real-world challenges, each carefully designed to build an intuitive knowledge of essential pandas tasks. In Pandas Workout you’ll learn how to: • Clean your data for accurate analysis • Work with rows and columns for retrieving and assigning data • Handle indexes, including hierarchical indexes • Read and write data with a number of common formats, such as CSV and JSON • Process and manipulate textual data from within pandas • Work with dates and times in pandas • Perform aggregate calculations on selected subsets of data • Produce attractive and useful visualizations that make your data come alive

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Pandas Workout: 200 Exercises to Make You a Stronger Data Analyst ## 【One-Line Pitch】 A hands-on, exercise-driven guide that builds practical pandas fluency through 200 real-world challenges, perfect for data analysts who learn best by doing rather than reading. If you know Python basics but want to move from "can follow tutorials" to "confidently solve data problems," this book is your training ground. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces the book's structure—13 chapters, each with exercises that build on prior techniques—and starts with the fundamentals of Series: creating them, retrieving values with `.loc` and `.iloc`, and understanding descriptive statistics like mean, median, and standard deviation. - **Early (~10%–23%)**: Moves into DataFrames—constructing them from lists of lists or dicts, performing column arithmetic, and using `value_counts` for frequency analysis. Exercises use realistic data like taxi passenger counts and test scores. - **Early-to-Middle (~23%–39%)**: Covers importing and exporting data, with deep dives into CSV quirks, dtype inference problems, and memory-efficient loading. Includes practical exercises on NYC taxi data, detecting outliers via IQR, and handling missing values. - **Middle (~39%–48%)**: Addresses indexes—setting, resetting, and creating hierarchical multi-indexes—plus pivot tables. Exercises include parking tickets, SAT scores, and Olympic games data, showing how flexible indexing enables more powerful queries. - **Middle-to-Late (~48%–100%)**: Continues into advanced grouping, joining, and sorting (two full chapters), then moves through string manipulation, datetime handling, visualization with pandas and Seaborn, performance optimization, and two large capstone projects (Python developer survey and American colleges data). ## 【Key Takeaways】 - **Series are the atomic unit of pandas** (Early): Master creating Series, retrieving values via `.loc` (label-based) and `.iloc` (position-based), and applying vectorized operations like `s + 10`. Understanding Series deeply makes DataFrame work intuitive. - **Descriptive statistics need context** (Early): The mean alone can mislead—a single outlier (the "Bill Gates walks into a bar" problem) skews it. Use median, quantiles, and standard deviation together; `s.mean()` equals `s.sum() / s.count()`, which helps you understand what's actually being calculated. - **`value_counts` is your best friend for categorical data** (Early): It returns a sorted frequency Series, supports `normalize=True` for percentages, and can be indexed like any Series (e.g., `s.value_counts()[[1,6]]`). This one method handles most "how common is X" questions. - **Creating DataFrames from scratch has four patterns** (Early): Lists of lists (positional), lists of dicts (key-based), dicts of lists, and dicts of Series. Each suits different data assembly scenarios—knowing all four prevents awkward workarounds. - **CSV loading requires active dtype management** (Middle): Pandas guesses dtypes by sampling, which can cause mixed-type columns and `DtypeWarning`. Pass explicit `dtype` dictionaries (e.g., `{'passenger_count': np.int8}`) to control memory usage and avoid surprises; remember NaN values force float dtypes. - **Outlier detection is a judgment call** (Early): The IQR method (1.5× IQR beyond quartiles) is standard, but alternatives like z-scores (>3) or trimming top/bottom 10% give different results. The exercise shows how to combine multiple columns (trip distance, passenger count) to find meaningful outliers. - **Indexes are query power tools** (Middle): Setting meaningful indexes (dates, names) and building multi-indexes enables hierarchical slicing that positional indexing can't match. Pivot tables (`df.pivot` and `df.pivot_table`) transform raw data into aggregate summaries—essential for reporting. - **Performance matters at scale** (Late): The book dedicates a full chapter to speed and memory optimization, covering techniques like dtype selection, `low_memory=False` trade-offs, and efficient column selection. These skills separate analysts who can handle big data from those who crash. ## 【Reading Tips】 - **Do every exercise before reading the solution**—the "Working it out" sections explain the reasoning, but the learning comes from struggling first. The "Beyond the exercise" extensions are where deeper understanding crystallizes. - **Skim the reference tables** (e.g., Table 3.1, Table 4.1) on first pass; they're dense but serve as excellent quick-reference when you're stuck later. Mark them for return visits. - **Pay special attention to the dtype discussion in Chapter 3** (around Exercise 17–19)—this is subtle, practical knowledge that prevents real-world headaches with large CSVs. The `low_memory` parameter and explicit dtype dicts are worth memorizing. - **The two project chapters (8 and 13) are capstones**—if you're short on time, prioritize the exercises in Chapters 6–7 (grouping/joining) and 11 (visualization) before attempting them, as they integrate everything. - **Keep a Python REPL open while reading**—many exercises use the NYC taxi dataset or generated random data; experimenting interactively as you read each problem statement builds muscle memory faster than passive reading. ## 【Coverage Limits】 This guide synthesizes excerpts covering roughly the first half of the book (through Chapter 4 on indexes). Detailed content on advanced grouping, strings, dates, visualization, performance, and the final projects is not covered here—those sections are summarized only from the table of contents. ##
Page 16
2 xiv ABOUT THIS BOOKHow this book is organized: A road map This book has 13 chapters, each focusing on a different aspect of pandas. Exercises in each chapt...
View in text
Excerpt 2
Feb 78.0 Mar 80.0 Apr 90.0 May 89.0 Jun 75.0 dtype: float64 Notice how this solution moves back and forth between scalar values and series, which is common i...
View in text
Excerpt 3
lue of the z-score is greater than 3. NaN and missing data So far, we have seen that analyzing data with pandas isn’t overly difficult. We need to know what...
View in text
Excerpt 4
er type you want. EXERCISE 19 ■ Bitcoin values 93in the URL. For those that don’t, you can’t retrieve directly via read_csv. Rather, you need to retrieve the...
View in text
Excerpt 5
ears 1980–2016, all seasons, and all sports from "Swimming" from pandas import IndexSlice as idx to "Table tennis" df.loc[idx[1980:2016, :, 'Swimming':'Table...
View in text
Excerpt 6
index Reorder the rows of a df = df.sort_index() http://mng.bz/wvB7 data frame based on the values in its index, in ascending order df.sort_values Reorder th...
View in text
Excerpt 7
2019-02-18 12:00:00 556 1 -1 Boston MA 12:00:00 2019-01-09 237 16 10 Los CA 15:00:00 Angeles date_time max_temp min_temp city state 2019-01-14 278 12 10 Los...
View in text
Excerpt 8
mbers, each representing one measure- ment of precipitation. What function can we write that will return a new series with the same length and index as the o...
View in text
Tags
AI categories
ProgrammingDataPython
Publish Year: 2024
Language: English
Pages: 442
File Format: PDF
File Size: 4.2 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…