Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Joel Grus

Rating No ratings yet

Data science libraries, frameworks, modules, and toolkits are great for doing data science, but they're also a good way to dive into the discipline without actually understanding data science. With this updated second edition, you'll learn how many of the most fundamental data science tools and algorithms work by implementing them from scratch. If you have an aptitude for mathematics and some programming skills, author Joel Grus will help you get comfortable with the math and statistics at the core of data science, and with hacking skills you need to get started as a data scientist. Today's messy glut of data holds answers to questions no one's even thought to ask. This book provides you with the know-how to dig those answers out.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Science from Scratch — Reading Guide ## 【One-Line Pitch】 A hands-on, code-first tour through the fundamental math, statistics, and algorithms of data science, teaching you to build core tools from scratch rather than blindly relying on libraries. Ideal for programmers with basic Python skills who want to truly understand what's under the hood before using scikit-learn, pandas, or TensorFlow. ## 【Book Arc】 - **Opening (~0%–10%)**: Sets up the "data science Venn diagram" (hacking skills + math/statistics + substantive expertise) and explains why the book focuses on the first two. Introduces a running example—a fictional social network called DataSciencester—and starts manipulating friendship data with basic Python structures. - **Early (~10%–23%)**: A crash course in Python essentials (functions, dictionaries, defaultdict, classes, and "dunder" methods), followed by visualization with matplotlib and the foundational building blocks of linear algebra: vectors, matrices, and their operations. - **Early-to-Middle (~23%–39%)**: Covers descriptive statistics (mean, median, quantiles, variance, correlation), then probability theory (PDFs, CDFs, normal distribution), hypothesis testing, p-hacking, and gradient descent as the workhorse optimization technique for machine learning. - **Middle (~39%–48%)**: Moves into practical data wrangling—reading/writing CSV files, scraping HTML/XML with Beautiful Soup, working with APIs (GitHub example), and cleaning/transforming data. Introduces dataclasses and more advanced data structures. - **Late (~48%–end)**: Dives into machine learning proper: PCA for dimensionality reduction, the critical distinction between training/validation/test sets, and why "accuracy" is a misleading metric for classification (the "Luke leukemia test" example). Excerpts suggest the book continues into regression, classification, and beyond. ## 【Key Takeaways】 - **Data science = hacking skills + math/statistics** (Opening): The author deliberately skips "substantive expertise" to keep the book focused; you learn how to code and think statistically, not domain knowledge. (Early) - **Python fundamentals are the foundation** (Early): Master dictionaries, defaultdict, Counter, classes, and first-class functions before touching any data science library—these are the tools you'll use to build everything else. (Early) - **Vectors and matrices are the language of data** (Early): Understanding shape, row/column operations, and matrix construction via functions like `make_matrix` is prerequisite for every algorithm later in the book. (Early) - **Mean vs. median is a real trade-off** (Early): The mean is smooth and calculus-friendly but outlier-sensitive (the Michael Jordan salary example); the median is robust but requires sorting. Choose based on whether outliers are signal or noise. (Early) - **Probability distributions are building blocks** (Early): The uniform and normal distributions, their PDFs and CDFs, underpin everything from random sampling to hypothesis testing—and you'll implement them yourself. (Early) - **p-hacking is a real danger** (Early): If you test enough hypotheses, 5% will appear "significant" by chance alone; the book demonstrates this with a coin-flip experiment that rejects fairness 46 out of 1000 times. (Early) - **Gradient descent is the workhorse** (Early): Maximizing/minimizing functions by stepping in the direction of the gradient is the core of most ML algorithms; minibatch variants make it practical for large datasets. (Middle) - **Data wrangling is unglamorous but essential** (Middle): CSV handling, HTML/XML scraping, and API calls (with date parsing via `dateutil`) are the messy realities of real data science work. (Middle) - **Accuracy is a trap for classification** (Late): A test that predicts "leukemia if named Luke" is >98% accurate but useless—you need precision, recall, and proper train/validation/test splits. (Late) ## 【Reading Tips】 1. **Skim the Python crash course if you're experienced** (Early): Chapters 2–3 cover basics you may know; focus instead on the defaultdict/Counter patterns and the vector/matrix implementations, which appear throughout. 2. **Deep-read the statistics and probability chapters** (Early): These are the conceptual core—don't skip the discussions of mean vs. median, quantiles, and the normal distribution even if the code seems simple. 3. **Code along with the DataSciencester example** (Opening–Early): The running social-network dataset makes abstract concepts concrete; implement the friendship-graph manipulations yourself. 4. **Watch for the "gotcha" examples** (Early–Late): The p-hacking experiment and the "Luke test" are memorable warnings about statistical pitfalls—internalize these lessons even if you skim the surrounding code. 5. **Expect to install third-party libraries** (Middle): The book uses matplotlib, requests, Beautiful Soup, and python-dateutil; set up a clean Python environment before starting the data-wrangling chapters. ## 【Coverage Limits】 This guide is based on excerpts covering roughly the first half of the book (through PCA and model evaluation). The later chapters on regression, classification algorithms, neural networks, and natural language processing are not covered in the sampled material. ##
Page 10
matics and statistics that are at the core of data science. This is a somewhat heavy aspiration for a book. The best way to learn hacking skills is by hackin...
View in text
Excerpt 2
s: (1, 2) # keyword args: {'key': 'word', 'key2': 'word2'} That is, when we define a function like this, args is a tuple of its unnamed arguments and kwargs...
View in text
Excerpt 3
ool]) -> bool: """Using the 5% significance levels""" num_heads = len([flip for flip in experiment if flip]) return num_heads < 469 or num_heads > 531 random...
View in text
Excerpt 4
losing_price: float def is_high_tech(self) -> bool: """It's a class, so we can add methods too""" return self.symbol in ['MSFT', 'GOOG', 'FB', 'AMZN', 'AAPL'...
View in text
Excerpt 5
magine that we have a sample of data v1, ..., vn that comes model is about that coefficient. Unfortunately, we’re not set up to do that kind of linear algebr...
View in text
Excerpt 6
ance=variance) for _ in range(dims[0])] assert shape(random_uniform(2, 3, 4)) == [2, 3, 4] assert shape(random_normal(5, 6, mean=10)) == [5, 6] And then wrap...
View in text
Excerpt 7
'MongoDB', 'Cassandra', 'HBase', 'Postgres'] databases 5 ['Python', 'scikit-learn', 'scipy', 'numpy', 'statsmodels', 'pandas'] Python and statistics 5 machin...
View in text
Excerpt 8
look at the first one.) return pr If we compute page ranks: pr = page_rank(users, endorsements) # Thor (user_id 4) has higher page rank than anyone else asse...
View in text
Tags
AI categories
ProgrammingPythonData
Publisher: O'Reilly Media
Publish Year: 2019
Language: English
File Format: PDF
File Size: 10.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…