Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Catherine Nelson

Rating No ratings yet

Data science happens in code. The ability to write reproducible, robust, scaleable code is key to a data science project's success—and is absolutely essential for those working with production code. This practical book bridges the gap between data science and software engineering, and clearly explains how to apply the best practices from software engineering to data science. Examples are provided in Python, drawn from popular packages such as NumPy and pandas. If you want to write better data science code, this guide covers the essential topics that are often missing from introductory data science or coding classes, including how to: Understand data structures and object-oriented programming Clearly and skillfully document your code Package and share your code Integrate data science code with a larger code base Learn how to write APIs Create secure code Apply best practices to common tasks such as testing, error handling, and logging Work more effectively with software engineers Write more efficient, maintainable, and robust code in Python Put your data science projects into production And more

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical bridge from notebook-based data science to production-grade Python, showing data scientists how to write reproducible, tested, and scalable code. Best for analysts, ML engineers, and self-taught practitioners who can already use pandas and scikit-learn but have never been taught software engineering discipline. 【Book Arc】 - **Opening (~0%–10%)**: Frames the core problem—data science happens in code, yet introductory courses skip the engineering skills needed for production. Defines the audience and prerequisites (Python, NumPy, pandas, basic ML) and previews the full topic map from documentation to APIs. - **Early (~10%–32%)**: Establishes foundational principles—simplicity, readability, modularity, error handling, and testing—then moves into performance analysis: timing with `time`/`timeit`, profiling with `cProfile` and SnakeViz, and Big O notation as a hardware-independent way to reason about how code slows as data grows. - **Middle (~32%–48%)**: Deepens into data structures and their performance trade-offs: Python lists, tuples, dictionaries, and sets; NumPy arrays and vectorization; pandas DataFrames and Series. Explains hash tables, O(1) dictionary lookups, NumPy's fixed-size allocation costs, and sparse matrices for memory efficiency. - **Late (~48%–70%)**: Shifts from analysis to engineering practice—object-oriented and functional paradigms, interfaces and contracts, coupling, refactoring notebooks into scripts, and documentation (names, comments, docstrings, READMEs, and documenting ML experiments). - **Ending (~70%–100%)**: Covers sharing and shipping code: version control with Git, dependency management, packaging, writing APIs with FastAPI, consuming external APIs (e.g., SDG API), security, and integrating data science code into larger production codebases. 【Key Takeaways】 - **Complexity is the enemy of maintainability** (Early): Ousterhout's definition—anything in a system's structure that makes it hard to understand and modify—anchors the book's argument that simplicity is an engineering choice, not an aesthetic one. - **Break large projects into discrete, self-contained components** (Early): Sketching a system as a flowchart of functions (load, clean, plot) with explicit inputs creates a skeleton you can fill in, making each piece independently testable and replaceable. - **Measure before you optimize** (Early): Timing and profiling tools reveal actual bottlenecks; Big O notation predicts how runtime scales with data size independent of hardware, so you can choose better algorithms rather than micro-tweaking. - **Choose data structures for both performance and clarity** (Middle): Dictionaries offer O(1) lookup, insertion, and deletion via hash tables; NumPy vectorized operations can be ~100x faster than Python loops, but appending to arrays is O(n), so pre-allocate with `np.zeros`. - **pandas inherits NumPy's principles but adds its own** (Middle): DataFrames are built from Series with indexes; the 2.0 release added PyArrow backends, and sparse matrices help when data is mostly empty—performance work here is about memory as much as speed. - **Notebooks are for exploration; scripts are for production** (Late): Refactoring notebooks into scripts with clear interfaces and low coupling is a deliberate workflow, not a rewrite—strategies and an example workflow are provided. - **Documentation is part of the codebase** (Late): Names, comments, docstrings, READMEs, and tutorials serve different audiences; ML experiments need their own documentation discipline so results are reproducible. - **Sharing code requires version control, dependency management, and packaging** (Ending): Git, dependency pinning, and packaging turn personal code into something others can install, run, and trust—and APIs (FastAPI) extend that reach to other systems. 【Reading Tips】 - **Skim the ML-specific sections if your job doesn't involve ML**: The author explicitly marks these with "ML" in the section name and says you can skip them without losing continuity. - **Deep-read Chapters 2–3 on performance and data structures**: These are the most technically dense and the most immediately applicable to everyday data science code; the timing/profiling examples are worth running yourself. - **Treat the refactoring chapter as a workflow to practice, not just read**: The notebook-to-script transition is where most readers will feel the gap between knowing and doing—try it on one of your own notebooks. - **Use the documentation and packaging chapters as reference**: Return to them when you're actually preparing code to share or deploy, rather than reading passively. - **Don't skip the "who this book is for" preface**: It clarifies prerequisites and helps you decide which chapters to prioritize based on your background. 【Coverage Limits】 This guide is based on stratified excerpts covering the preface, table of contents, and portions of the early and middle chapters; the later chapters on Git, packaging, APIs, and security are represented mainly through the table of contents and brief mentions, so specific techniques and examples from those chapters are not detailed here.
Page 10
120 Creating Scripts from Notebooks 121 Refactoring 124 Strategies for Refactoring 124 An Example Refactoring Workflow 125 Key Takeaways 127 9. Documentation...
View in text
Excerpt 2
general note. This element indicates a warning or caution. xvi | Preface This adaptability becomes more important as your codebase grows. With a single small...
View in text
Excerpt 3
SnakeViz with the following command: $ pip install snakeviz Then, if you’re working in the Jupyter Notebook you can use the SnakeViz extension. You can load...
View in text
Excerpt 4
the zeros with the new elements instead of appending to the array. You also can save a lot of memory space with NumPy arrays by taking advantage of NumPy’s d...
View in text
Excerpt 5
ance of it; that just adds extra complexity you don’t need. If you find yourself wanting to do new things to some data that remains fixed, FP might be a good...
View in text
Excerpt 6
{running_total}") return (running_total/len(num_list)) 72 | Chapter 5: Errors, Logging, and Debugging Table 5-2 lists command shortcuts for pdb based on this...
View in text
Excerpt 7
run the tests, and confirm that your code also works there. Testing provides some assurance that your code does what you say it should do. This helps other p...
View in text
Excerpt 8
’t know exactly what model you’ll get from a given dataset, because most machine learning algorithms include randomization in some way. But this doesn’t mean...
View in text
Tags
AI categories
DataPythonSoftware
ISBN: 1098136209
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 258
File Format: PDF
File Size: 6.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…