Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Catherine Nelson

Rating No ratings yet

Data science happens in code. The ability to write reproducible, robust, scaleable code is key to a data science project's success—and is absolutely essential for those working with production code. This practical book bridges the gap between data science and software engineering,and clearly explains how to apply the best practices from software engineering to data science. Examples are provided in Python, drawn from popular packages such as NumPy and pandas. If you want to write better data science code, this guide covers the essential topics that are often missing from introductory data science or coding classes, including how to: Understand data structures and object-oriented programming Clearly and skillfully document your code Package and share your code Integrate data science code with a larger code base Learn how to write APIs Create secure code Apply best practices to common tasks such as testing, error handling, and logging Work more effectively with software engineers Write more efficient, maintainable, and robust code in Python Put your data science projects into production And more

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Software Engineering for Data Scientists: From Notebooks to Scalable Systems ## 【One-Line Pitch】 A practical bridge between data science and software engineering, this book teaches Python-focused best practices—from writing clean, documented code to packaging, testing, and deploying ML systems—for data scientists who want their work to survive outside the notebook. ## 【Book Arc】 - **Opening (~0%–9%)**: Sets the stage by defining "good code" through the lens of complexity, readability, and maintainability. Introduces the core problem: data science code often works in isolation but fails when integrated into larger systems. Includes a chapter-by-chapter roadmap covering version control, APIs, CI/CD, security, and team practices. - **Early (~9%–25%)**: Dives into code quality fundamentals—naming conventions, cleaning up commented-out code, the "Broken Window Theory" of code quality, and refactoring. Transitions into performance analysis, emphasizing that code must work correctly before being optimized, then introduces profiling tools like `timeit`, `cProfile`, and `SnakeViz`. - **Early–Middle (~25%–38%)**: Explores algorithmic complexity (Big O notation) and data structure performance. Compares Python lists vs. NumPy arrays, explaining memory allocation, view vs. copy semantics, and why single-type arrays enable faster operations. Introduces Dask for distributed computing and touches on arrays in machine learning. - **Middle (~38%–47%)**: Covers object-oriented programming (OOP) in Python—classes, `__init__`, `self`, attributes, and methods—with concrete data science examples. Explains abstraction, encapsulation, and polymorphism, using scikit-learn's unified `.fit()` interface as a canonical example. Briefly contrasts functional programming paradigms. - **Late (~47%–100%)**: Moves to production concerns: documentation at multiple levels (comments, docstrings, READMEs), packaging and sharing code, version control with Git, dependency management, building APIs with FastAPI, CI/CD with GitHub Actions, Docker deployment, security risks (including ML-specific threats), and working effectively in software teams with Agile practices. ## 【Key Takeaways】 - **Complexity is the enemy of maintainable data science code** (Early): When a change breaks something unrelated—like forgetting a preprocessing step in inference code—your system has become accidentally complex. Distinguish this from essential complexity inherent to your ML problem. - **Correctness before optimization** (Early): The most important thing is that your code solves the problem and returns expected outputs. Apply performance techniques only after code is working correctly—premature optimization is a common trap. - **Profiling tools reveal where time and memory actually go** (Early): `timeit` measures execution time, `cProfile` shows function-level CPU usage, and `Memray` generates flamegraphs for memory. These tools tell you where to focus optimization efforts instead of guessing. - **Big O notation is an approximation, not an exact measure** (Early): O(n) vs. O(n²) helps compare approaches at scale. Python lists have O(1) append but O(n) insert/delete due to contiguous memory; NumPy arrays trade mixed-type flexibility for significant performance gains. - **NumPy views vs. copies change behavior subtly** (Early–Middle): Slicing a NumPy array creates a view, not a copy—mutating the view changes the original. This is faster and more memory-efficient but requires awareness to avoid bugs. - **OOP with polymorphism simplifies ML code** (Middle): scikit-learn's classifiers all share a `.fit()` method, letting you swap models without rewriting surrounding code. This is the practical payoff of understanding classes, encapsulation, and interfaces. - **Documentation is a multi-level investment** (Late): From inline comments to docstrings to READMEs, each level serves different readers. Clean code signals quality standards—the Broken Window Theory applies: untidy code invites more untidiness. - **Production requires packaging, APIs, and automation** (Late): Turning scripts into Python packages, exposing functionality via FastAPI, and automating deployment with CI/CD and Docker are the steps that take data science from notebook to scalable system. ## 【Reading Tips】 - **Skim the OOP and functional programming sections** if you already write Python classes; focus instead on the scikit-learn polymorphism example and how it applies to your own code design. - **Deep-read the performance chapters** (Big O, lists vs. NumPy, profiling tools)—these have the most immediate payoff for everyday data science work and are rarely covered in intro courses. - **Use the profiling tools as you read**: Run `timeit`, `cProfile`, and `Memray` on your own scripts to internalize where bottlenecks actually occur rather than just reading about them. - **Treat the late chapters as a deployment checklist**: Version control, packaging, APIs, CI/CD, Docker, and security are each deep topics; read them to know what exists, then return when you need to implement a specific piece. - **Pay attention to the "Data in This Book" examples** (UN Sustainable Development Goals data) used throughout—they provide consistent, reproducible context for the techniques. ## 【Coverage Limits】 This guide synthesizes the first half of the book (through OOP and functional programming). The later chapters on version control, APIs, CI/CD, Docker, security, and team practices are summarized from the table of contents but not detailed from excerpts. ##
Page 5
ial Consulting, LLC Proofreader: Krsta Technology Solutions Indexer: WordCo Indexing Services, Inc. Interior Designer: David Futato Cover Designer: Karen Mon...
View in text
Excerpt 2
using to see commented out sections in someone else’s code. When you see untidy sections of code, it sends a message that poor code quality is acceptable in...
View in text
Excerpt 3
ortional to 2 to the factor of the size of the dataset, and recursive algorithms often give you a time complexity of O(2n). Logarithmic time means that the r...
View in text
Excerpt 4
an instance of the class. You can learn more about these in Introducing Python by Bill Lubanovic (O’Reilly, 2019). Abstraction is closely linked to encapsula...
View in text
Excerpt 5
log messages that state a task has happened and the result. An example of this could be logging the accuracy of a machine learning model when it has finished...
View in text
Excerpt 6
e is working correctly for at least these input data values. If your test fails, you need to check two things. First, confirm that your test is accurate and...
View in text
Excerpt 7
el_analysis.py │ └── test_utils.py Let’s break this down. First, there are some standard files that every project should have: ├── README.md ├── requirements...
View in text
Excerpt 8
d comment adds caveats, summarizes information, or explains something not already in the
View in text
Tags
AI categories
ProgrammingPythondata science
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 400
File Format: PDF
File Size: 5.9 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…