Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Katharine Jarmul

Rating No ratings yet

Between major privacy regulations like the GDPR and CCPA and expensive and notorious data breaches, there has never been so much pressure to ensure data privacy. Unfortunately, integrating privacy into data systems is still complicated. This essential guide will give you a fundamental understanding of modern privacy building blocks, like differential privacy, federated learning, and encrypted computation. Based on hard-won lessons, this book provides solid advice and best practices for integrating breakthrough privacy-enhancing technologies into production systems.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Practical Data Privacy ## 【One-Line Pitch】 A hands-on guide for data scientists, engineers, and architects who want to move beyond checkbox compliance and actually build privacy-preserving systems using modern technologies like differential privacy, federated learning, and encrypted computation. If you design, build, or test systems that handle personal data, this book gives you the mental models and practical tools to make privacy a first-class engineering concern. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces the concept of privacy engineering as a growing discipline that blends data science with systems architecture. Defines what counts as sensitive data—expanding beyond obvious PII to include person-related data and proprietary information—and establishes the core questions around data security and privacy documentation. - **Early (~10%–23%)**: Lays the theoretical foundation for differential privacy, explaining why older "anonymization" approaches fail and why process-based guarantees matter. Uses the US Census Bureau's re-identification attacks as a cautionary tale about why result-focused anonymization is a fallacy. - **Early (~23%–32%)**: Walks through the mechanics of differential privacy—the formal definition, the role of noise, epsilon budgets, and how to allocate limited privacy budgets across multiple queries. Includes toy implementations and practical guidance on choosing between Laplace and Gaussian noise mechanisms. - **Middle (~39%–48%)**: Moves into production practice with concrete examples: building privacy-aware data pipelines, using libraries like Great Expectations for data validation, and implementing differential privacy with Tumult Analytics in Spark sessions. Introduces the concept of trust boundaries and how to think about data movement across security domains. - **Late (~48%–end)**: Addresses the harder problem of re-identification risk in expansive datasets, using cardinality analysis and hashing techniques to detect when supposedly anonymized data can actually be linked back to individuals. The book closes with organizational advice on embedding privacy practices into team workflows. ## 【Key Takeaways】 - **Privacy engineering is a distinct discipline** (Early): It sits at the intersection of data science and software engineering, requiring you to architect solutions rather than just explore data. This means working directly with data engineering, application teams, and architects to ensure privacy is built into both products and internal systems. - **Sensitive data is broader than you think** (Early): Beyond obvious PII, sensitive data includes person-related information (interests, locations, behaviors) and proprietary/confidential business data. Context matters—your phone location reveals your home address when you're at home, making it personally identifiable in ways you might not expect. - **Differential privacy focuses on process, not results** (Early): Older anonymization methods try to inspect output data and judge if it's "safe enough"—a fallacy proven by re-identification attacks. Differential privacy instead guarantees bounds on privacy loss through the algorithm itself, allowing you to measure and control risk over time. - **Epsilon is a budget, not a magic number** (Early): You allocate your privacy budget across queries like any limited resource, spending more on important calculations and less on exploratory ones. Libraries can track this automatically, but you need to understand the trade-offs to make intelligent allocation decisions. - **Noise choice matters for downstream analysis** (Early): Gaussian noise is often preferable to Laplace noise because scientific workflows typically assume normal error distributions. There's no "ground truth" data—you're modifying one error-prone version of reality into another, so choose mechanisms that align with your analysis approach. - **Production pipelines need privacy-aware testing** (Middle): Tools like Great Expectations can validate that your privacy transformations actually worked—for example, checking that outliers were removed from price columns. Testing isn't optional; it's how you know your pipeline is properly protecting data. - **Trust boundaries define where privacy guarantees hold** (Middle): A trust boundary is where security changes fundamentally—like when user input crosses from a web form to an application server, or when data moves from a secure environment to a less secure one. Understanding these boundaries is essential for knowing where your privacy protections apply. - **Cardinality analysis reveals hidden re-identification risk** (Late): Large datasets with many unique combinations of seemingly innocuous attributes (browser agents, app settings, locations) can leak identity even when individual fields seem harmless. Hashing-based techniques can help you detect when data is easily linkable before you release it. ## 【Reading Tips】 - **Skim the math, focus on the intuition**: The differential privacy chapters include formal definitions and proofs, but the author explicitly says you don't need to master the math. Focus on understanding why the mechanisms work and how epsilon budgets behave, then move on. - **Do the toy implementations**: The book includes runnable examples (like the Laplace mechanism for average age) that make abstract concepts concrete. Run them, tweak the parameters, and see how noise affects results—this will build intuition faster than reading alone. - **Jump to the production chapters if you're an engineer**: If you're already comfortable with privacy fundamentals, the middle-to-late chapters on pipeline implementation, library usage (Tumult Analytics, Great Expectations), and cardinality analysis are where the practical value lives. - **Use the book repository**: The author references Jupyter notebooks and code examples throughout. Download them and adapt the examples to your own data and frameworks—this is a book meant to be worked through, not just read. - **Watch for the trust boundary concept**: This is a mental model worth internalizing. When you're designing systems, map out where data crosses trust boundaries and ensure privacy protections are explicitly considered at each crossing. ## 【Coverage Limits】 The excerpts cover the book's opening through roughly the halfway point in detail, with lighter coverage of the later chapters on federated learning and encrypted computation. The guide focuses on the differential privacy and data governance material that is well-documented in the source excerpts. ##
Page 18
architects at your company to ensure privacy is built into the product as well as the internal applications. This covers all consumer and employee data flows...
View in text
Excerpt 2
rsonal sensitive. identifiers are removed or pseudonymized. In my experience, the biggest argument against using pseudonymization would be that it creates a...
View in text
Excerpt 3
ber of contributions a user can make, inadvertently running computations more than once, or even incorrectly calculating sensitivity. Most important —now you...
View in text
Excerpt 4
ave a few good examples of workflows that have added proper privacy measures, ensure they are properly documented. Then, link to them or give a short interna...
View in text
Excerpt 5
ve been significant advances since the PATE paper in adding differential privacy noise during training. As you learned in Chapter 2, you can measure sensitiv...
View in text
Excerpt 6
e you are not looking at individual data points. If you are familiar with the concept of independent and identically distributed (IID) random variables, you...
View in text
Excerpt 7
rger computation, of which multiplication is only one piece. The building blocks in this chapter demonstrate how these protocols work but are not meant to be...
View in text
Excerpt 8
in action for private inference, where the model and values remain encrypted, take a look at the Zama + Hugging Face notebook and the Yann Dupis’s Moose + Hu...
View in text
Tags
AI categories
data privacyprivacy engineeringdifferential privacy
ISBN: 1098129466
Publish Year: 2023
Language: English
Pages: 500
File Format: PDF
File Size: 8.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…