Whether you're part of a small startup or a multinational corporation, this practical book shows data scientists, software and site reliability engineers, product managers, and business owners how to run and establish ML reliably, effectively, and accountably within your organization. You'll gain insight into everything from how to do model monitoring in production to how to run a well-tuned model development team in a product organization.
By applying an SRE mindset to machine learning, authors and engineering professionals Cathy Chen, Kranti Parisa, Niall Richard Murphy, D. Sculley, Todd Underwood, and featured guest authors show you how to run an efficient and reliable ML system. Whether you want to increase revenue, optimize decision making, solve problems, or understand and influence customer behavior, you'll learn how to perform day-to-day ML tasks while keeping the bigger picture in mind.
You'll examine:
• What ML is: how it functions and what it relies on
• Conceptual frameworks for understanding how ML "loops" work
• How effective productionization can make your ML systems easily monitorable, deployable, and operable
• Why ML systems make production troubleshooting more difficult, and how to compensate accordingly
• How ML, product, and production teams can communicate effectively
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical guide to running machine learning systems that stay trustworthy in production, not just in notebooks. It's for data scientists, SREs, software and product engineers, and managers who need ML to be monitorable, deployable, and accountable at organizational scale.
【Book Arc】
- **Opening (~0%–10%)**: Sets the premise—most ML material teaches how to build a model, but almost nothing teaches how to build a reliable ML *system*. Introduces SRE as the lens and previews the lifecycle, data, and feature topics ahead.
- **Early (~10%–30%)**: Frames the ML lifecycle and the SRE mindset: SLOs, launch monitoring, golden signals, basic model-health signals, and feedback loops. Establishes that ML systems inherit all distributed-system failure modes plus novel ones.
- **Early–Middle (~30%–50%)**: Moves into data management principles—collection policy, privacy, PII, pseudonymization, access control, bias and fairness processes, dataset augmentation, and the tension between personalization and private data.
- **Middle (~50%–65%)**: Covers features and training data: feature selection and engineering, the lifecycle of a feature, feature systems, labels (including human-generated labels and annotation quality), and a basic model-creation workflow.
- **Late (~65%–90%)**: Productionization and operations—making ML monitorable, deployable, and operable; why ML troubleshooting is harder than ordinary system debugging; and how ML, product, and production teams communicate.
- **Ending (~90%–100%)**: Case studies and takeaways (e.g., ad click prediction "databases versus reality," testing and measuring dependencies in ML workflows) that consolidate the reliability lessons.
【Key Takeaways】
- **Reliability is the overlooked ML attribute** (Early): The book's central claim is that building a model and building a reliable ML *system* are different problems, and SRE thinking—holistic, sustainable, customer-focused—fills the gap.
- **ML systems are systems first** (Early): They share all the failure modes of distributed systems plus novel ones, so don't skip the basics—golden signals, process health, data arrival—while chasing model sophistication.
- **Monitoring needs layered signals** (Early): Distinguish system health (golden signals) from basic model health (generic ML signals) from domain-specific signals; each answers a different question about whether things are working.
- **Data decisions are policy decisions** (Middle): Whether and how to collect data, handle PII, pseudonymize, and restrict access is governance, not just engineering—and access should be granular, logged, and justified.
- **Bias requires an explicit process** (Early–Middle): Bias enters at many stages and can't be fully prevented; the practical first step is adopting a repeatable process (e.g., Model Cards) and reviewing it continuously.
- **Features and labels are first-class artifacts** (Middle): Feature engineering, feature lifecycles, feature systems, and label quality (especially human annotation) determine whether a model can be trusted in production.
- **Productionization is what makes ML operable** (Late): Effective productionization is what makes systems monitorable, deployable, and operable—without it, models remain experiments.
- **Troubleshooting ML is harder, so compensate** (Late): ML systems make production debugging more difficult; the book argues for deliberate practices and cross-team communication to offset that.
【Reading Tips】
- **Deep-read the early lifecycle and SRE chapters** (~10%–30%): this is where the book's distinctive framing lives; skim the preface and table-of-contents material.
- **Treat the data and feature chapters as reference** (~30%–65%): useful when designing your own pipelines, but dense; read for principles rather than memorizing specifics.
- **Read the case studies at the end carefully** (~90%–100%): they show how the abstract reliability principles play out in real incidents and workflows.
- **Bring your own system**: the book is most valuable if you map each recommendation (SLOs, monitoring layers, access controls, bias processes) onto your current ML stack as you read.
- **Don't expect model architecture guidance**: the excerpts note architecture selection is largely out of scope, so pair this with a modeling-focused text if you need that.
【Coverage Limits】
This guide is based on stratified excerpts (33 indexed chunks, 22 sampled) that include the preface, table of contents, and portions of the early and middle chapters; later chapters and case studies are only partially represented, so specific chapter-level detail beyond the excerpted material is not covered.
Excerpt 1
34 Version Control 35 Performance 36 Availability 36 Data Integrity 36 Security 37 Privacy 37 Policy and Compliance 40 Conclusion 41 3. Basic Introduction to...
n-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and...
nymization requires access to an additional data or system. This protects the data from casual inspection by engineers working on the pipeline but permits di...
ring, using, and eventually deleting that private data. The most thorough structural approach generally requires creating per user datastores that are encryp...
swers to some of the important questions listed previously, but we will also dive deeply into specific areas of this example in later chapters. Yarn Product...
ble place to store completed labels is in the feature store. By treating human annotations as their own columns, we can take advantage of all the other funct...
if positives are just a small fraction of the overall data. That said, it is important to notice that the metrics are in tension with each other in an intere...
de-identified information. For example, for most Americans, it turned out to be possible to identify them in a de-identified dataset merely from knowing thei...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Reliable Machine Learning Applying SRE Principles to ML in Production (Cathy Chen, Niall Richard Murphy, Kranti Parisa etc.)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Reliable Machine Learning Applying SRE Principles to ML in Production (Cathy Chen, Niall Richard Murphy, Kranti Parisa etc.)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment