With this practical book, AI and machine learning practitioners will learn how to successfully build and deploy data science projects on Amazon Web Services. The Amazon AI and machine learning stack unifies data science, data engineering, and application development to help level upyour skills. This guide shows you how to build and run pipelines in the cloud, then integrate the results into applications in minutes instead of days. Throughout the book, authors Chris Fregly and Antje Barth demonstrate how to reduce cost and improve performance.
• Apply the Amazon AI and ML stack to real-world use cases for natural language processing, computer vision, fraud detection, conversational devices, and more
• Use automated machine learning to implement a specific subset of use cases with SageMaker Autopilot
• Dive deep into the complete model development lifecycle for a BERT-based NLP use case including data ingestion, analysis, model training, and deployment
• Tie everything together into a repeatable machine learning operations pipeline
• Explore real-time ML, anomaly detection, and streaming analytics on data streams with Amazon Kinesis and Managed Streaming for Apache Kafka
• Learn security best practices for data science projects and workflows including identity and access management, authentication, authorization, and more
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical, end-to-end field guide for AI/ML practitioners who want to build, train, deploy, and secure machine learning pipelines on AWS, covering everything from data ingestion to MLOps with real-world examples.
【Book Arc】
- **Opening (~0%–13%)**: Introduces the AWS AI/ML stack, the benefits of cloud computing, and the concept of MLOps. Sets the stage for the book’s core promise: unifying data science, data engineering, and application development into a single workflow.
- **Early (~13%–31%)**: Surveys a wide range of real-world use cases—from product recommendations and fraud detection to NLP and predictive maintenance—and introduces SageMaker Autopilot for automated machine learning. This section helps readers map business problems to AWS services.
- **Middle (~31%–50%)**: Dives into the data lifecycle: ingesting data into data lakes (S3, Athena, Redshift), exploring and visualizing datasets (SageMaker Studio, QuickSight), and preparing data for training (feature engineering, SageMaker Processing, Feature Store). Concludes with a deep dive into training a BERT-based NLP model, including fine-tuning, debugging, and interpretation.
- **Middle (~50%–63%)**: Covers scaling model training (hyperparameter tuning, distributed training) and deploying models to production (real-time endpoints, batch transform, Lambda, edge deployment). Also introduces SageMaker Pipelines for orchestrating repeatable MLOps workflows.
- **Late (~63%–69%)**: Explores streaming analytics and real-time ML with Amazon Kinesis and Kafka, then shifts to security best practices—IAM, encryption, network isolation, and governance—to close the book with a production-ready mindset.
【Key Takeaways】
- **MLOps is the backbone of production ML** (Early): The book emphasizes that successful data science isn’t just about modeling—it requires disciplined pipelines, monitoring, and automation. Expect to learn how to tie together training, deployment, and monitoring into a repeatable workflow.
- **SageMaker Autopilot democratizes AutoML** (Early): For a subset of use cases, you can automate model selection, training, and tuning without writing custom code. The book walks through a text classifier example to show how to track experiments and deploy results.
- **Data lakes are the foundation for scalable analytics** (Middle): Learn to build and query a data lake on Amazon S3 using Athena and Redshift Spectrum, and how to choose between Athena and Redshift based on cost and performance trade-offs.
- **Feature engineering is a team sport** (Middle): SageMaker Processing Jobs, Feature Store, and Data Wrangler are introduced to scale feature engineering and share features across teams, while tracking artifact and experiment lineage for reproducibility.
- **BERT is the canonical deep-dive example** (Middle): The book dedicates significant space to fine-tuning a pre-trained BERT model for NLP, covering the Transformer architecture, training scripts, debugging with SageMaker Debugger, and model interpretability.
- **Deployment is more than just hosting** (Middle): Real-time endpoints, batch transforms, and edge deployment are covered, along with auto-scaling, model versioning, and monitoring for drift in data quality, model quality, and bias.
- **Streaming ML enables real-time decisions** (Late): Using Kinesis and Kafka, you can build pipelines that classify product reviews or detect anomalies in real time, with windowed queries and online learning concepts explained.
- **Security is a shared responsibility** (Late): The book closes with a thorough treatment of IAM, encryption at rest and in transit, securing SageMaker notebooks and jobs, and governance/auditability—essential for any production deployment.
【Reading Tips】
- **Skim Chapter 2** (use cases) if you already know your problem domain; it’s a catalog of examples rather than deep technical content. Return to it later as a reference for matching services to scenarios.
- **Deep-read Chapters 4–7** for the data lifecycle and BERT training walkthrough—this is the heart of the book. Follow along with the code if you have AWS access; the hands-on examples are the real value.
- **Pay attention to "Reduce Cost and Increase Performance" sections** at the end of most chapters—they contain practical, often overlooked tips on instance selection, tagging, and budgeting.
- **Treat Chapter 10 (Pipelines and MLOps) as the capstone**: If you’re short on time, read this after Chapter 7 to understand how the pieces fit together before diving into deployment details.
- **Use the book as a reference, not a cover-to-cover read**: The table of contents is detailed enough to jump to specific topics (e.g., security, streaming) as needed.
【Coverage Limits】
This guide synthesizes the book’s table of contents and early endorsements; the excerpts do not include detailed code samples, specific AWS console walkthroughs, or the full text of later chapters. For hands-on implementation, refer to the book’s companion code repository.
Page 3
ld examples to get you started on your data science journey." —Jeff Barr, Vice President & Chief Evangelist, Amazon Web Services “It’s very rare to find a bo...
and services that could be used for each step along the way.” —Rustem Feyzkhanov, Machine Learning Engineer, Instrumental, Amazon Web Services ML Hero “Chris...
ze and Manage Models at the Edge 357 Table of Contents | ix Deploy a PyTorch Model with TorchServe 357 TensorFlow-BERT Inference with AWS Deep Java Library 3...
or anyone who uses data to make critical business decisions. The guid‐ ance here will help data analysts, data scientists, data engineers, ML engineers, rese...
buting examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. I...
to the way we will present technical material in the future. You helped elevate this book from good to great, and we really enjoyed working with you all on t...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Data Science on AWS Implementing End-to-End, Continuous AI and Machine Learning Pipelines (Chris Fregly, Antje Barth)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Data Science on AWS Implementing End-to-End, Continuous AI and Machine Learning Pipelines (Chris Fregly, Antje Barth)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment