Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Chip Huyen

Rating No ratings yet

Recent breakthroughs in AI have not only increased demand for AI products, they've also lowered the barriers to entry for those who want to build AI products. The model-as-a-service approach has transformed AI from an esoteric discipline into a powerful development tool that anyone can use. Everyone, including those with minimal or no prior AI experience, can now leverage AI models to build applications. In this book, author Chip Huyen discusses AI engineering: the process of building applications with readily available foundation models. The book starts with an overview of AI engineering, explaining how it differs from traditional ML engineering and discussing the new AI stack. The more AI is used, the more opportunities there are for catastrophic failures, and therefore, the more important evaluation becomes. This book discusses different approaches to evaluating open-ended models, including the rapidly growing AI-as-a-judge approach. AI application developers will discover how to navigate the AI landscape, including models, datasets, evaluation benchmarks, and the seemingly infinite number of use cases and application patterns. You'll learn a framework for developing an AI application, starting with simple techniques and progressing toward more sophisticated methods, and discover how to efficiently deploy these applications. Table of Contents: 1. Introduction to Building AI Applications with Foundation Models 2. Understanding Foundation Models 3. Evaluation Methodology 4. Evaluate AI Systems 5. Prompt Engineering 6. RAG and Agents 7. Finetuning 8. Dataset Engineering 9. Inference Optimization 10. AI Engineering Architecture and User Feedback Epilogue Index

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, end-to-end guide for software engineers and technical leaders who want to build production-grade AI applications using foundation models, covering everything from model selection and evaluation to prompt engineering, RAG, finetuning, and deployment. 【Book Arc】 - **Opening (~0%–5%)**: Introduces AI engineering as a discipline distinct from traditional ML engineering, framing the model-as-a-service shift that lowers barriers to entry. Sets up the "new AI stack" and the central role of evaluation in mitigating failure risks. - **Early (~5%–20%)**: Builds foundational knowledge of foundation models—how they work, their capabilities and limitations—and establishes a rigorous evaluation methodology, including metrics for open-ended tasks and the emerging AI-as-a-judge approach. - **Middle (~20%–50%)**: Moves into core application techniques: prompt engineering (zero-shot, few-shot, chain-of-thought), retrieval-augmented generation (RAG), and agent design. Emphasizes a progression from simple to sophisticated methods. - **Late (~50%–80%)**: Covers finetuning and dataset engineering, explaining when to finetune versus prompt, how to curate and clean data, and the trade-offs involved. Includes inference optimization for latency and cost. - **Ending (~80%–100%)**: Ties everything together with AI engineering architecture patterns and user feedback loops, discussing how to design systems that learn from real-world usage and iterate safely. 【Key Takeaways】 - **AI engineering is a new discipline** (Early): Unlike traditional ML engineering, it focuses on composing and orchestrating pre-trained models rather than training from scratch. This shifts the core skill from model building to system design and evaluation. - **Evaluation is the linchpin of AI applications** (Early): Open-ended models fail in unpredictable ways, so rigorous evaluation—using benchmarks, human feedback, and automated judges—is essential. The AI-as-a-judge approach is rapidly growing as a scalable alternative to human annotation. - **Prompt engineering is the first lever to pull** (Middle): Techniques like few-shot prompting, chain-of-thought, and role prompting can dramatically improve output quality without any model changes. Start here before considering more expensive interventions. - **RAG and agents extend model capabilities** (Middle): Retrieval-augmented generation grounds responses in external knowledge, while agents enable multi-step reasoning and tool use. These patterns are the building blocks for most practical AI applications. - **Finetuning is a targeted tool, not a default** (Late): Finetuning is most valuable when you need to change model behavior, style, or domain knowledge. It requires careful dataset engineering—curation, cleaning, and balancing—to avoid degrading general performance. - **Dataset engineering is as important as model choice** (Late): The quality and composition of your training and evaluation data often matter more than the model architecture. Garbage in, garbage out applies with full force to foundation models. - **Inference optimization is a cost and latency game** (Late): Techniques like quantization, distillation, and batching can reduce inference costs by orders of magnitude. These decisions directly impact whether your application is economically viable at scale. - **Architecture and feedback loops close the loop** (Ending): Production AI systems need thoughtful architecture for logging, monitoring, and user feedback. This data is the fuel for continuous improvement—both for prompts and for future finetuning rounds. 【Reading Tips】 - **Skim the opening chapters** (~0–10%) if you already have ML experience; the value is in the framing of AI engineering, not the basics. Deep-read the evaluation chapters (3–4) since they are the book's most distinctive contribution. - **Treat Chapters 5–7 (Prompt, RAG/Agents, Finetuning) as a decision tree**: Read them together to understand when to use each technique. The progression from simple to sophisticated is the book's core practical framework. - **Pay special attention to the AI-as-a-judge material** in the evaluation chapters—this is a fast-moving area and the book's treatment is both current and practical. - **Use the dataset engineering and inference optimization chapters as reference material** rather than reading them cover-to-cover. They are dense with specifics you'll want to consult when you hit those problems. - **The epilogue and architecture chapter are worth a careful read** before you deploy anything; they contain the hard-won lessons about user feedback and system design that most tutorials skip. 【Coverage Limits】 The excerpts focus on the book's structure and high-level themes; detailed technical content, code examples, and specific case studies from the chapters are not covered in this guide.
Excerpt 1
书名: 大语言模型-基础与前沿 [转换版] (熊涛) (Z-Library) 作者: 熊涛 本书深入阐述了大语言模型的基本概念和算法、研究前沿以及应用,涵盖大语言模型的广泛主题,从基础到前沿,从方法到应用,涉及从方法论到应用场景方方面面的内容。 首先,本书介绍了人工智能领域的进展和趋势;其次,探讨了语言模型的基本...
View in text
Excerpt 2
ons and stored procedures in a high-performance environment. Readers will learn how to leverage WebAssembly to optimize data processing tasks, run computatio...
View in text
Excerpt 3
. Thank you for always being there. —Dinesh Kumar Chemuduru v About the Authors ...
View in text
Excerpt 4
1 What Is Artificial Intelligence?
View in text
Excerpt 5
e paper Naga Santhosh Reddy Vootukuri Seattle, WA, USA
View in text
Excerpt 6
ate pipeline development 216 real-time data streaming 215 Database Migration Service (DMS) 213 database passwords 37 delete function 166 deployment process 1...
View in text
Excerpt 7
72 Conclusion 73 Chapter 3...
View in text
Tags
AI categories
AIArtificial IntelligenceBackend
ISBN: 1098166302
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 535
File Format: PDF
File Size: 31.9 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…