Generative AI enables powerful new capabilities, but they come with some serious limitations that you'll have to tackle to ship a reliable application or agent. Luckily, experts in the field have compiled a library of 32 tried-and-true design patterns to address the challenges you're likely to encounter when building applications using LLMs, such as hallucinations, nondeterministic responses, and knowledge cutoffs. This book codifies research and real-world experience into advice you can incorporate into your projects. Each pattern describes a problem, shows a proven way to solve it with a fully coded example, and discusses trade-offs.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Generative AI Design Patterns — Reading Guide
## 【One-Line Pitch】
A practical catalog of 32 battle-tested design patterns for building reliable LLM applications, tackling hallucinations, nondeterminism, and knowledge cutoffs with fully coded examples. Essential reading for AI engineers who want to move beyond prompt tinkering to production-grade systems.
## 【Book Arc】
- **Opening (~0%–10%)**: Establishes the AI engineering paradigm—building on foundational models (GPT-4, Gemini, Llama, etc.) rather than training custom models. Introduces core concepts like in-context learning, the landscape of foundational models, and why design patterns are needed for the unique challenges of generative AI.
- **Early (~10%–23%)**: Covers the first cluster of patterns focused on **controlling generation output**—Logits Masking for strict rule enforcement, Style Transfer via few-shot examples, and evaluation techniques for comparing generated content. Includes hands-on code with Transformers' LogitsProcessor.
- **Early-to-Middle (~23%–39%)**: Shifts to **grounding and retrieval**—addressing hallucinations and citation inability through retrieval-augmented generation (RAG). Covers embedding strategies, hybrid search with structured data, and the limitations of naive chunk retrieval.
- **Middle (~39%–48%)**: Deepens RAG with advanced techniques—HyDE (Hypothetical Document Embeddings) for logical queries, query expansion, node postprocessing, and the "chicken-and-egg" problem of knowing what to retrieve before you know the answer.
- **Late (~48%+ and beyond)**: Moves into **system-level reliability**—evaluation frameworks (Ragas metrics), reflection patterns for self-correction, dependency injection for testable LLM chains, and prompt optimization on example datasets. The book culminates in Deep Search systems that synthesize information across multiple sources.
## 【Key Takeaways】
- **Logits Masking gives you hard guarantees** (Early): Unlike prompt engineering or few-shot examples, intercepting the sampling process lets you enforce rules that cannot be violated—the model literally cannot output disallowed tokens. Trade-off: it only works when valid continuations can be cleanly censored; complex cases may need backtracking.
- **Style Transfer via in-context learning is simple but fragile** (Early): Providing few-shot examples works for many style conversions, but examples can contradict each other, and longer prompts degrade inference speed. When latency matters, fine-tuning a smaller model often beats cramming examples into the context window.
- **Evaluation requires a robust test, not vibes** (Early): Direct measurement (A/B testing content) and matching prompts (pairing similar queries) are practical ways to compare generated outputs. But "teaching to the test" is a real risk—your evaluation must reflect reality or you'll optimize for the wrong thing.
- **RAG solves hallucinations but introduces retrieval problems** (Early-to-Middle): Grounding generation in retrieved chunks prevents fabrication, but naive chunking fails on questions requiring holistic interpretation—you can't retrieve what you don't know you need. This is the fundamental chicken-and-egg problem of RAG.
- **HyDE and query expansion rescue logical queries** (Middle): Generating a hypothetical answer first, then retrieving chunks that match its structure, dramatically improves retrieval for multi-step reasoning questions. Query expansion with domain context (e.g., "Darius III" for "the Persian king") surfaces far more relevant chunks.
- **More retrieved nodes ≠ better answers** (Middle): Simply increasing top_k often compounds errors—you get more irrelevant context that still misses the point. Node postprocessing (compressing chunks to relevant bits, reranking) is what actually improves answer quality.
- **Evaluation must be multi-dimensional and weighted** (Late): A single accuracy score is insufficient. Production RAG systems need weighted composites covering relevance, comprehensiveness, factual correctness, coherence, citation quality, and efficiency—plus a curated dataset of reference answers for the sniff test.
- **Reflection and dependency injection make systems maintainable** (Late): Enabling models to correct earlier responses based on feedback significantly improves reliability in complex tasks, while independently developing and testing each LLM chain component keeps systems robust as dependencies change.
## 【Reading Tips】
- **Skim Chapter 1's model landscape** (~0%–10%): The LMArena leaderboard discussion and model taxonomy are useful context but date quickly. Focus instead on the in-context learning explanation—it's the conceptual foundation for everything that follows.
- **Deep-read the Logits Masking pattern** (~10%–13%): The code example with LogitsProcessor is the clearest illustration of how to enforce hard constraints. If you only implement one "advanced" technique, this is it.
- **Pay special attention to the RAG failure modes** (~39%–48%): The Grand Canyon example (where the model describes all of Northern Arizona instead of the canyon itself) is a perfect illustration of why naive retrieval fails. Understanding this failure mode will save you weeks of debugging.
- **Treat the evaluation chapter as a checklist** (~48%): The Ragas metrics list (relevance, comprehensiveness, accuracy, coherence, citation quality, efficiency) is worth copying into your project docs. You'll thank yourself later.
- **Skip the code you don't need, but read every "Considerations" and "Alternatives" section**: The trade-off discussions are where the real wisdom lives—each pattern's limitations are as important as its implementation.
## 【Coverage Limits】
This guide covers the first ~48% of the book in detail (patterns 1–20 roughly), including generation control, RAG fundamentals, and evaluation. The later patterns (21–32) covering advanced agent architectures, multi-modal patterns, and deployment considerations are not covered in the sampled excerpts.
##
Page 15
model an appropriate text input, which is known as a prompt. However, you will face certain common problems—the generated content may not match the style you...
onstraint is therefore working properly. The pipe separator Suppose that you want the LLM to extract three pieces of information and output them, separated f...
on relevance, then the top two nodes will have directly TIP A number of ML libraries offer guardrail tooling for LLMs, such as the following: Guardrails AI T...
'text': 'You are a food influencer.'}]}, {'role': 'user', 'content': [{'type': 'text', 'text': 'Suggest 3 ways to improve the flavor of ice cream.'}, {'role'...
e LLM in Step 1 in order to get the Critique object to pass to the second step—so there seems to be no way to develop and test Step 2 independently of Step 1...
ic cache that uses federated learning to honor user privacy. Automatic prefix caching was introduced in a class project by Jha and Wang (2023) and implemente...
irst consider whether Pattern 29, Template Generation, will suit your needs—its ability to review all templates provides an extra safeguard. Choose Assembled...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Generative AI Design Patterns (Valliappa Lakshmanan, Hannes Hapke)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Generative AI Design Patterns (Valliappa Lakshmanan, Hannes Hapke)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment