Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Dominik Polzer

Rating No ratings yet

As businesses race to unlock the full potential of large language models (LLMs), a critical challenge has emerged: How do you connect these tools to real-time, external data to solve real-world problems? Retrieval-augmented generation (RAG) is the answer. By combining LLMs with information retrieval, RAG empowers you to build everything from intelligent chatbots to autonomous, task-solving agents. Packed with over 70 practical recipes, this go-to guide tackles a wide range of GenAI applications through structured hands-on learning. Author Dominik Polzer provides the tools you need to design, implement, and optimize RAG systems for your unique use cases. Whether you're working with simple data retrieval or designing cutting-edge autonomous agents, this cookbook will help you stay ahead of the curve. Learn core RAG components including embedding, retrieval, and generation techniques Understand advanced workflows like semantic-aware chunking and multi-query prompting Build custom solutions such as chatbots and autonomous agents for specific data challenges Continuously evaluate and optimize systems for accuracy, relevance, and performance

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# RAG-Ready Patterns for Data Platforms ## 【One-Line Pitch】 A practical, recipe-driven guide for building production-grade Retrieval-Augmented Generation (RAG) systems in Python—from data ingestion and chunking strategies to agentic architectures and evaluation—ideal for developers and data engineers who want hands-on patterns rather than theory. ## 【Book Arc】 - **Opening (~0%–10%)**: Sets up the RAG landscape—why RAG matters for connecting LLMs to real-time external data, how to choose your IDE and coding environment, and the core three-component architecture (embedding model, vector store, LLM). Includes a survey of the RAG library ecosystem (LangChain, LlamaIndex, Chroma, etc.) and a first end-to-end chatbot build. - **Early (~10%–23%)**: Covers foundation model selection—comparing OpenAI, Google, Anthropic, and local models via Ollama—with practical guidance on context windows, latency trade-offs, and cost considerations. Emphasizes that model tier choice matters more than provider price differences. - **Early (~23%–32%)**: Dives into data loading, addressing the reality that ~80% of enterprise information is unstructured. Covers PDF parsing, OCR (Tesseract, EasyOCR, PaddleOCR), multimodal approaches for complex documents, and structured data loading from PostgreSQL—including a decision guide for OCR versus multimodal models based on document type and volume. - **Middle (~32%–48%)**: Focuses on data preparation and chunking strategies. Explains why chunks must stand alone without surrounding context, covers metadata filtering to prevent false positives from overlapping terminology, and introduces progressive chunking techniques: recursive, semantic (topic-based grouping), and agentic (decontextualizing and atomizing text into standalone propositions). - **Late (~48%–100%)**: Extends into advanced retrieval patterns (auto-merging retrievers, sentence window retrievers, reranking, multi-query decomposition), agentic RAG (custom tools, workflow patterns, function calling, asyncio acceleration, MCP tools, LangGraph), Graph RAG with Neo4j knowledge graphs, and systematic evaluation of RAG systems. ## 【Key Takeaways】 - **RAG's core value is grounding LLMs in external data** (Early): The three essential components—embedding model, vector store, and LLM—form a pipeline where documents are chunked, embedded, and stored, then retrieved by similarity and passed to the LLM as context. This solves the problem of LLMs lacking access to real-time or proprietary information. - **Model selection should prioritize tier over provider** (Early): Cost differences between model tiers range from $0.05 to $30 per 1M input tokens, but provider price variations within the same tier are marginal. Choose based on reasoning depth, latency needs, and context window requirements—not on small price differences between OpenAI, Google, or Anthropic. - **Unstructured data handling requires a format-specific strategy** (Early): With ~80% of enterprise data unstructured, you need a decision framework: simple text PDFs under 1,000 docs/month → local OCR (Tesseract); mixed content with tables/images → multimodal models; sensitive data → local deployment only. Never send regulated data to external APIs. - **Chunk quality determines retrieval quality** (Middle): Each chunk should contain exactly one piece of information and be understandable independently. The progression from recursive (separator-based) to semantic (topic-based) to agentic (standalone proposition) chunking represents increasing sophistication—agentic chunking is the only approach that actively resolves references and pronouns. - **Metadata filtering prevents semantic false positives** (Middle): Documents can use the same words for different contexts (e.g., "release" in product docs vs. press releases). LLM-generated metadata adds preprocessing cost ($1–$5 and several minutes for 1,000 documents) but significantly improves retrieval precision by separating semantically similar but contextually irrelevant content. - **Hypothetical questions improve retrieval** (Middle): Generate hypothetical questions from text chunks during ingestion, then search those questions at runtime while passing the underlying full text to the LLM. This bridges the gap between how users ask questions and how documents are written. - **Agentic RAG extends beyond simple retrieval** (Late): Building agents requires custom tool design, workflow patterns, and framework choices (OpenAI Agents SDK, LangGraph, etc.). Function calling enables agents to act on retrieved information, and asyncio accelerates concurrent agent operations. - **Graph RAG adds relational understanding** (Late): Knowledge graphs (e.g., Neo4j) capture relationships between entities that vector similarity alone misses. Cypher queries enable structured retrieval, and semantic search can be layered on top for hybrid approaches. ## 【Reading Tips】 - **Skim the opening chapters (1–2)** if you're already familiar with RAG basics—the IDE recommendations and library overviews are useful references but not deep content. Focus instead on the model selection guidance and the Ollama local-model setup. - **Deep-read Chapter 3 (Loading Data)** if you deal with messy real-world documents—the OCR versus multimodal decision table is worth bookmarking as a reference for production decisions. - **Pay special attention to Chapter 4 (Data Preparation)**—chunking strategy is the single highest-leverage decision in RAG system quality. The progression from recursive to semantic to agentic chunking is the conceptual core of the book. - **Treat the later chapters (7–10) as a menu** rather than a linear read—pick the advanced retrieval pattern, agent framework, or evaluation approach that matches your current project. The recipes are self-contained. - **Watch for the cost/latency trade-off notes** scattered throughout—the book is honest about the operational costs of techniques like LLM-generated metadata and multimodal processing, which is rare and valuable in RAG literature. ## 【Coverage Limits】 This guide synthesizes the first ~48% of the book in detail (foundations, data loading, and chunking strategies). The later sections on advanced retrieval, agentic RAG, Graph RAG, and evaluation are mapped from the table of contents but not deeply summarized from the excerpts. ##
Page 7
nto Multiple Subqueries 206 8. Agentic RAG. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 2
the Frameworks and Libraries for Your RAG Applications | 19 This structure guides the model toward grounded responses. By explicitly instructing the model to...
View in text
Excerpt 3
and extracted text (least privilege, auditing, encryption). Start with local OCR (Tesseract, EasyOCR, PaddleOCR). If you need layout-aware extraction (tables...
View in text
Excerpt 4
um size, it looks for the next separator to split the text: from langchain_text_splitters import RecursiveCharacterTextSplitter import PyPDF2 file_path = ".....
View in text
Excerpt 5
uments, reducing retrieval latency and improving precision. Use classification routing when your vector store combines drastically different con‐ tent types...
View in text
Excerpt 6
or, as shown in Figure 6-11. The query vector (shown as the blue dot) is compared to the data points in the selected partition to identify the clos‐ est matc...
View in text
Excerpt 7
mpt template that instructs the LLM to rerank the retrieved text chunks according to their relevance to the user’s question: import textwrap from openai impo...
View in text
Excerpt 8
stems, reduce this with batching or asynchronous execution, as shown in Recipe 8.5. See Also • The OpenAI function calling guide provides detailed documentat...
View in text
Tags
AI categories
PythonAIBackend
ISBN: 8341600560
Publish Year: 2026
Language: English
Pages: 378
File Format: PDF
File Size: 17.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…