Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: SAM BHAGWATMICHELLE GIENOW

Rating No ratings yet

Here at Mastra, the open source Typescript framework forbuilding AI agents, we’ve had a front-row seat to how peopleare building agents.Back in February 2025 (practically ancient times, in AIworld), we published the popular guide Principles of BuildingAI Agents.In May, we updated the guide to include MCP, agenticRAG, and a few other emerging principles.But our work was far from finished.2025 is the year of agents and, over the summer, we began tosee a set of stories and guides emerging from prominent AIcompanies, model labs, and early-stage AI startups.The people pushing agents into production werepublicly describing their successes (and failures).Principles was textbook-style knowledge. This wasmessier and rough, expressed in ah-ha! moments, lessonslearned, and retrospectives — knowledge that sprawledvi Introductionacross social media posts, Substacks, eng blogs, and Gitrepos.Collecting and wrangling these lessons for our userseventually led to this book: Patterns for Building AI Agents —now Volume 2 of an eventual trilogy. Principles are conceptual, patterns are pragmatic.While Principles of Building AI Agents covered what to build,Patterns covers how to build.Principles will get you through the first few weeks ofbuilding, but Patterns should be on your desk until itscontents are imprinted in your mind.We start off by sharing patterns for agent design andarchitecture, then dive into the art and science of contextengineering. Next, we dig into the discipline of evals, thestandard way for iterating on and refining agent quality.Finally, we talk about agent security, a field evolving inresponse to novel attack patterns. Agents are in the hands ofearly adopters — and attackers are enthusiastic earlyadopters!Thanks for coming along as we all learn together. Thisbook is a work in progress and a living document.Perhaps as you build, you’ll discover a new pattern thatmakes it into our next edition!

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A pragmatic field guide to shipping AI agents, drawn from the hard-won lessons of teams already running them in production. Read this if you've moved past "what should an agent do?" and are now wrestling with architecture, context, evaluation, and security. 【Book Arc】 - **Opening (~0%–9%)**: Frames the book's purpose — moving from the conceptual "what to build" of *Principles* to the pragmatic "how to build" of *Patterns*, and previews the four-part structure: agent design, context engineering, evals, and security. - **Early (~9%–30%)**: Part I, "Configure Your Agents." Covers turning a wishlist of capabilities into a coherent architecture: whiteboarding capabilities, evolving from monolithic mega-agents to grouped specialists, dynamic runtime configuration, and human-in-the-loop review. - **Middle (~30%–43%)**: Part II, "Engineer Agent Context." Tackles the Goldilocks problem of context — parallelizing carefully, sharing context between subagents, avoiding failure modes, compressing context, and feeding errors back into the loop. - **Late (~43%–52%)**: Part III, "Evaluate Agent Responses." Shifts from building to measuring: listing failure modes, defining critical business metrics, cross-referencing the two, and iterating against evals with SME-labeled and production datasets. - **Ending (~52%+)**: The excerpts trail off mid-Part III; the introduction also promises a final section on agent security, but the excerpts do not cover its contents. 【Key Takeaways】 - **Architecture emerges through iteration, not upfront design** (Early): The book's running example shows a content agent evolving across five iterations from a single writer into a coordinator → router → specialist chain. The lesson is to solve one problem at a time rather than build a "master agent." - **Monolithic mega-agents fail predictably** (Early): The more tools an agent has, the higher the odds it picks the wrong one. Grouping capabilities by department, task type, or business process keeps each agent focused and reliable. - **Dynamic agents trade simplicity for customization** (Early): Runtime signals like user tier, preferences, or system state can drive tool selection, memory, and model choice — but this introduces complexity in logic, testing, and consistency. - **Full autonomy is often untenable; design for human fallback** (Early): Heterogeneous agent performance and legal/ethical stakes mean humans should review risky or low-confidence outputs. Patterns include post-processing review, in-the-loop confirmation, and deferred tool execution. - **Context engineering is both art and science** (Middle): It draws on prompt engineering, RAG, tool calling, and agent state, but also on intuition about how LLMs behave. Too little context starves the agent; too much derails it. - **Parallel subagents create incompatible outputs** (Middle): When subagents work in isolation, they can produce conflicting intermediate results. Solutions include single-threaded linear threads, sequential execution, or sharing full traces between agents. - **Compression and error feedback keep long-running agents coherent** (Middle): Periodic context compression (at thresholds, at tool-call boundaries, or via recursive summarization) prunes irrelevant tokens, while feeding error messages back into context lets agents self-correct instead of crashing. - **Evals are the discipline that turns an MVP into a production agent** (Late): Because LLM outputs are nondeterministic, teams must classify failure modes, define business-aligned success metrics, cross-reference the two, and iterate with SME-labeled and production datasets. 【Reading Tips】 - **Deep-read Part I if you're starting a new agent.** The whiteboarding and architecture-evolution patterns are the most immediately actionable and require no prior infrastructure. - **Treat Part II as a reference during debugging.** Context failure modes, compression strategies, and error-feedback patterns are best consulted when you hit specific reliability problems, not read cover-to-cover. - **Don't skip Part III even if you're an engineer.** The eval workflow (SME labeling → PM prioritization → engineering iteration → validation) is the book's bridge between technical metrics and business outcomes. - **Watch for the running examples.** The content-agent iterations and the medical claim review agent carry the most concrete lessons; the excerpts are thin on other case studies. - **Note the disagreements.** The book flags that not everyone agrees on context-sharing approaches (e.g., Devin vs. Claude Code), so treat patterns as options, not dogma. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book (through Part III's eval workflow). The promised security section and the remainder of Part III are not represented in the excerpts, so their specific patterns are not summarized here.
Page 7
ether. Thisbook is a work in progress and a living document.Perhaps as you build, you’ll discover a new pattern thatmakes it into our next edition! for Build...
View in text
Page 13
rned by the same API call Figure out the natural divisions: Different responsible departments Type of task (e.g., data fetching vs. synthesis vs. triggering...
View in text
Excerpt 3
kflow: 1. Agent breaks its work down into multiple parts. 2. Initiates subagents to work on those parts. 3. Combines those results in the end. Frequently, th...
View in text
Excerpt 4
inistic Most software tests have clear pass/fail conditions. That isn’t true for AI. Because an LLM can return different results when queried with the exact...
View in text
Excerpt 5
suite — which is often just one test running in a loop with LLM-as-judge1 comparing agent answers to your base- line/benchmark standard on each of the evalua...
View in text
Excerpt 6
s, Iterate Against Your Evals, Have SMEs Label Data PART IV SECURE YOUR AGENTS Traditional software security was built around a predictable model. Humans cli...
View in text
Excerpt 7
tensive. We’re excited about approaches to automate some of this work with specialized eval-writing agents. The patterns we talk about in this book are curre...
View in text
Excerpt 8
in, Hamel. "Hot take: your software engineers are the worst candidates for annotating/labeling AI outputs in most cases." LinkedIn, September 20, 2025. https...
View in text
Tags
AI categories
Artificial IntelligenceAISoftware
Publish Year: 2026
Language: English
File Format: PDF
File Size: 3.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…