Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Benjamin Labaschin, Jim Allen Wallace, Andrew Brookins, Manvinder Singh

Rating No ratings yet

As AI agents become increasingly essential to daily workflows, a major limitation on their usefulness is their inability to retain and meaningfully recall information across time. While today's agents excel at processing vast amounts of data within a single conversation, they suffer from digital amnesia—forcing users into endless loops of re-explanation and lost context. And unlike traditional databases with predictable storage and retrieval, agent memory operates in a non-deterministic world where the same query might pull different information based on subtle changes in phrasing.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical field guide to why AI agents forget, and how to engineer memory that persists across sessions without drowning in cost or complexity. Best for engineers, architects, and technical leads building or evaluating agentic systems who need to reason about memory as an architectural decision, not a feature toggle. 【Book Arc】 - **Opening (~0%–10%)**: Frames the core problem—agents suffer "digital amnesia," losing context across sessions, and unlike deterministic databases they retrieve fuzzily through semantic space, so the same query can yield different results. - **Early (~10%–32%)**: Builds the technical foundation: memory classifications (sensory, short-term/working, long-term with episodic/semantic/procedural subtypes), vector embedding and storage, context-window limits, summarization loss, retrieval imprecision, and persistence via checkpointing. - **Middle (~32%–48%)**: Moves into implementation—how long-term memory works across frameworks (LangGraph, Mem0/Mem0g, Redis LangCache, Google ADK MemoryService), vector database alternatives, and enhancement techniques like named entity recognition (NER) for structured, entity-aware retrieval. - **Late (~48% onward)**: Shifts to economics and architecture choices—token costs, model selection, and the build-vs-framework-vs-hosted tradeoff that shapes every downstream memory decision. (Excerpts thin here.) - **Ending**: Extends to collective memory—how teams and organizations share knowledge through agents via Transactive Memory Systems, Zep, and MCP—closing on why getting agentic memory right matters organizationally. (Excerpts thin here.) 【Key Takeaways】 - **Agent memory is nondeterministic, not a database** (Opening): The same query can pull different results based on phrasing; storage is embedded vectors, retrieval is fuzzy semantic search, and relevance is calculated rather than guaranteed. - **Human memory is a useful but limited metaphor** (Early): Agents weigh recency, compress, and forget like humans, but they don't learn continuously—their core knowledge is frozen, so dynamic memory must be simulated through engineering. - **Memory types shape storage and retrieval decisions** (Middle): Episodic (past events), semantic (facts and profiles), and procedural (skills and learned behaviors) memory each demand different implementation approaches, and the field is trending toward flexible hybrid models. - **Summarization and context limits always cost information** (Early): FIFO eviction, intelligent pruning, and LLM summarization all lose detail—a lost negation or case reference can invert meaning—while larger context windows still suffer degraded recall for early-passed information. - **Semantic caching and RAG are complementary retrieval strategies** (Early): RAG constrains the corpus and forces sourced answers; semantic caching prioritizes frequently retrieved content and is cost-effective for shared corpora, though it breaks down in multiturn conversations. - **Framework choice encodes a memory philosophy** (Middle): LangGraph favors simplicity and rapid prototyping; Mem0 keeps concise selective entries; Redis LangCache targets repetitive queries; Google ADK offers enterprise reliability but ties you to its ecosystem. - **NER turns fuzzy memory into structured, queryable knowledge** (Middle): Entity extraction, metadata storage, and entity-aware retrieval let agents answer "What did John say about the budget?" by filtering on person and topic rather than broad keyword search. - **Memory has an economics and a build-vs-buy dimension** (Late): Every stored token costs money, and the choice between custom builds, frameworks, and hosted solutions shapes everything about how memory is managed. 【Reading Tips】 - Deep-read the Early chapters on memory classification and retrieval mechanics—these concepts underpin every later architectural decision and are the book's technical core. - Skim the framework comparison in the Middle if you already have a stack preference; use it as a decision matrix rather than reading linearly. - Pay close attention to the economics and tradeoff chapters (Late)—they're where conceptual understanding converts into budget and architecture choices. - Treat the NER section as a practical toolkit: it's the most concrete bridge from theory to improved retrieval accuracy. - Keep the "human agent" framing in mind throughout—the book repeatedly returns to the idea that the most important agent is the human one. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book; the economics, build-vs-buy tradeoffs, collective memory, and conclusion are only lightly represented, so those sections are summarized from chapter descriptions rather than detailed content.
Page 4
related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s vie...
View in text
Page 9
th: the most important agent will always be the human agent. We are the conductors guiding these powerful orchestras. The systems we’ll explore—from vector d...
View in text
Page 14
ke those in Gemini 2.5, which can handle millions of tokens. We can stuff more information into a model, but we may not have effective recall for the first-p...
View in text
Page 20
tives Each serves varying levels of demand. Pinecone offers managed services that are excellent for scale. Weaviate provides open source with hybrid search c...
View in text
Excerpt 5
he marginal cost curve shifts downward and to the right. In practical terms, each additional unit of task complexity now carries a lower cost, so the interse...
View in text
Excerpt 6
his challenge is the “LLM-as-a-judge” approach. This method involves using a powerful, state-of-the-art LLM to act as an impartial evaluator. The judge LLM i...
View in text
Excerpt 7
t severely derisk the reliance on such tools, especially by enterprise staff who have less experience with generative AI and machine learning engineering. Co...
View in text
Excerpt 8
ence that such centralized systems of knowledge benefit not only existing team members but newly onboarded and novice employees as well. In a recent 2023 cal...
View in text
Tags
AI categories
Artificial IntelligenceAIData
ai
Publish Year: 2025
Language: English
File Format: PDF
File Size: 1.4 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…