It's a new day in search. Before ChatGPT, combing the web was simple—powerful search engines dominated for 25 years. That changed with conversational search powered by chatbots. Data catalogs, the search engines for your company's data, have evolved as well. The Enterprise Data Catalog explores how AI is transforming enterprise-wide data search.
In this second edition, you'll explore how the role of the data catalog has changed in the age of AI. Data catalogs no longer serve as tools to find and use data—they now deliver essential metadata for AI projects. Author Ole Olesen-Bagneux explains how metadata organized as enterprise ontologies, delivered through knowledge graphs, provides the context required by large language models, Model Context Protocol, and Agent2Agent Protocol. By drawing on data management and library and information science, the book shows why information science methodology is critical to successful catalog implementations.
• Understand the role of the enterprise data catalog in discovery, governance, and AI
• Organize data and sources using metadata, domains, and enterprise ontologies
• Search and browse data across domains, lineage, and graphs
• Leverage data catalog knowledge graphs for AI use cases
• Apply data catalogs to data products and data contracts
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# The Enterprise Data Catalog: Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture
## 【One-Line Pitch】
A practical guide for data leaders and practitioners on how enterprise data catalogs have evolved from simple search tools into the metadata backbone that powers AI initiatives—covering domain architecture, knowledge graphs, and the emerging LLM + knowledge graph pattern. Read this if you're implementing or scaling a data catalog and want to understand its strategic role in AI-driven organizations.
## 【Book Arc】
- **Opening (~0%–9%)**: Establishes the core thesis—data catalogs are the "search engines for your company's data," just as web search engines transformed the internet. The preface recounts the author's journey from realizing this potential to now grappling with how ChatGPT and conversational AI fundamentally change what data catalogs can do.
- **Early (~9%–25%)**: Introduces the foundational architecture of data catalogs: how they crawl IT landscapes, organize metadata into domains and subdomains, and create visibility across the entire organization at the metadata level (without exposing actual data values). Emphasizes that how you organize data determines how effectively you can search it.
- **Early–Middle (~25%–38%)**: Explores the search functionality in depth, positioning data catalogs as both search engines and AI assistants. Introduces the concept of "ambient findability"—the target state where anyone can find any data asset from anywhere at any time—and distinguishes between searching *for* data (discovery) versus searching *in* data (data science).
- **Middle (~38%–47%)**: Covers the human and organizational dimensions: the data discovery team (not just a "catalog team"), the ideal placement under a Chief Data Officer, and the three categories of end users (data analytics/AI, governance, and everyday users). Details specialized roles like asset stewards, term owners, and term stewards.
- **Late (~47%–end)**: The book's trajectory points toward advanced topics covered in later chapters (not fully available in this excerpt): data products and data contracts (survivors of the data mesh movement), the LLM + knowledge graph pattern, and how data catalogs become a source in themselves for AI systems.
## 【Key Takeaways】
- **Data catalogs are now AI infrastructure, not just search tools** (Early): The role has shifted from helping humans find data to delivering essential metadata context for LLMs, MCP, and agentic architectures. This reframing helps secure executive buy-in for governance and compliance investments.
- **Metadata-level visibility solves the governance paradox** (Early): Data catalogs let everyone see everything about company data—column names, descriptions, lineage—without exposing actual values. This makes comprehensive data overviews possible where raw data access would be technically and legally impossible.
- **Domain architecture is the organizing principle** (Early): Data must "belong somewhere." You design domains and subdomains that vertically organize assets by business function (e.g., finance), making it possible to pinpoint exactly where any data originates.
- **Search is the underutilized superpower** (Early): Most organizations treat search as a feature, not a strategy. Thinking of your data catalog as an enterprise search engine—and increasingly an AI assistant—unlocks its full value for data discovery.
- **Ambient findability is the goal** (Middle): Borrowed from Peter Morville's work on web search, this concept means finding "anyone or anything from anywhere at anytime" within your enterprise data landscape. It's the standard data catalogs should strive toward.
- **Data discovery teams, not catalog teams** (Middle): Implementation, maintenance, and promotion require dedicated cross-functional ownership. The ideal setup places this team under a CDO, where the executive data strategy is grounded in empirical data visibility.
- **End users fall into three categories** (Middle): Data analytics/AI users (who deliver ROI through innovation), governance users (who ensure compliance), and everyday efficiency users. Each has different needs and should be supported accordingly.
- **AI enables natural language interaction with catalogs** (Early): Enterprise AI assistants can now query both catalog metadata and other sources (like Slack) via MCP servers, enabling entirely new patterns of data discovery and contextual understanding.
## 【Reading Tips】
- **Skim the preface** (~0%–9%) if you're already convinced of the data catalog's strategic value; it's most useful for building the business case and understanding the author's intellectual journey.
- **Deep-read Chapter 1** (~9%–47%) for the conceptual foundation: domain organization, metadata principles, search mechanics, and team structures. This is where the practical architecture guidance lives.
- **Pay special attention to the "searching for data vs. searching in data" distinction** (~38%): This clarifies why data catalogs are discovery tools, not analytics platforms—a confusion that undermines many implementations.
- **Note the organizational guidance** (~44%–47%): The discussion of where to place the data catalog team (CDO vs. analytics unit) involves real trade-offs between innovation and governance that you'll likely face.
- **Watch for the AI integration patterns** throughout: The examples of enterprise AI assistants querying catalogs via MCP servers foreshadow the book's later focus on LLM + knowledge graph architectures.
## 【Coverage Limits】
This guide covers the available early-release content (Preface through Chapter 2). The excerpts do not cover later chapters on data products/contracts, the LLM + knowledge graph pattern, standards, or lifecycle management—these are listed in the table of contents but their content is unavailable in this sample.
##
ny about the potential of a data catalog and it was crystal clear in my head, I was then faced with the battle of explaining the features to the important st...
w you to search that landscape. The main difference is that while a web search engine covers the web as a landscape, a data catalog covers your organization’...
expertise on a particular subset of assets (an entire data source or parts of data sources) in a domain. Term owner Term owners typically own a large subpart...
riven Value at Scale (Sebastopol, CA: O’Reilly, 2021). Pie‐ thein Strengholt, Data Management at Scale (Sebastopol, CA: O’Reilly, 2023, second edition), p. 3...
end to create hierarchies between terms, some being broader than others; for example, daily clothes is a narrower term than clothes but broader than daily cl...
d the DPO have correctly classified the data—it’s often the case that data is classified as highly confidential and not sensitive at all, at the same time. T...
L. The difference lies in what data layer the languages are applied on: you use DQLs to search in data, and you use IRQLs to search for data. To match DQL an...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
The Enterprise Data Catalog Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture (Early Release) (Ole Olesen-Bagneux)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
The Enterprise Data Catalog Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture (Early Release) (Ole Olesen-Bagneux)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment