Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ole Olesen-Bagneux

It's a new day in search. Before ChatGPT, combing the web was simple—powerful search engines dominated for 25 years. That changed with conversational search powered by chatbots. Data catalogs, the search engines for your company's data, have evolved as well. The Enterprise Data Catalog explores how AI is transforming enterprise-wide data search. In this second edition, you'll explore how the role of the data catalog has changed in the age of AI. Data catalogs no longer serve as tools to find and use data—they now deliver essential metadata for AI projects. Author Ole Olesen-Bagneux explains how metadata organized as enterprise ontologies, delivered through knowledge graphs, provides the context required by large language models, Model Context Protocol, and Agent2Agent Protocol. By drawing on data management and library and information science, the book shows why information science methodology is critical to successful catalog implementations. • Understand the role of the enterprise data catalog in discovery, governance, and AI • Organize data and sources using metadata, domains, and enterprise ontologies • Search and browse data across domains, lineage, and graphs • Leverage data catalog knowledge graphs for AI use cases • Apply data catalogs to data products and data contracts

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# The Enterprise Data Catalog: Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture ## 【One-Line Pitch】 A practical guide for data leaders and practitioners on how enterprise data catalogs have evolved from simple search tools into the metadata backbone that powers AI initiatives—covering domain architecture, knowledge graphs, and the emerging LLM + knowledge graph pattern. Read this if you're implementing or scaling a data catalog and want to understand its strategic role in AI-driven organizations. ## 【Book Arc】 - **Opening (~0%–9%)**: Establishes the core thesis—data catalogs are the "search engines for your company's data," just as web search engines transformed the internet. The preface recounts the author's journey from realizing this potential to now grappling with how ChatGPT and conversational AI fundamentally change what data catalogs can do. - **Early (~9%–25%)**: Introduces the foundational architecture of data catalogs: how they crawl IT landscapes, organize metadata into domains and subdomains, and create visibility across the entire organization at the metadata level (without exposing actual data values). Emphasizes that how you organize data determines how effectively you can search it. - **Early–Middle (~25%–38%)**: Explores the search functionality in depth, positioning data catalogs as both search engines and AI assistants. Introduces the concept of "ambient findability"—the target state where anyone can find any data asset from anywhere at any time—and distinguishes between searching *for* data (discovery) versus searching *in* data (data science). - **Middle (~38%–47%)**: Covers the human and organizational dimensions: the data discovery team (not just a "catalog team"), the ideal placement under a Chief Data Officer, and the three categories of end users (data analytics/AI, governance, and everyday users). Details specialized roles like asset stewards, term owners, and term stewards. - **Late (~47%–end)**: The book's trajectory points toward advanced topics covered in later chapters (not fully available in this excerpt): data products and data contracts (survivors of the data mesh movement), the LLM + knowledge graph pattern, and how data catalogs become a source in themselves for AI systems. ## 【Key Takeaways】 - **Data catalogs are now AI infrastructure, not just search tools** (Early): The role has shifted from helping humans find data to delivering essential metadata context for LLMs, MCP, and agentic architectures. This reframing helps secure executive buy-in for governance and compliance investments. - **Metadata-level visibility solves the governance paradox** (Early): Data catalogs let everyone see everything about company data—column names, descriptions, lineage—without exposing actual values. This makes comprehensive data overviews possible where raw data access would be technically and legally impossible. - **Domain architecture is the organizing principle** (Early): Data must "belong somewhere." You design domains and subdomains that vertically organize assets by business function (e.g., finance), making it possible to pinpoint exactly where any data originates. - **Search is the underutilized superpower** (Early): Most organizations treat search as a feature, not a strategy. Thinking of your data catalog as an enterprise search engine—and increasingly an AI assistant—unlocks its full value for data discovery. - **Ambient findability is the goal** (Middle): Borrowed from Peter Morville's work on web search, this concept means finding "anyone or anything from anywhere at anytime" within your enterprise data landscape. It's the standard data catalogs should strive toward. - **Data discovery teams, not catalog teams** (Middle): Implementation, maintenance, and promotion require dedicated cross-functional ownership. The ideal setup places this team under a CDO, where the executive data strategy is grounded in empirical data visibility. - **End users fall into three categories** (Middle): Data analytics/AI users (who deliver ROI through innovation), governance users (who ensure compliance), and everyday efficiency users. Each has different needs and should be supported accordingly. - **AI enables natural language interaction with catalogs** (Early): Enterprise AI assistants can now query both catalog metadata and other sources (like Slack) via MCP servers, enabling entirely new patterns of data discovery and contextual understanding. ## 【Reading Tips】 - **Skim the preface** (~0%–9%) if you're already convinced of the data catalog's strategic value; it's most useful for building the business case and understanding the author's intellectual journey. - **Deep-read Chapter 1** (~9%–47%) for the conceptual foundation: domain organization, metadata principles, search mechanics, and team structures. This is where the practical architecture guidance lives. - **Pay special attention to the "searching for data vs. searching in data" distinction** (~38%): This clarifies why data catalogs are discovery tools, not analytics platforms—a confusion that undermines many implementations. - **Note the organizational guidance** (~44%–47%): The discussion of where to place the data catalog team (CDO vs. analytics unit) involves real trade-offs between innovation and governance that you'll likely face. - **Watch for the AI integration patterns** throughout: The examples of enterprise AI assistants querying catalogs via MCP servers foreshadow the book's later focus on LLM + knowledge graph architectures. ## 【Coverage Limits】 This guide covers the available early-release content (Preface through Chapter 2). The excerpts do not cover later chapters on data products/contracts, the LLM + knowledge graph pattern, standards, or lifecycle management—these are listed in the table of contents but their content is unavailable in this sample. ##
Page 7
. . . . . . . . . . . . . . . . . . . . . . . vii Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 17
ny about the potential of a data catalog and it was crystal clear in my head, I was then faced with the battle of explaining the features to the important st...
View in text
Excerpt 3
w you to search that landscape. The main difference is that while a web search engine covers the web as a landscape, a data catalog covers your organization’...
View in text
Excerpt 4
expertise on a particular subset of assets (an entire data source or parts of data sources) in a domain. Term owner Term owners typically own a large subpart...
View in text
Excerpt 5
riven Value at Scale (Sebastopol, CA: O’Reilly, 2021). Pie‐ thein Strengholt, Data Management at Scale (Sebastopol, CA: O’Reilly, 2023, second edition), p. 3...
View in text
Excerpt 6
end to create hierarchies between terms, some being broader than others; for example, daily clothes is a narrower term than clothes but broader than daily cl...
View in text
Excerpt 7
d the DPO have correctly classified the data—it’s often the case that data is classified as highly confidential and not sensitive at all, at the same time. T...
View in text
Excerpt 8
L. The difference lies in what data layer the languages are applied on: you use DQLs to search in data, and you use IRQLs to search for data. To match DQL an...
View in text
Tags
AI categories
DataAITechnology
ISBN: 0642572284909
Publish Year: 2027
Language: English
Pages: 120
File Format: PDF
File Size: 2.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…