Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ole Olesen-Bagneux

Rating No ratings yet

It's a new day in search. Before ChatGPT, combing the web was simple—powerful search engines dominated for 25 years. That changed with conversational search powered by chatbots. Data catalogs, the search engines for your company's data, have evolved as well. The Enterprise Data Catalog explores how AI is transforming enterprise-wide data search. In this second edition, you'll explore how the role of the data catalog has changed in the age of AI. Data catalogs no longer serve as tools to find and use data—they now deliver essential metadata for AI projects. Author Ole Olesen-Bagneux explains how metadata organized as enterprise ontologies, delivered through knowledge graphs, provides the context required by large language models, Model Context Protocol, and Agent2Agent Protocol. By drawing on data management and library and information science, the book shows why information science methodology is critical to successful catalog implementations.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical guide to how enterprise data catalogs have evolved from simple data-finding tools into essential metadata infrastructure for AI, showing data professionals how to organize, search, and govern data in the age of LLMs and agentic architectures. Read this if you work in data engineering, analysis, science, management, or governance—and need to understand why catalogs now matter for AI, not just for humans. 【Book Arc】 - **Opening (~0%–10%)**: Sets up the central disruption—ChatGPT and conversational search upended 25 years of stable web search, and the same shift is now hitting enterprise data catalogs. The problem: how does AI change what a catalog is and does? - **Early (~10%–30%)**: Introduces the core thesis that how you organize data defines how you can search it, drawing on the author's experience as an enterprise architect and his realization that a data catalog is essentially a search engine for a company's data. - **Early–Middle (~30%–40%)**: Explains the dual role of AI: catalogs are now *augmented by* AI (easier to implement, use, and scale) and *serve as a source for* AI (metadata feeding AI assistants and LLMs). Introduces the shift from catalog-as-tool to catalog-as-source. - **Middle (~40%–50%)**: Covers the core mechanics—what a data catalog actually is (an organized metadata inventory), how data is organized into domains, assets, and data sources, and why metadata visibility solves the data silo problem without exposing sensitive data. - **Late (~50%–70%)**: Moves into advanced territory: data products and data contracts as matured components from the data mesh movement, plus the role of knowledge graphs and enterprise ontologies in providing context for LLMs. - **Ending (~70%–100%)**: Explores the emerging standards and architecture for AI-ready catalogs—Model Context Protocol, Agent2Agent Protocol, the LLM + knowledge graph pattern—and closes with a future vision for data catalogs as sources in themselves. (Note: later chapters are marked unavailable in the excerpts.) 【Key Takeaways】 - **The catalog's role has fundamentally shifted from finder to feeder** (Early–Middle): Data catalogs no longer just help humans locate data—they now supply essential metadata to AI projects, making them a source in themselves rather than merely a pointer to sources. - **How you organize data defines how you can search it** (Early): This is the book's foundational principle, drawn from library and information science, and it applies equally to web search engines and enterprise data catalogs. - **Metadata visibility solves the data silo problem without exposing data** (Middle): Everyone can see everything at the metadata level, which creates organizational awareness of data assets while keeping actual data values hidden and secure. - **Knowledge graphs are the bridge between catalogs and LLMs** (Late): Organizing metadata as enterprise ontologies delivered through knowledge graphs provides the context that large language models, MCP, and A2A protocols require to work effectively. - **Data products and data contracts survived the data mesh hype cycle** (Early): These two components proved durable and are now key elements of enterprise data catalogs, covered in this second edition because they matter for real implementations. - **AI makes catalog implementation and adoption easier, not just more complex** (Early–Middle): AI-supported features reduce the friction of implementing, using, and scaling catalogs, which helps win executive buy-in for data governance and compliance initiatives. - **Different stakeholders search for different things** (Early): The author's breakthrough was framing catalog value through specific stakeholder stories—data scientists, data protection officers, CISOs—each with distinct information needs, rather than listing features. - **Information science methodology is critical to successful catalog implementations** (Opening): The book draws on library and information science to argue that cataloging principles, not just technology, determine whether a catalog actually works. 【Reading Tips】 - **Deep-read Part I if you're implementing or evaluating a catalog**: It covers the core mechanics—domains, assets, search, roles, and permissions—and explains why search is the most underutilized and undervalued catalog capability. - **Deep-read Part II if you're working on AI projects**: The knowledge graph + LLM pattern, MCP, and agentic architecture discussions are where the book's most forward-looking and strategically valuable content lives. - **Skim the preface and first-edition retrospective if you already know the basics**: The opening chapters spend time on the author's personal journey and the first edition's context, which is useful for framing but not essential for practitioners. - **Pay attention to the stakeholder framing**: The author's approach of explaining catalog value through specific user stories (data scientist, DPO, CISO) is a practical communication tool you can reuse in your own organization. - **Note what's marked unavailable**: Later chapters on the LLM + KG pattern, standards and AI, and the future vision are not covered in the excerpts—seek out the full book for those sections. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book, with later chapters (10–13) marked unavailable. The synthesis of Part II's advanced AI architecture content is therefore limited to what the preface and early chapters preview.
Page 3
ment Editor: Sara Hunter Production Editor: Katherine Tozer Interior Designer: David Futato Interior Illustrator: Kate Dullea February 2023: First Edition Ju...
View in text
Page 10
of AI, it’s also a new day in search for data catalogs. We are now beginning to understand what that day is about, and that is what we will unfold in the fol...
View in text
Page 14
are all searching for something. And we search all the time. The only thing is, at work, it is very difficult to search for whatever we are trying to find. A...
View in text
Page 20
ther metadata repositories) and organized in glossaries. We will discuss glossary terms in Chapter 2 and how to search with them in Chapter 3. Besides glossa...
View in text
Excerpt 5
tem and, ideally, how the data is transformed as it travels. Data lineage can be subdivided into many different layers, in Figure 1-3 below, you see the laye...
View in text
Excerpt 6
purpose of searching for data, for example, a data catalog.4 The difference between searching for data and searching in data may strike you as not very impor...
View in text
Excerpt 7
ance end users primarily search the data catalog for either confidential data or sensitive data—or both—in order to protect that data. They do so both as the...
View in text
Excerpt 8
ion to Modern Information Retrieval (New York: Neal-Schuman Publishers, 2010), chaps. 1 and 2. 5 Peter Morville, Ambient Findability: What We Find Changes Wh...
View in text
Tags
AI categories
DataArtificial IntelligenceData Governance
Publish Year: 2026
Language: English
File Format: PDF
File Size: 4.6 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…