Digital Library

The Enterprise Data Catalog (for Raymond Rhine) (Ole Olesen-Bagneux) (z-library.sk, 1lib.sk, z-lib.sk)

Ole Olesen-Bagneux

The Enterprise Data Catalog (for Raymond Rhine) (Ole Olesen-Bagneux) (z-library.sk, 1lib.sk, z-lib.sk)

Author Ole Olesen-Bagneux

数据

It's a new day in search. Before ChatGPT, combing the web was simple—powerful search engines dominated for 25 years. That changed with conversational search powered by chatbots. Data catalogs, the search engines for your company's data, have evolved as well. The Enterprise Data Catalog explores how AI is transforming enterprise-wide data search. In this second edition, you'll explore how the role of the data catalog has changed in the age of AI. Data catalogs no longer serve as tools to find and use data—they now deliver essential metadata for AI projects. Author Ole Olesen-Bagneux explains how metadata organized as enterprise ontologies, delivered through knowledge graphs, provides the context required by large language models, Model Context Protocol, and Agent2Agent Protocol. By drawing on data management and library and information science, the book shows why information science methodology is critical to successful catalog implementations.

Format PDF
Size 4.6 MB
9
Views
0
Downloads
0.00
Total Donations

Text Preview (First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
The Enterprise Data Catalog SECOND EDITION Scale AI with Metadata Using LLMs, Model Context Protocol, and Agentic Architecture With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. Ole Olesen-Bagneux
Page 3
The Enterprise Data Catalog by Ole Olesen-Bagneux Copyright © 2027 O’Reilly Media, Inc. All rights reserved. Published by O’Reilly Media, Inc., 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (https://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Aaron Black Development Editor: Sara Hunter Production Editor: Katherine Tozer Interior Designer: David Futato Interior Illustrator: Kate Dullea February 2023: First Edition June 2027: Second Edition Revision History for the Early Release 2026-02-18: First Release See https://oreilly.com/catalog/errata.csp?isbn=9798341672949 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. The Enterprise Data Catalog, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc.
Page 4
The views expressed in this work are those of the author(s) and do not represent the publisher’s views. While the publisher and the author(s) have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the author(s) disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. 979-8-341-67290-1 [LSI]
Page 5
Brief Table of Contents (Not Yet Final) Preface (available) Chapter 1: Introduction to Data Catalogs (available) Chapter 2: Organize Data: Design a Robust Architecture for Search (unavailable) Chapter 3: Search Data: Concepts, Features, Mechanics, Patterns (unavailable) Chapter 4: Access and Observe Data (unavailable) Chapter 5 Empower End Users and Engage Stakeholders (unavailable) Chapter 6 Data Domains (unavailable) Chapter 7 Data Architecture, Providers and Consumers (unavailable) Chapter 8 Data Products and Data Contracts (unavailable) Chapter 9 Manage Data: Improve Lifecycle Management (unavailable) Chapter 10 The Data Catalog Is Now a Source in Itself (unavailable) Chapter 11 The LLM + KG Pattern (unavailable) Chapter 12 standards and AI (unavailable) Chapter 13 Looking Ahead, a Revisit (unavailable)
Page 6
Preface A NOTE FOR EARLY RELEASE READERS With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. This will be the Preface of the final book. If you’d like to be actively involved in reviewing and commenting on this draft, please reach out to the editor at shunter@oreilly.com. It’s A New Day in Search “So what’s your take on ChatGPT? How does it change data catalogs? Do you have an opinion about that?” Matt Housley was putting me on the spot. I was invited on the podcast Monday Morning Data Chat1 and we were discussing my soon to be published book The Enterprise Data Catalog, the first edition of this book that you are now reading in the second edition. I wasn’t properly mic’ed up, I was not yet used to being on tech podcasts and painfully aware of my clunky sound. I pulled myself together and managed to answer. I said that my book ended with a future vision for data catalogs, and that if I was to go into more depth about that future vision then, obviously I would look into the potential of ChatGPT and AI in general. It was a paradoxical moment, really. As I was publishing the first edition of The Enterprise Data Catalog, in early 2023, the core message of the book
Page 7
was that data catalogs are like search engines, just for data in companies. That’s why data catalogs need to be built on knowledge graphs, just like search engines are. TIP Don’t worry, we will get to knowledge graphs, they play a major role in this book! And yet, search engines themselves, the real ones, for the web, were unexpectedly challenged, right there in early 2023, as I published The Enterprise Data Catalog. For 25 years, the biggest business on the web - search - had been completely stable. Not technologically, of course it had evolved, but in terms of how you searched the web, as an end user. You had one, simple search bar, that would provide the best, the freshest, the most relevant search hits to you, with the blink of an eye. For 25 years, we had been using search engines as the most natural extension of our mind to search for everything from our absolute basic needs to the most complex types of curiosity the human brain can produce. And then, all of a sudden, it was over. In November 2022, ChatGPT 3.5 was released by OpenAI. Suddenly, we asked ourselves if the way we search was in fact obsolete. Would we talk with the web instead? Do conversational search through a chatbot? In those very early months of 2023 Microsoft CEO Satya Nadella said: “It’s a new day for search” when Microsoft joined forces with OpenAI. Nadella, Microsoft, was challenging Google: Microsoft was powering their search engine Bing with conversational search in a chatbot supported by OpenAI.2 The race for winning a disruption for search, as a business, was on - it was a spectacular turn of events after more than two decades of complete domination by Google. Nothing captured the zeitgeist better than the cover
Page 8
of the Economist 11th Feb 2023, as depicted in figure x.1. The battle for search was on. Figure P-1. The Economist nailed what AI did to search engines And here I was, arguing that data catalogs were like search engines, just for the data in your organization. And so, back to Matt’s question: How did this new technological push forward in AI change data catalogs? At the time I was asked the question, it was too early to provide an answer. No one knew. But something was clearly going to change. I saw things that I knew from my academic background would never fly, but I also saw interesting experiments.
Page 9
The first edition of my book was very well received, I was lucky to have thousands of readers all over the planet. The book was praised in reviews and featured in many podcasts, and I got to travel the world explaining my ideas at big conferences, events and for large corporations - the book even got translated. Its success was its message. Data catalogs are engines to search for data, so that you could later search in data, directly in databases. Static, preconceived metamodels in data catalogs led to catastrophic implementations, sky-rocketing costs and no ROI - because that is basically not how a search engine should work, not on the web, not in a company. As a PhD and an aspiring professor, I had taught knowledge organization, information retrieval and adjacent topics in my Library- and Information Science classes at the University of Copenhagen in Denmark, where I live. I went into industry, but I never forgot what I learned and taught. And it was on that basis, that I argue, that data catalogs, essentially, are like search engines. And I wrote The Enterprise Data Catalog from that perspective. And then came AI. All of a sudden there was a new dimension to data cataloging that I had not covered in my book. Why I Wrote This Book Therefore, it is time for a second edition of The Enterprise Data Catalog. All the insights from the first edition are kept - it all holds true still. The AI perspective has been added on top, and it falls into two parts: Part I: How AI augments data catalogs Part II: How data catalogs (ontologies) are a source for AI You will notice that these two parts are also the parts of this book - because we can implement and use data catalogs at scale with AI, and we can use the very structure of the data catalog, the metadata, as a source in itself, to fuel AI.
Page 10
Basically, because of AI, it’s also a new day in search for data catalogs. We are now beginning to understand what that day is about, and that is what we will unfold in the following pages. This book is about how AI augments the core message of my book, namely; how you organize data, defines how you can search it. And this message also has a new dimension, since the data catalog is in itself a source. Finally, the data mesh movement was at its peak when I published the first edition of The Enterprise Data Catalog. At the time of publication, it was difficult to say what would remain from the movement, after the buzz would fade. Now, that has become clear. No one talks about data mesh anymore, but two of the components in the data mesh complex stood the test of time: data products and data contracts. Therefore, this second edition covers data products and data contracts, because they are absolute key components in data catalogs. Who Should Read This Book This book is for everyone working in data, meaning Data engineering Data analysis Data science Data management Data governance These groups of employees all work with data and they all need to know where data is, who owns data, the quality of data, how data moves, where data ends up, and who uses it. The data engineers facilitate the storage, transformation, and movement of data, and they can monitor that in the data catalog. The data analysts and scientists use the data catalog to search for interesting data - and then, once found, they will search in those data
Page 11
sources. Data managers and data governors use the catalog as a strategic tool to ensure that data governance policies are successfully implemented. On top of that, a wealth of groups can benefit from using a data catalog, e.g. Architects Information security Data protection officers Architects can use a data catalog for effective data migration projects, information security and data protection to ensure an empirical validation of their anticipation of what data the organization has. Navigating This Book This book has two parts, Part I: The AI augmented Data Catalog, and Part II: Data Catalogs as a resource for AI. Part I: The AI Augmented Data Catalog is all about how data catalogs work and how they have improved since the first edition - mainly thanks to AI. The chapters introduce data catalogs, explains how data is organized in data catalogs and how search is performed, furthermore how data is accessed and observed. All these aspects are improved, made easier and faster, thanks to AI. Also, we will dive deeper into data domains, data architectures, and especially focus on data products and data contracts, as these matured significantly since the first edition. Part II: Data catalogs as a resource for AI uncovers a new perspective: Data catalogs are now becoming sources themselves, not only a catalog of sources. This new role for metadata is discussed in the light of the combination of knowledge graphs and large language models. Furthermore, the emerging standards of effective AI are explained, such as Model Context Protocol and Agent2Agent. This second ends with a future vision for data catalogs, and this is substantially updated since the first edition.
Page 12
Preface for the First Edition “This simply can’t be all there is to a data catalog. What does it really do?” About five years ago, I sat alone in the office among 20 empty desks. My company had shut off the air-conditioning to go green, so I was uncomfortably warm on top of being perplexed by the bunch of white papers, both printed and on my laptop, that were sitting in front of me. The papers explained a new technology called a data catalog. As an enterprise architect, I had been asked to implement a data catalog for our company. But first, I had to understand it. The papers I was looking at described cool, advanced features: column- based data lineage, graph visualizations of ontologies, and workflows to access virtualized data. Useful. Mesmerizing, really. But what was the overall point of a data catalog? I was sweating, physically and mentally, trying to draw upon my experiences to figure out the potential of this new technology. I have a BA, MA, and PhD in library and information science (LIS). I have taught LIS in university courses and been to conferences all over the world. I’ve seen a lot of things in this field, both good and bad. During my first job in pharma, senior management regularly called me late at night because inspections from the authorities were going haywire. The inspectors were asking them a multitude of questions: What was the temperature of this tube, in that machine, in June 1992? Where is the proof that the fermentation tank was cleaned according to the standard operating procedure (SOP)? When the data managers searched and couldn’t answer, they called my team—the Records and Information Management team. We employed our searching superpowers to find the information they needed. We were adept with our queries, used intuition and creativity to plan our moves, and drew on our knowledge to guide our search. We were able to do this because we had one guiding principle: How you organize data defines how you can search it. Because we knew how the data was organized, we knew ways to begin searching, modifying how we searched, broadening, changing focus, excluding hits, and finding the information we
Page 13
needed. Sometimes this was easy, and sometimes it was hard, but we would get there every time. This guiding principle has followed me throughout my career. I have cataloged furniture, weapons, human tissue, a lot of paper, and massive amounts of data. I know how to structure and operate a physical card catalog in a library. I know records and information management systems with both physical and digital storage. We both stored and cataloged data on premises, and then, later, in the cloud. Throughout everything I experienced, I saw that if you have a poorly organized data landscape, searching for the information you need will be a terrible experience. You will have to guess where to search and what to query. If your data is logically and systematically organized, however, you will know exactly where to look and what to query. It will be a much better experience. The idea that how you organize data defines how you search for it is reflected in our web habits as well. We never really think about how we search it anymore; it’s so intuitive. At work, within our company’s IT landscape, it can be a different story. We search in vain—companies hardly know their own data, let alone how it is processed. Data is undiscoverable and unmanageable. If only we had an enterprise search engine… On that hot summer day, alone in the empty office, surrounded by physical papers and dozens of open PDFs on my laptop, it suddenly hit me. “This data catalog has the potential to become a search engine for companies! We are finally getting an engine that will be able to do for companies what search engines have done for the web. The data catalog is a search engine!” That realization led to another a few years down the road. All of the papers I read that fateful day, along with all of the documentation that followed it, focused on explaining the complex features that are in data catalogs. They did not explain the data catalog itself and how it could revolutionize how we organize and search data. Nowhere has anyone talked about the future of the data catalog as an enterprise search engine. That realization has brought us to the book you are reading today.
Page 14
Although I had the epiphany about the potential of a data catalog and it was crystal clear in my head, I was then faced with the battle of explaining the features to the important stakeholders of my company. Although I knew they would see the benefits of this tool if they only took the time to understand it, they were simply not interested, nor did they have time to study it. I had to come up with a way to reach them. I went back to the idea of using the data catalog as an enterprise search engine. So, I asked myself, “What are people searching for? What would a data scientist be searching for? A data protection officer? A chief information security officer?” I decided to build demonstrations of the most vital data catalog features into small stories about specific stakeholders. Each slide deck had one central picture: a minimalist search bar with the company’s logo above it. I would explain the information need of a specific stakeholder, show the search in the search bar, reveal the result, then close with how the results could be used. In this way, I showed simple searches, complex searches, how to browse back and forth in the lineage of data, up and down in domains, and relationally in the graph that depicted our company. It had the same content as my previous demonstrations, but this time, it was explained from a stakeholder point of view: a specific person who was searching for something specific. And that worked. The stakeholders not only got interested, but they also got excited. They now wanted the data catalog, because they understood that this tool was not just a collection of fancy features for data geeks. No, this tool was something way more fundamental: the data catalog could help them search and find the data they were looking for. I explained that, implemented with care, a data catalog has relevance for many of the employees in a company. This approach worked for me and my colleagues, and I hope that it will work for you and yours as well. At the end of the day, we are all searching for something. And we search all the time. The only thing is, at work, it is very difficult to search for
Page 15
whatever we are trying to find. And we take that for granted, as something that we must just accept. I’m assuming you’re reading this book because you’re involved with planning to implement a data catalog, improve an existing one, sunset it, or simply trying to understand what kind of technology a data catalog is: what it does, how it should be used, and if it can help you in a certain way. You might be part of the offices of the legal counsel, chief data officer, data protection officer, or chief information officer. You might be a data engineer, data scientist, or data manager, or you might be part of the data governance team. If you are, then this book will help you understand what a data catalog is and how it will enable you to find exactly what you are searching for. However, you may also be a data catalog provider. In my book, I put forward a vision for the future of data catalogs, which you could benefit from when planning the future development of your data catalog technology. 1 Monday Morning Data Chat was a podcast by Joe Reis and Matt Housley, I was interviewed in the episode The Future of Data Catalogs, Matt asks the question 52:10 2 NPR, Feb 7th, 2023: Microsoft revamps Bing search engine to use artificial intelligence
Page 16
Chapter 1. Introduction to Data Catalogs A NOTE FOR EARLY RELEASE READERS With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. This will be the 1st chapter of the final book. If you’d like to be actively involved in reviewing and commenting on this draft, please reach out to the editor at shunter@oreilly.com. In this chapter, you’ll learn how a data catalog works, who uses them, and why. Before we dive into that, we will focus on why data catalogs have become relevant in a new way for Artificial Intelligence (AI). Pay close attention to this part, as it will ensure you strategic buy-in at the executive level, also for data governance and compliance, which is usually hard to get - as you are perhaps well aware. Then, we’ll go over the core functionalities of a data catalog and how it creates an overview of the data in your organization’s IT landscape. I’ll learn how the data can be organized in a data catalog, and how it makes searching for your data easy. Search is often underutilized and undervalued as part of a data catalog, which is a huge detriment to data catalogs. As such, we’ll talk about your data catalog as a search engine - and an AI assistant! - for your enterprise data that will unlock the potential for success. In this chapter, you’ll also learn about the benefits of a data catalog in an organization: a data catalog improves data discoverability, subsequently
Page 17
ensuring data governance and enhancing data-driven innovation. Moreover, you’ll learn about how to set up a data discovery team and you’ll learn who the users of your data catalog are. I’ll wrap up this chapter by explaining the roles and responsibilities in the data catalog. OK, off we go. The AI Data Catalog and as Source for AI Before we jump into the nuts and bolts of what a data catalog is, let’s take a moment and focus on why data catalogs have become relevant both with AI, and for AI. With AI. Data catalogs are now substantially easier to implement, use and scale, thanks to many of the features in them being supported by AI. This will become clear already in this chapter, in the search examples below, and in we will go deeper into this in the other chapters of Part I. For AI. As you will also see in this chapter, data catalogs are now part of a bigger search infrastructure, via AI assistants. AI assistants not only search on data catalogs but in all sources connected to the AI assistant. This has created a remarkable shift: The data catalog is not only a tool leading to sources, it has become a source in itself. NOTE You might know the term AI assistant under the term chatbot. Think Claude, Mistral, OpenAI and others, only for dedicated enterprise use. Throughout this book, we use AI assistants. Furthermore, data catalogs - if they are built on knowledge graphs - are a rich source for AI. Graphs have proved to increase precision greatly when applying Large Language Models (LLM) for generative AI projects.
Page 18
Furthermore, agentic architectures are also executing tasks more effectively when supplemented by a knowledge graph.1 Overall, the combination of AI and knowledge graphs is promising and we will discuss this in detail in Part II. You may think, If you are in data governance or data engineering: “Why does all this AI stuff matter to me?” I need my data catalog anyway. The answer is simple. It matters to you because this evolution has made it significantly easier to get your enterprise decision makers on board in implementing a data catalog, and you can also expect a smoother experience in rolling it out for new end users. The Core Functionality of a Data Catalog At its core, a data catalog is an organized inventory of the data in your company. That’s it. The data catalog provides an overview at a metadata level only, and thus no actual data values are exposed. This is the great advantage of a data catalog: you can let everyone see everything without fear of exposing confidential or sensitive data. In Figure 1-1, you can see a high-level description of a data catalog.
Page 19
Figure 1-1. High-level view of a data catalog A data catalog is basically a database with metadata that has been extracted from data sources in the IT landscape of a given company. The data catalog
Page 20
also has a search engine and an AI assistant inside it that allows you to search the metadata collected from the data sources. A data catalog will almost always have a lot more features, but Figure 1-1 illustrates the necessary core components. And in this book, I argue that the search capability is the single most important feature of data catalogs. In this section, we will discuss the three key features of the data catalog, namely that it creates an overview of the data in your IT landscape, it organizes your data, and it allows you to search your data. Let’s take a brief look at how data catalogs do this. NOTE With a data catalog, your entire organization is given the ability to see the data it has. Used correctly, that transparency can be very useful. For example, data scientists will no longer spend half their time searching for data, and they will have a much better overview of data that can really deliver value. Imagine the possibilities. They could be using their newfound time to analyze that data and discover insights that could lead the enterprise to developing better products! Create an Overview of the Data in the IT Landscape Creating an overview of the data in your IT landscape involves finding and displaying all the data sources in it, along with listing the people or roles attached to it. A data catalog can pull metadata with a connector that scans your IT landscape - each data source will have its own connector. Furthermore, as data product architectures are becoming more frequent and mature, data products can automatically publish themselves at the metadata layer, to the data catalog. We will discuss data products and data contracts in depth in Chapter 8. The IT landscape that is reflected in your data catalog will get business terminology added to it terms that are created in the data catalog (or inherited from other metadata repositories) and organized in glossaries. We will discuss glossary terms in Chapter 2 and how to search with them in
The above is a preview of the first 20 pages. Register to read the complete e-book.

Support Author

0.00
Total Amount (¥)
0
Donation Count

Recommended for You

Loading recommended books...
Failed to load, please try again later
Back to List