Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ole Olesen-Bagneux

Rating No ratings yet

Combing the web is simple, but how do you search for data at work? It's difficult and time-consuming, and can sometimes seem impossible. This book introduces a practical solution: the data catalog. Data analysts, data scientists, and data engineers will learn how to create true data discovery in their organizations, making the catalog a key enabler for data-driven innovation and data governance. Author Ole Olesen-Bagneux explains the benefits of implementing a data catalog. You'll learn how to organize data for your catalog, search for what you need, and manage data within the catalog. Written from a data management perspective and from a library and information science perspective, this book helps you: Learn what a data catalog is and how it can help your organization Organize data and its sources into domains and describe them with metadata Search data using very simple-to-complex search techniques and learn to browse in domains, data lineage,...

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A vendor-agnostic guide to enterprise data catalogs that treats them as search engines for organizational knowledge, not just metadata repositories. Best for data analysts, engineers, scientists, and governance leads who need to make data findable, understandable, and trustworthy at scale. 【Book Arc】 - **Opening (~0%–10%)**: Frames the core problem—searching for data at work is hard—and introduces the catalog as the practical solution, with endorsements positioning it as a foundational, vendor-neutral reference. - **Early (~10%–35%)**: Grounds the catalog in library and information science, using the ancient Library of Alexandria and Callimachus's Pinakes as an analogy for metadata, indexing, and the danger of losing institutional memory when teams turn over. - **Middle (~35%–55%)**: Walks through the book's structure: Part I covers organizing data into domains, building glossaries, and searching; Part II addresses data democratization, governance, and analytics use cases; Part III presents the vision of the catalog evolving into a company search engine. - **Late (~55%–85%)**: Develops the practical mechanics of domains, metadata description, and simple-to-complex search techniques, including browsing lineage and graph relationships. (Excerpts do not cover this range in detail.) - **Ending (~85%–100%)**: Argues that catalogs will shift from cataloging data to cataloging knowledge and "works," enabling true company memory. (Excerpts do not cover this range in detail.) 【Key Takeaways】 - **How you organize data defines how you can search it** (Early): The book's guiding principle, drawn from records management experience—organization and searchability are inseparable. - **The catalog is fundamentally a search engine, not a feature collection** (Early): Framing it around stakeholder information needs (data scientists, DPOs, CISOs) makes adoption succeed where feature demos fail. - **Metadata is the foundation of discovery** (Early): Callimachus's Pinakes shows that describing content without reading it is an ancient, still-valid solution to information overload. - **Governance failures stem from people and process, not technology** (Early): Data lakes became swamps through absent governance, while over-bureaucratic tools killed agility—the catalog must balance both. - **Domains and glossaries structure the catalog** (Middle): Organizing data sources into domains and describing them with metadata is the practical backbone of Part I. - **Data democratization depends on the catalog** (Middle): More employees can discover, access, and manage data independently, reducing reliance on a central team. - **The future is a company search engine** (Middle): The author envisions catalogs expanding from data to knowledge and "works," curing collective organizational amnesia. 【Reading Tips】 - Deep-read the early LIS chapters even if you're an engineer—the Alexandria analogy is the book's conceptual key, not decoration. - Skim the preface and endorsements quickly; the real substance starts with the organizing-and-searching chapters. - If you're evaluating vendors, focus on the vendor-agnostic framing and the stakeholder-search demonstrations rather than any specific tool. - Treat Part III as a vision essay, not an implementation manual—read it for direction, not step-by-step guidance. - Keep the guiding principle ("how you organize defines how you search") in mind as a checklist when designing your own catalog. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book (front matter, preface, and structural overview); detailed chapters on search techniques, lineage, governance workflows, and the full company-search-engine vision are not covered in the excerpts.
Excerpt 1
les data-driven organizations to reach their full potential. This book is a must-read for all IT professionals, data authorities, and data enthusiasts. Ann F...
View in text
Excerpt 2
or of Demetrius Phalereus, as the head of the Great Library. Demetrius, considered to be one of the greatest Greek thinkers, had been the creator and archite...
View in text
Excerpt 3
d cataloged data on premises, and then, later, in the cloud. Throughout everything I experienced, I saw that if you have a poorly organized data landscape, s...
View in text
Excerpt 4
e through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-dept...
View in text
Excerpt 5
suring data governance and enhancing data-driven innovation. Moreover, you’ll learn about how to set up a data discovery team and you’ll learn who the users...
View in text
Excerpt 6
ext relevant to the knowledge universe of your organization. This will make your asset more searchable. We’ll talk more about how to organize it in Chapter 2...
View in text
Excerpt 7
markably more effective with a data catalog than without it. Data discovery for data, in a data catalog, has a distinct target state: ambient findability . T...
View in text
Excerpt 8
atalog, based on conceptual metadata structures. Figure 1-9. Example of a metamodel in a data catalog Consider the metamodel in Figure 1-9 . In this hypothet...
View in text
Tags
AI categories
DataBig DataDatabase
ISBN: 149209871X
Publisher: O'Reilly Media
Publish Year: 2023
Language: English
Pages: 216
File Format: EPUB
File Size: 7.0 MB