AI guide
# RAG-Ready Patterns for Data Platforms
## 【One-Line Pitch】
A practical blueprint for evolving traditional enterprise data platforms into AI-ready foundations that can support retrieval-augmented generation (RAG) systems—essential reading for data architects, platform engineers, and technical leaders facing pressure to deliver grounded generative AI on enterprise data.
## 【Book Arc】
- **Opening (~0%–9%)**: Defines the "enterprise context gap"—the disconnect between what legacy data platforms were built to deliver (dashboards, aggregates, pre-shaped questions) and what LLMs actually need (raw, contextual, retrievable knowledge). Introduces RAG as a fundamentally different consumption pattern where models, not humans, are the primary consumers.
- **Early (~9%–25%)**: Diagnoses why the "RAG readiness gap" exists—fragmentation, shallow semantics, governance built for reports, metadata as documentation, quality trapped in dashboards. Introduces the three-pillar framework (Data Assets, Infrastructure, Trust Layer) with diagnostic checklists and maturity models.
- **Early (~25%–34%)**: Deep dive into Pillar 1 (Data Assets)—the entity kernel concept (User, Tenant, Subscription, Contract, etc.), three-layer semantic architecture (source → conformed → product), and the four-level maturity progression from ad hoc entities to a full semantic fabric.
- **Middle (~34%–47%)**: Explores Pillar 2 (Infrastructure)—metadata-first pipeline engineering where contracts precede code, pipelines act as policy enforcers, and quality becomes a runtime service. Includes the "Yesterday's Numbers" case study showing how freshness must be embedded in retrieval, not dashboards.
- **Middle (~47%+)**: Covers Pillar 3 (Trust Layer)—scenario-based access control (SBAC) as an evolution beyond RBAC/ABAC, quality signals as retrieval inputs, and the IDEAS case study where PII incidents drove rapid SBAC adoption. Concludes with a self-assessment framework and transitions to the semantic backbone patterns in Chapter 2.
## 【Key Takeaways】
- **The enterprise context gap is the core problem** (Opening): LLMs don't consume dashboards or KPIs—they need raw, semantically rich, retrievable knowledge across structured and unstructured sources. Most platforms were built for descriptive analytics, not for the open-ended, cross-domain questions AI systems must answer.
- **RAG readiness is diagnosable and measurable** (Early): The three-pillar framework (Data Assets, Infrastructure, Trust Layer) with its checklist of ✅/⚠️/❌ indicators lets organizations honestly assess where they stand. Most will find gaps—that's the starting point, not a failure.
- **Semantics must start at the source** (Early): "Fixing it in the semantic layer" is too late for RAG. Core entity meanings (User, Tenant, Subscription, Partner) must be captured at acquisition and preserved through layered models, or everything downstream becomes guesswork.
- **Metadata must compile** (Early): Contracts, lineage, quality rules, and policy can't be documentation—they must be machine-readable, validated automatically, and enforced at creation time. Treat metadata as code and pipelines as the first line of governance.
- **Trust is a runtime signal, not a quarterly report** (Early): Quality isn't a dashboard artifact; it's a live signal retrieval must see. RAG systems should filter, rerank, and augment using freshness, completeness, and confidence—automatically, not through manual review.
- **Purpose beats permission** (Early): "Who can see what" is necessary but insufficient in the RAG era. Scenario-based access control (SBAC) explains why access is needed, under what scenario, and with which permissible joins—and it's surprisingly simple to implement alongside RBAC/ABAC.
- **The entity kernel is the foundation** (Early): Start with a small set of core business entities, make each explicit (identifiers, relationships, privacy tier, ownership), and bind your glossary to concrete attributes. Keep ontology "lite"—pragmatic, not academic.
- **Quality should be treated like an API** (Middle): The "Yesterday's Numbers" case study shows that freshness embedded in retrieval code (with enforcing thresholds) beats freshness displayed on dashboards. Stale queries get dropped automatically, restoring user confidence.
## 【Reading Tips】
- **Deep-read the opening chapters (~0%–25%)**: The three-pillar framework and the four hard-won lessons from Microsoft's IDEAS platform are the conceptual backbone. The diagnostic checklists in Tables 1-1 through 1-5 are worth returning to as self-assessment tools.
- **Skim the maturity models initially, return later**: The Level 1–4 progressions (for both data assets and trust) are useful for positioning your organization, but don't get bogged down in the details on first read—the patterns matter more than the taxonomy.
- **Pay special attention to the case studies**: "Yesterday's Numbers" (freshness as a retrieval signal) and the IDEAS PII incident (SBAC adoption) are concrete, memorable illustrations of abstract principles. They're the closest thing to worked examples in the available material.
- **Note what's unavailable**: Chapters 4–11 (data quality as retrieval signal, curated grounding, SBAC deep dive, RAG-ready data products, MCP orchestration, enterprise RAG stack, scorecard, and practices/pitfalls) are listed but not included in this excerpt sample. The full book will contain the detailed playbooks and artifacts (YAML/JSON contracts, glossary templates, policy specs) that this guide only previews.
- **Read with your own architecture in mind**: As you go through the three pillars, map each diagnostic question to your own systems. The value is in the self-assessment, not just the concepts.
## 【Coverage Limits】
This guide synthesizes only the available excerpt sample (~0%–53% of the book), covering Chapters 1–2 in depth. Chapters 3–11 (metadata-first pipeline engineering details, data quality as retrieval signal, curated grounding, SBAC implementation, RAG-ready data products, MCP orchestration, enterprise RAG stack design, readiness scorecard, and future practices) are listed in the table of contents but their content is not covered here.
##
Passage locations
Excerpt 1
ttil. All rights reserved. Published by O’Reilly Media, Inc. , 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educati...
View in text
Excerpt 2
thousands of internal teams and millions of external users. For years, our job was clear: create a single source of truth and run it like a mission-critical...
View in text
Excerpt 3
what your glossary editor and SME loop will evolve to meet. Case Study: Three Definitions of a “User” When we began grounding agents at Microsoft, one of the...
View in text
Excerpt 4
ibutes? Infrastructure Pipeline contracts as code? Enforced? Do you have automated lineage/data-quality monitors? Unified indexes: structured & unstructured?...
View in text