Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Page
1
EXPERT INSIGHT The Architecture Handbook for Milvus Vector Database Design and implement high-performance vector search systems with Milvus Yudong Cai | Jeremy Zhu | Xuan Yang | Bang Fu Yudong C ai | Jerem y Zhu X uan Yang | Bang Fu The A rchitecture H andbook for M ilvus Vector D atabase www.packtpub.com The Architecture Handbook for Milvus Vector Database Things you will learn: • Deploy Milvus using Docker, Kubernetes, and Helm • Confi gure Milvus and monitor system health with Prometheus, Grafana, and Loki • Understand core components like Knowhere, indexes, time sync, compaction, and garbage collection • Design and optimize schema, queries, and data modifi cation fl ows • Benchmark performance and simulate real-world failure recovery • Scale Milvus clusters to support large datasets and high-concurrency traff ic • Apply security hardening, rate-limiting, and role-based access control • Build AI applications using Milvus with LangChain Download a free PDF copy of this book https://packtpub.com/unlock/9781835881705 The rapid adoption of LLMs demands eff icient storage and lightning-fast retrieval of unstructured data. Designed as a vector database, Milvus has earned widespread recognition in the community and support from tech giants like Apple and NVIDIA. Yet, many developers only scratch the surface of what Milvus is truly capable of. Writt en by the contributors of the Milvus project, this handbook gives you an insider’s perspective on its design and how it handles large-scale, high-dimensional vector data. Starting with the basics, you’ll learn about everything from service deployment and SDK usage to Milvus’ layered architecture and how its components interact. You’ll learn how the indexing, replication, compaction, and garbage collection systems work and how to apply them to real scenarios. Through practical demos and confi guration exercises, you’ll learn how to monitor, scale, and secure Milvus in production and then advance to performance evaluation and scalability testing using tools like VectorDBBench. You’ll also explore Milvus’ integration with LangChain for use cases such as vector search and RAG-based chatbots. By the end of this book, you’ll be able to analyze Milvus internals, fi ne-tune for performance, ensure system stability, and integrate it into next-generation AI solutions.
Page
2
The Architecture Handbook for Milvus Vector Database Design and implement high-performance vector search systems with Milvus Yudong Cai Jeremy Zhu Xuan Yang Bang Fu
Page
3
The Architecture Handbook for Milvus Vector Database Copyright © 2026 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the authors, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. This book was written by Yudong Cai, Xuan Yang, Jeremy Zhu, and Bang Fu. Generative AI tools were used only to assist with ideation, phrasing, and diagram drafts, and all technical content and code were created, verified, and tested by the author and Packt's editorial team. Packt does not accept AI-generated content that replaces expert authorship. Portfolio Director: Gebin George Relationship Lead: Gebin George/Sonia Chauhan Project Manager: Prajakta Naik Content Engineer: Aditi Chatterjee Technical Editor: Rahul Limbachiya Indexer: Manju Arasan Proofreader: Aditi Chatterjee Production Designer: Ponraj Dhandapani Growth Lead: Nimisha Dua First published: March 2026 Production reference: 1270326 Published by Packt Publishing Ltd. Grosvenor House 11 St Paul's Square Birmingham B3 1RB, UK. ISBN 978-1-83588-170-5 www.packtpub.com
Page
4
I would like to dedicate this book to my wife, Chai Lei, and my daughter, Cai Chaifei, who are the most important people in my life. They have brought me endless joy and strength, and without their support, this book would not have been possible. I also wish to dedicate this work to Starlord Xie, CEO of Zilliz, who gave me the opportunity to join Zilliz and collaborate with the team to build the remarkable Milvus project. – Cai Yudong I would like to dedicate this book to my husband, Tailang Li, whose unwavering support and encouragement made this journey possible. His patience and belief in me gave me the strength to see this work through to completion. – Xuan Yang
Page
5
(This page has no text content)
Page
6
Contributors About the authors Yudong Cai is a senior software engineer with over 20 years of experience in large-scale system development. As one of the founding members of the Milvus project, he helped build Milvus from the ground up and has been involved in the development and iteration of every version since its initial open-source release. His key contributions include delivering the first production-grade Range Search implementation, as well as the refactoring of the entire Milvus configuration system, alongside the design and implementation of numerous other critical features. He is also the original developer and key maintainer of Knowhere, Milvus' core vector computation engine, where he designed its architecture to support multiple hardware acceleration frameworks and a wide range of vector search algorithms. Jeremy Zhu is a quality assurance engineer at Zilliz, focused on ensuring the robustness and high performance of the Milvus vector database. His core responsibilities include designing comprehensive test cases, developing automated system test pipelines for diverse scenarios, and executing rigorous stress, recovery, and performance testing. Jeremy possesses deep expertise in chaos engineering, distributed systems testing, and test automation frameworks, playing a key role in maintaining Milvus' high-quality standards. Xuan Yang is a senior software engineer at Zilliz in China, passionate about designing high- performance, scalable distributed database systems. As a core Milvus contributor, she architected the DataNode module, implemented the compaction process, and led the L0 segment design. She is the primary maintainer of PyMilvus, the official Python SDK, and VectorDBBench, an open- source benchmarking framework for vector databases. She cares deeply about system stability and performance and is always eager to collaborate with the community to push the boundaries of large-scale AI and vector data infrastructure.
Page
7
Bang Fu is a senior software engineer at Zilliz. With extensive experience in both Go and Python, he has actively contributed to the development of several key features for Milvus, including permission verification, request interception, incremental synchronization, and serverless metering functionalities. He is also interested in AI technology and led the development of the GPTCache project, which focuses on caching LLM responses to improve speed and reduce costs. In addition, he has participated in the development of the DeepSearcher project.
Page
8
About the reviewers Anwar Hermuche is a senior AI engineer and founder of DascIA, with six years of experience in Data and AI solutions. He is dedicated to the mission of democratizing AI education, believing in technology as a robust, ethical, and accessible production infrastructure to transform the global ecosystem. Andrew Zhu is a software engineer and AI enthusiast with deep expertise in vector databases, machine learning systems, and distributed computing. With over a decade of experience in technology development, Andrew brings practical insights from both industry applications and research implementations. Currently based in the Seattle area, Andrew works extensively with modern AI infrastructure, including Milvus for scalable similarity search. His background includes hands-on experience building retrieval-augmented generation (RAG) systems, implementing approximate nearest neighbor (ANN) algorithms, and optimizing vector database performance at scale. When he's not working with vector databases, Andrew contributes to open-source projects, experiments with LLM applications, and helps mentor others entering the AI engineering field. Priyanka Mohekar is an AI engineer specializing in generative AI, RAG, and data engineering. She builds scalable, enterprise-grade AI systems using LLMs, agentic frameworks, and modern data platforms. With experience in designing end-to-end data pipelines and intelligent automation solutions, she is passionate about solving real-world problems through AI. Priyanka also actively shares her knowledge through mentoring and content creation, aiming to make advanced AI concepts more accessible and impactful. Hitesh Chaudhari is an AI/ML engineer specializing in generative AI, large-scale vector search, and distributed AI systems. He has deep expertise in Milvus, including cluster deployment, indexing strategies, and performance optimization for high-scale RAG pipelines. His work spans multi-agent LLM systems, hybrid search, embeddings, and production-grade AI architectures. He focuses on building scalable, low-latency, and fault-tolerant AI systems and provides enterprise AI consulting to design and deploy cutting-edge AI platforms.
Page
9
Subscribe for a free eBook New frameworks, evolving architectures, research drops, production breakdowns—AI_Distilled filters the noise into a weekly briefing for engineers and researchers working hands-on with LLMs and GenAI systems. Subscribe now and receive a free eBook, along with weekly insights that help you stay focused and informed. Subscribe at https://packt.link/8Oz6Y or scan the QR code below.
Page
10
Table of Contents Preface xxi Free benefits with your book ........................................................................... xxvii Part 1: Getting Started with Milvus 1 Chapter 1: Introduction to Milvus 3 Technical requirements ........................................................................................ 4 Understanding a vector database .......................................................................... 5 Types of databases • 5 Unstructured data • 6 Exploring the encoding process • 6 One-hot encoding • 7 TF-IDF encoding • 8 Similarity search algorithms • 10 LSH • 11 Introducing Milvus .............................................................................................. 15 Summary ............................................................................................................ 17 Chapter 2: Deploying Milvus in Multiple Ways 19 Technical requirements ....................................................................................... 19 Understanding common service deployment methods ........................................ 20 Introducing Docker and Docker Compose • 20 Introducing Docker Compose • 21 Introducing the Kubernetes operator • 23 Introducing Helm charts • 24 Comparing Helm charts and Kubernetes operators • 25 Deploying Milvus ............................................................................................... 26 Deploying Milvus with Docker Compose • 26
Page
11
Deploying Milvus through Kubernetes • 28 Using operators to deploy Milvus • 29 Using Helm charts to deploy Milvus • 33 Summary ........................................................................................................... 37 Chapter 3: Interacting with Milvus 39 Technical requirements ...................................................................................... 39 Introducing the core objects of Milvus ................................................................ 39 Fields • 40 Schemas • 43 Collection • 44 Index • 45 Interacting with Milvus via the Python SDK ........................................................ 45 Step 1: Connecting to Milvus • 46 Step 2: Creating a collection • 46 Step 3: Viewing collections • 47 Step 4: Inserting data • 49 Step 5: Querying the data with an ID of 2 • 50 Step 6: Upserting the data with an ID of 2 • 50 Step 7: Running a query with filtering • 51 Step 8: Deleting an entity • 51 Step 9: Carrying out vector search • 52 Step 10: Deleting the collection • 53 Interacting with Milvus via REST APIs ................................................................ 53 Step 1: Creating a collection • 54 Step 2: Describing the collection • 54 Step 3: Listing all collections • 55 Step 4: Checking the index information and load status of the collection • 55 Step 5: Carrying out a data query or search • 56 Step 6: Dropping the collection • 56 Summary ........................................................................................................... 57 Table of Contents x
Page
12
Chapter 4: Configuring the Milvus System 59 Technical requirements ...................................................................................... 59 Understanding dependency services and startup configurations for Milvus ......... 60 Managing core system and service configurations in Milvus ................................ 63 Adjusting configurations before starting Milvus • 65 Modifying configuration ..................................................................................... 68 Monitoring access and system logs • 70 Deploying Prometheus and Grafana • 71 Deploying Loki and Promtail • 76 Summary ........................................................................................................... 79 Part 2: System Architecture and Data Lifecycle Management 81 Chapter 5: Understanding the Milvus Data Model and Architecture 83 Technical requirements ...................................................................................... 83 Milvus data model .............................................................................................. 84 Collection, entity, and schema • 84 Entity composition • 88 Partition • 93 Partition APIs • 93 Partition key field • 94 Shards • 95 Typical number of shards • 98 Milvus' layered architecture ............................................................................... 98 Access layer • 100 Coordination layer • 102 Worker layer • 103 Storage layer • 105 MQ – the lifeline of streaming data ..................................................................... 107 Global event ordering • 109 Timestamp in Milvus • 110 xi Table of Contents
Page
13
Timetick mechanism • 111 Summary .......................................................................................................... 112 Chapter 6: Data Modification and Maintenance in Milvus 115 Technical requirements ..................................................................................... 115 The core of data organization – the segment ....................................................... 116 Data states – streaming vs historical • 116 Internal structure – the file composition of a segment • 116 The lifecycle of a segment ................................................................................. 120 Phase one – the growing state • 120 Phase two – the sealed state • 121 Phase three – the flushed state • 121 Phase four – the dropped state • 122 The journey behind DML requests ..................................................................... 123 Insert request • 123 Synchronous processing (during the API call) • 127 Asynchronous processing (after the API returns) • 128 Visibility and persistence points • 130 Delete request • 130 Synchronous processing (during the API call) • 131 Asynchronous processing (after the API returns) • 133 Upsert request • 134 Channel checkpoints ......................................................................................... 136 Summary .......................................................................................................... 138 Chapter 7: Reading Data in Milvus 139 The query scatter-gather pattern ...................................................................... 140 Query participants • 141 The scatter phase – distributing the workload • 142 The gather phase • 142 Loading and balancing ...................................................................................... 143 Replicas • 144 Table of Contents xii
Page
14
Loading and releasing replicas • 146 Dynamic load balancing • 147 Handling deletes ............................................................................................... 151 The delegator's role and Bloom filters • 152 Ensuring multi-level consistency ....................................................................... 153 How Milvus ensures consistency • 155 Summary .......................................................................................................... 158 Chapter 8: Compaction and Garbage Collection 159 Technical requirements .................................................................................... 160 Why compaction matters in Milvus? ................................................................. 160 Segment fragmentation • 160 How fragmentation impacts query performance • 161 Mix compaction's performance optimization • 161 Clustering compaction enables intra-partition segment pruning • 163 Storage and performance overhead of soft deletions • 164 L0 and mix compaction address soft delete overhead • 164 Compaction in detail ......................................................................................... 165 The common process • 165 Taking a closer look at the scheduling process • 166 Mix compaction • 167 L0 compaction • 169 Clustering compaction • 171 Triggers • 172 Managing obsolete data with garbage collection ................................................ 172 GC working principles • 173 Safety first – multi-layer protection ensuring data integrity • 173 Performance priority – intelligent scheduling and query protection • 174 Fine-tuning GC behavior • 175 Summary .......................................................................................................... 176 177 xiii Table of Contents
Page
15
Part 3: Vector Computing Engine, Indexing, and Advanced Query Mechanisms Chapter 9: Exploring Milvus' Vector Engine 179 Technical requirements ..................................................................................... 179 Introducing Milvus' core engine: Knowhere ...................................................... 180 Understanding the functional scope of Knowhere • 180 Milvus data model • 183 Knowhere architecture • 185 Working with Knowhere's main APIs ................................................................ 186 Knowhere parameters • 187 Train(), Add(), and Build() • 188 Serialize() and Deserialize() • 189 Search() and RangeSearch() • 190 HasRawData() and GetVectorByIds() • 191 Understanding Knowhere's functionality .......................................................... 191 Thread management • 192 Version management • 193 Prometheus monitor • 194 Summary .......................................................................................................... 194 Chapter 10: How to Select a Vector Index 197 Introducing performance metrics ..................................................................... 198 Recall and accuracy • 198 Vector per second and latency • 199 Understanding vector types and metric types .................................................... 200 FloatVector series • 201 BinaryVector • 203 SparseFloatVector • 205 Theory behind indexes • 207 Measuring the performance of indexes .............................................................. 208 FLAT index and BruteForce series • 210 Table of Contents xiv
Page
16
IVF series indexes • 212 SCANN • 216 HNSW series indexes • 217 DiskANN • 220 RAFT series indexes • 221 SPARSE series indexes • 226 Recommendations for index selection • 226 Summary .......................................................................................................... 227 Chapter 11: Handling Complicated Search Requests 229 Handling attribute filtering .............................................................................. 230 Boolean or logical expression • 232 AST • 235 Plan AST generation • 236 AST execution • 238 Using bitsets to soft delete vectors ..................................................................... 238 What's a bitset? • 238 How to calculate a bitset • 239 Using SIMD to accelerate distance calculations ................................................. 243 SIMD APIs' dynamic hooking • 244 Connecting the dots – an overview of Milvus query execution ............................ 246 Summary ......................................................................................................... 247 Part 4: System Evaluation, Tuning, and Ecosystem Application 249 Chapter 12: Getting Started with Milvus Performance Benchmarking 251 Technical requirements ..................................................................................... 251 Understanding ANN search and Milvus performance metrics ............................. 252 What is ANN search? • 252 How ANN search works in Milvus • 253 Why Milvus performance matters • 253 xv Table of Contents
Page
17
Common performance challenges in vector search • 255 Measuring Milvus performance • 256 Factors influencing Milvus' performance • 261 Benchmarking tools overview ........................................................................... 266 Standardized benchmarks vs. real-world workloads • 266 Common benchmarking tools • 268 Running your benchmark with VectorDBBench and Locust ............................... 270 Using Locust for custom performance testing • 277 Summary .......................................................................................................... 281 Chapter 13: Stability and Reliability Evaluation for the Milvus Vector Database 283 Technical requirements .................................................................................... 284 Understanding Milvus stability and reliability ................................................... 285 Understanding the relationship between stability and reliability • 287 Setting up the Milvus reliability testing environment ........................................ 287 Configuring Milvus for high-availability testing • 287 Key Grafana monitoring panels for reliability testing • 290 How to interpret Grafana metrics for cluster health • 292 Load testing for Milvus reliability ..................................................................... 294 Designing load scenarios • 294 Example scenario: e-commerce product recommendation system • 294 Executing load tests • 298 Step 1: preparing test data • 298 Step 2: writing load scripts with Locust • 298 Step 3: writing your own custom load shape • 300 Analyzing the results and deriving action items • 300 Long-term running (soak) test • 302 Methods and goals • 302 E-commerce soak test scenario • 303 Complex workload testing simulating real-world scenarios • 304 Simulating Milvus component failures and verifying recovery ........................... 305 Table of Contents xvi
Page
18
Prerequisites: installing Chaos Mesh • 305 Simulating active MixCoord failure • 306 Test steps • 308 Verification points • 310 Simulating worker node failure • 310 Understanding the Milvus WAL mechanism • 311 Test steps • 311 Testing WAL data durability • 312 Verification points for QueryNode failure • 313 Example test output • 313 Summary .......................................................................................................... 314 Chapter 14: Scalability Evaluation for the Milvus Vector Database 317 Technical requirements ..................................................................................... 318 Scaling strategies in Milvus ............................................................................... 318 Vertical scaling (scale-up) • 319 Horizontal scaling (scale-out) • 320 Worker nodes and their scaling benefits • 320 Making the right scaling decision • 321 Introduction to scalability testing methodology ................................................ 322 Define clear test objectives • 322 Prepare test environment and data • 323 Dataset requirements • 323 Design basic test workloads • 324 Focusing on key performance indicators • 325 Utilizing basic testing and monitoring tools • 327 Hands-on lab: vertical scaling with testing methodologies ................................ 328 Hands-on lab: horizontal scaling (performance improvement type) .................... 332 Scenario – single in-memory replica test (verify basic horizontal scaling effects) • 335 Scenario – multiple in-memory replicas test • 336 Hands-on lab: capacity improvement testing with methodology ........................ 338 xvii Table of Contents
Page
19
Scenario – verifying the cluster's ability to linearly handle larger datasets after scaling • 338 Using scaling to resolve Milvus performance bottlenecks .................................. 342 Monitoring scaling effectiveness • 345 Summary ......................................................................................................... 346 Chapter 15: Getting Started with Milvus Performance Tuning 349 Technical requirements .................................................................................... 350 Understanding performance trade-offs ............................................................. 350 Goal 1: making searches faster • 350 Search speed vs. accuracy • 350 Search speed vs. memory usage • 351 Goal 2: speeding up data ingestion • 352 Write throughput vs. search performance • 352 Write throughput vs. memory usage • 352 Write throughput vs. search freshness • 352 Goal 3: maximizing load capacity • 353 Capacity vs. search accuracy • 353 Capacity vs. search speed • 353 Making the right trade-off decisions • 353 When search speed is your primary goal • 354 When data ingestion is your primary goal • 354 When load capacity is your primary goal • 355 Making your searches faster .............................................................................. 355 The vector search execution path • 356 Optimization technique 1: choosing and tuning the right index • 357 Optimization technique 2: leveraging batch search processing • 360 Optimization technique 3: scaling with multiple replicas • 364 Combining techniques for maximum performance • 365 Speeding up your data ingestion ....................................................................... 366 Understanding data ingestion performance challenges • 366 Optimization technique 1: batch insertion for immediate impact • 367 Table of Contents xviii
Page
20
Optimization technique 2: multi-shard parallelization • 370 Optimization technique 3: bulk import for massive datasets • 373 Combining ingestion techniques for maximum throughput • 375 Maximizing queryable data capacity ................................................................. 376 Understanding queryable capacity bottlenecks • 376 Optimization technique 1: IVF_SQ8 quantization for substantial capacity expansion • 377 Optimization technique 2: mmap for intelligent hot/cold data management • 380 Optimization technique 3: DiskANN for breaking the billion-vector barrier • 386 Combining capacity techniques for maximum efficiency • 390 Summary ......................................................................................................... 390 Chapter 16: Implementing Multi-Tenancy in Milvus 393 Technical requirements .................................................................................... 394 How Milvus enables multi-tenancy ................................................................... 394 Available multi-tenancy strategies • 395 Strategy 1: Database-level isolation • 395 Strategy 2: Collection-level isolation • 399 Strategy 3: Partition-based multi-tenancy • 404 Strategy 4: Resource group-based performance isolation • 411 Multi-tenancy strategy comparison ................................................................... 414 Choosing the right multi-tenancy strategy • 416 Summary .......................................................................................................... 417 Chapter 17: How Milvus Works in AI 419 Technical requirements .................................................................................... 420 Identifying Milvus' role in AI applications ......................................................... 420 Milvus with unstructured data • 421 Milvus with RAG • 422 Milvus with multi-modal AI • 424 Image search with Milvus ................................................................................. 425 Building a chatbot with Milvus and LlamaIndex ............................................... 432 xix Table of Contents
The above is a preview of the first 20 pages. Register to read the complete e-book.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
【One-Line Pitch】
A contributor-written, insider tour of how Milvus actually works—from Docker deployment and SDK basics to the internals of indexing, compaction, and query execution—for engineers who need to run vector search in production rather than just call an API. Best suited to backend, ML platform, and infrastructure engineers building RAG or large-scale similarity search systems.
【Book Arc】
- **Opening (~0%–10%)**: Frames the problem—LLMs need fast storage and retrieval of unstructured data—and sets up the toolchain (Google Colab, VS Code, Docker, GitHub repo) plus a first look at vector-search fundamentals such as LSH.
- **Early (~10%–32%)**: Hands-on deployment and interaction: Docker, Docker Compose, Kubernetes/Minikube, and the Milvus Operator; then the Python SDK workflow (connect, create schema/collection, insert, query, filter, delete, search) and configuration via milvus.yaml, including monitoring with Prometheus and Grafana.
- **Middle (~32%–50%)**: The architecture core—Milvus' layered design, etcd's role in metadata and service discovery, the log broker as WAL (RocksMQ standalone, Pulsar/Kafka cluster), object storage (MinIO/S3/GCS) enabling compute-storage separation, and the data model of entities, fields, and partitions.
- **Late (~50%–80%)**: Data lifecycle and query internals: DML request journeys with synchronous/asynchronous phases, segment lifecycle, delete semantics, the scatter-gather query pattern, loading and balancing, and multi-level consistency.
- **Ending (~80%–100%)**: Production operations and resilience—compaction and garbage collection, benchmarking with VectorDBBench, failure simulation with Chaos Mesh (MixCoord, worker nodes, WAL durability), vertical vs. horizontal scaling, security hardening, rate limiting, RBAC, and LangChain integration for RAG.
【Key Takeaways】
- **Milvus is a four-layer, disaggregated system** (Middle): etcd handles metadata and service discovery, the log broker acts as WAL, and object storage is the source of truth—this separation is what makes stateless workers and elastic scaling possible.
- **The log broker is the reliability backbone** (Middle): every insert, upsert, or delete must be persisted to the WAL before the proxy returns success, so acknowledged DML is never lost.
- **Deployment choice drives the log broker** (Early): standalone uses embedded RocksMQ; cluster defaults to Pulsar, with Kafka supported for teams already invested in that stack.
- **DML is two-phase** (Late): the synchronous phase resolves primary keys and writes to the message queue; actual persistence happens asynchronously in the background—understanding this explains delete/insert visibility behavior.
- **Queries follow a scatter-gather pattern** (Late): data loaded into memory is queried across distributed nodes and results are merged, which is central to how Milvus scales reads.
- **Consistency is multi-level, not binary** (Late): the book treats consistency as a tunable spectrum tied to loading and balancing behavior.
- **Operations are first-class** (Ending): monitoring (Prometheus/Grafana/Loki), benchmarking (VectorDBBench), and chaos testing (Chaos Mesh) are presented as normal production practice, not afterthoughts.
- **Security and integration close the loop** (Ending): rate limiting, RBAC, and LangChain-based RAG show how Milvus fits into real AI applications.
【Reading Tips】
- **Deep-read the Middle and Late sections** on architecture, WAL, and query flow—these are the "insider" chapters that distinguish this book from SDK documentation.
- **Skim the Early deployment walkthroughs** if you already use Kubernetes; treat them as reference commands rather than narrative.
- **Run the code**: the book leans on a GitHub repo and Colab, and the DML/query internals land better after you've executed the SDK examples yourself.
- **Hard spots**: the log broker/WAL and timestamp model reward slow reading; revisit them after the DML chapter for a second pass.
- **Take away a mental model**, not memorized configs—know which component owns which responsibility, and the configs become self-explanatory.
【Coverage Limits】
This guide is synthesized from stratified excerpts covering roughly the first half of the book plus table-of-contents and preface material; later chapters (compaction, GC, chaos testing, scaling, security) are represented mainly by headings and summaries, so specifics there are not detailed.
Passage locations
Excerpt 1
lvus’ integration • Scale Milvus clusters to support large with LangChain for use cases such as vector search and datasets and high-concurrency traff ic RAG-...
View in text
Excerpt 2
download images, making sharing and management easier. Many companies also set up their own private Docker image repositories to provide faster upload and do...
View in text
Excerpt 3
=/etc/prometheus/prometheus.yml' depends_on: - "standalone" grafana: container_name: milvus-grafana image: grafana/grafana ports: - "3002:3000" environment:...
View in text
Excerpt 4
DML requests are executed. The journey behind DML requests After understanding 8 9 5 1 2 e 4 1 _ xtd i he static structure and dynamic lifecycle of segments,...
View in text
Recommended for You
{{#thumbnailUrl}}
{{/thumbnailUrl}}
{{^thumbnailUrl}}
{{/thumbnailUrl}}
Loading recommended books...
Failed to load, please try again later
Tip the Site
Scan the WeChat Pay or Alipay code to tip. No login required.
WeChat Pay
Alipay