Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Indrajit Kar

Rating No ratings yet

Generative AI and agentic AI are reshaping how we interact with data, enabling intelligent systems that can reason, generate, and autonomously act across multiple modalities. From text and images to voice and structured data, these technologies are increasingly essential in enterprise and research applications today. - This book offers a complete roadmap to mastering multimodal generative AI and agentic AI systems. It covers foundational concepts, vision-language models, retrieval-augmented generation, human-in-the-loop and multi-agent workflows, text-to-SQL, OCR, and hybrid AI integrations. Each chapter combines theory, practical guidance, code implementations, and real-world case studies, helping readers understand architectures, pipelines, and production-grade deployments. - By the end of this book, readers will be capable of designing, implementing, and scaling robust multimodal and agentic AI systems. They will gain hands-on expertise in reasoning, generation, retrieval, agent orchestration, and Ops, equipping them to build production-ready AI applications and excel in their roles.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on roadmap for engineers and technical leaders who want to move from generative AI fundamentals to production-grade multimodal and agentic systems, spanning retrieval, vision-language models, voice, text-to-SQL, and operations. It is best suited to practitioners who learn by building rather than by reading theory alone. 【Book Arc】 - **Opening (~0%–15%)**: Frames the shift toward multimodal and agentic AI and lays out the book's 18-chapter, concept-to-code promise, positioning retrieval, generation, and orchestration as the core building blocks. - **Early (~15%–32%)**: Introduces foundational concepts—RAG, tokens, vector databases, reranking, bi- vs. cross-encoders, guardrails, agents, and Model Context Protocols—then moves into vision-language models and multimodal system classification. - **Early–Middle (~32%–47%)**: Turns to implementation: local GPU and Ollama setups, OpenAI API-based systems, human-in-the-loop and multi-agent RAG workflows, and two/multi-stage retrieval with dense retrieval interactions and reranking. - **Middle (~47%–60%)**: Builds bidirectional and multimodal RAG systems, adds reranking with cross-encoders, and covers retrieval optimization techniques such as modality routing, query expansion, hybrid retrieval, and adaptive index refresh. - **Late (~60%–85%)**: Extends into voice-enabled RAG, advanced reasoning and prompting, text-to-SQL systems (including agentic variants), OCR and multimodal document handling, and hybrid integration of traditional ML (e.g., XGBoost) into GenAI workflows. - **Ending (~85%–100%)**: Closes on operations and evaluation—LLM/RAG evaluation methods, RagOps, monitoring and observability, graph-enhanced RAG, and experiment management with MLflow. 【Key Takeaways】 - **Retrieval, generation, and orchestration are the three pillars** (Early): the book treats RAG, vector stores, reranking, and orchestration as the foundation everything else builds on. - **Multimodal means more than text and images** (Early–Middle): vision-language models, voice input, and structured data are all treated as first-class modalities with distinct architectures. - **Agentic systems differ from simple AI agents** (Middle): the book draws an explicit distinction and shows how human-in-the-loop workflows evolve into multi-agent RAG systems. - **Retrieval quality is an engineering problem, not a given** (Middle): optimization techniques like modality routing, query expansion, hybrid retrieval, and adaptive index refresh are presented as practical fixes to common retrieval drawbacks. - **Reasoning and prompting are levers for reliability** (Late): advanced multimodal systems are framed around reasoning types, benchmarks, and prompting strategies that improve dependability. - **Text-to-SQL is a hard, multi-faceted problem** (Late): the book covers entity extraction, architecture design, and a full set of accuracy, latency, and human-evaluation metrics. - **Hybrid AI is a legitimate production pattern** (Late): traditional ML models such as XGBoost can be wrapped inside LLM-powered systems, illustrated through a telecom fraud-detection case study. - **Ops and evaluation decide whether systems survive** (Ending): RagOps, continuous monitoring, observability, and MLflow-based experiment management are treated as essential, not optional. 【Reading Tips】 - **Skim the opening chapters if you already know RAG basics**, but read the vision-language and multimodal classification material carefully—it sets vocabulary used throughout. - **Deep-read the implementation chapters (local GPU, API-based, HITL, multi-stage RAG)** with the code bundle open; the excerpts indicate these are step-by-step and exercise-driven. - **Treat the "to do" sections as required practice**, not optional—they appear repeatedly and are clearly designed to reinforce each chapter's concepts. - **Pay extra attention to retrieval optimization and evaluation chapters** if you are targeting production; these are where the book shifts from building to sustaining systems. - **Use the case studies (fraud detection, retail text-to-SQL, recommendation systems)** as templates for mapping the techniques onto your own domain. 【Coverage Limits】 This guide is based on stratified excerpts covering the book's front matter, chapter outlines, and selected sections; detailed code, full case-study results, and the complete text of later chapters are not fully represented, so specific implementation details may differ from this summary.
Excerpt 1
lications, India ISBN: 978-93-65898-385 All Rights Reserved. No part of this publication may be reproduced, distributed or transmitted in any form or by any...
View in text
Excerpt 2
s on vision-language models and their role in multimodal AI. It explains what vision-language models are, compares different implementation approaches, and e...
View in text
Excerpt 3
able at https://github.com/bpbpublications . Check them out! Errata We take immense pride in our work at BPB Publications and follow best practices to ensure...
View in text
Excerpt 4
M-as-a-judge Rationale and functionality To do Conclusion 9. Building GenAI Systems with Reranking Introduction Structure Objectives Reranking Reranking in i...
View in text
Excerpt 5
ML model integration in GenAI workflows To do Conclusion 18. LLM Operations and GenAI Evaluation Techniques Introduction Structure Objectives Importance of O...
View in text
Excerpt 6
isticated, learning-driven, and memory-augmented techniques. Prior to understanding modern retrieval systems, it is helpful to trace their evolution briefly,...
View in text
Excerpt 7
ure shows the types of LLMs and generation models: Figure 1.3 : Types of LLMs and generation models Types of generation systems GenAI systems span multiple m...
View in text
Excerpt 8
to improve the quality before passing them to the generator. Reduces hallucination by focusing the generation only on the most relevant documents. The follow...
View in text
Tags
AI categories
Artificial IntelligenceGenerative AIMultimodal AI
ISBN: 9365898382
Publisher: BPB Publications
Publish Year: 2026
Language: English
File Format: EPUB
File Size: 6.8 MB