Share E-Book

AuthorArindam Ganguly

Artificial Intelligence (AI) is the bedrock of today's applications, propelling the field towards Artificial General Intelligence (AGI). Despite this advancement, integrating such breakthroughs into large-scale production-grade enterprise applications presents significant challenges. This book addresses these hurdles in the domain of large language models within enterprise solutions. By leveraging Big Data engineering and popular data cataloguing tools, you'll see how to transform challenges into opportunities, emphasizing data reuse for multiple AI models across diverse domains. You'll gain insights into large language model behavior by using tools such as LangChain and LLamaIndex to segment vast datasets intelligently. Practical considerations take precedence, guiding you on effective AI Governance and data security, especially in data-sensitive industries like banking. This enterprise-focused book takes a pragmatic approach, ensuring large language models align with broader enterprise goals. From data gathering to deployment, it emphasizes the use of low code AI workflow tools for efficiency. Addressing the challenges of handling large volumes of data, the book provides insights into constructing robust Big Data pipelines tailored for Generative AI applications. Scaling Enterprise Solutions with Large Language Models will lead you through the Generative AI application lifecycle and provide the practical knowledge to deploy efficient Generative AI solutions for your business. What You Will Learn Examine the various phases of an AI Enterprise Applications implementation. Turn from AI engineer or Data Science to an Intelligent Enterprise Architect. Explore the seamless integration of AI in Big Data Pipelines. Manage pivotal elements surrounding model development, ensuring a comprehensive understanding of the complete application lifecycle. Plan and implement end-to-end large-scale enterprise AI applications with confidence. Who This Book Is For Enterprise Architects

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

Tags
AI categories
ai后端数据
No tags
ISBN: 8868811537
Publisher: Apress
Publish Year: 2025
Language: English
Pages: 458
File Format: PDF
File Size: 10.1 MB
Support Statistics
¥.00 · 0times
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

(This page has no text content)
Scaling Enterprise Solutions with Large Language Models Comprehensive End-to-End Generative AI Solutions for Production-Grade Enterprise Solutions Arindam Ganguly
Scaling Enterprise Solutions with Large Language Models: Comprehensive End-to-End Generative AI Solutions for Production-Grade Enterprise Solutions ISBN-13 (pbk): 979-8-8688-1153-1 ISBN-13 (electronic): 979-8-8688-1154-8 https://doi.org/10.1007/979-8-8688-1154-8 Copyright © 2025 by Arindam Ganguly This work is subject to copyright. All rights are reserved by the Publisher, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed. Trademarked names, logos, and images may appear in this book. Rather than use a trademark symbol with every occurrence of a trademarked name, logo, or image we use the names, logos, and images only in an editorial fashion and to the benefit of the trademark owner, with no intention of infringement of the trademark. The use in this publication of trade names, trademarks, service marks, and similar terms, even if they are not identified as such, is not to be taken as an expression of opinion as to whether or not they are subject to proprietary rights. While the advice and information in this book are believed to be true and accurate at the date of publication, neither the authors nor the editors nor the publisher can accept any legal responsibility for any errors or omissions that may be made. The publisher makes no warranty, express or implied, with respect to the material contained herein. Managing Director, Apress Media LLC: Welmoed Spahr Acquisitions Editor: Aditee Mirashi Desk Editor: James Markham Editorial Project Manager: Jacob Shmulewitz Copy Editor: Kezia Endsley Distributed to the book trade worldwide by Springer Science+Business Media New York, 1 New York Plaza, New York, NY 10004. Phone 1-800-SPRINGER, fax (201) 348-4505, e-mail orders-ny@springer-sbm.com, or visit www.springeronline.com. Apress Media, LLC is a Delaware LLC and the sole member (owner) is Springer Science + Business Media Finance Inc (SSBM Finance Inc). SSBM Finance Inc is a Delaware corporation. For information on translations, please e-mail booktranslations@springernature.com; for reprint, paperback, or audio rights, please e-mail bookpermissions@springernature.com. Apress titles may be purchased in bulk for academic, corporate, or promotional use. eBook versions and licenses are also available for most titles. For more information, reference our Print and eBook Bulk Sales web page at http://www.apress.com/bulk-sales. Any source code or other supplementary material referenced by the author in this book is available to readers on GitHub. For more detailed information, please visit https://www.apress.com/gp/ services/source-code. If disposing of this product, please recycle the paper Arindam Ganguly Howrah, West Bengal, India
Dedicated to my better half, Sugandha Ghosh, and to my mother
(This page has no text content)
(This page has no text content)
(This page has no text content)
(This page has no text content)
(This page has no text content)
(This page has no text content)
(This page has no text content)
xiii Arindam Ganguly is an experienced data scientist at a leading multi-national software service firm, where he is responsible for developing and designing intelligent solutions by leveraging his expertise in AI and data analytics. He has over nine years of experience delivering enterprise products and applications and has proven skillsets in developing and managing a number of software products with various technical stacks. Arindam also is well-versed in developing automation and hyper- automation solutions that leverage automated workflow engines and integrating them with AI. Additionally, he is the author of Build and Deploy Machine Learning Solutions Using IBM Watson, which teaches readers how to build AI applications using the popular IBM Watson toolkit. About the Author
xv Varunsaagar Saravanan is an experienced AI/ML engineer specializing in Generative AI, Large Language Models, and NLP. With over six years of experience, he has contributed to groundbreaking AI solutions across the media, entertainment, and e-governance sectors. Recognized for award-winning innovations like the ePaarvai AI Cataract Application and his leadership in AI-driven local news broadcasting, Varunsaagar has published research and mentored teams in AI strategy. His work integrates cutting-edge AI techniques into impactful, scalable products, driving transformation and innovation. About the Technical Reviewer
xvii This book would not be possible without the constant support of my partner, Sugandha Ghosh. There were many times when she pushed me to complete the book when I thought it was not within me to do so. A big hug and a thanks to my mother, Moli Ganguly, for always putting up with me when I’m too busy working and writing and not being able to give her the time she deserves. It would not be just if I didn't mention my childhood pal, Dipanjan Bosu, for being on my side when I needed motivation. A lot of this book is derived from public documentation available on various frameworks, including LangChain and LangGraph. The Scikit- Learn and OpenAI documentation are worth mentioning here, as most of the book relies on them. Last, but not the least, a big thanks to the team at Apress, especially Aditee and Shobana, for helping me during the project. Acknowledgments
xix Introduction It will not be long before the world adapts AI into daily life for even the simple things. The information industry will be overwhelmed with the demands of consumers. Many AI enthusiasts and self-proclaimed experts have good knowledge of certain parts of AI, but they usually fail when attempting to put all these concepts together. I had been a long-term enthusiast when AI was still a buzzword and the information age was gearing up for Big Data. I have seen my colleagues and seniors struggle to manage large datasets, let alone analyze and play around with them. After Big Data was sorted with technologies such as Hadoop and Spark, and machine learning was established using Scikit-Learn, anticipation ran wild with the kinds of opportunities it could bring to the market. Developers soon realized that, although these AI techniques are good tools for solving smaller isolated problems, there was no way to marry them together to create a large enterprise application. Soon interest dropped, and AI became dormant in areas dealing with large scale applications. I have seen some of the biggest players throw away machine learning just because it couldn't address these real-world datasets. There are two types of people working with AI—the scientists who give their heart and soul in bettering the algorithms and the practitioners who try to use these algorithms in the real world. Although it was going good for the scientists, the constant setbacks of the practitioners led the AI industry to a halt—that is, until the introduction of the Transformers architecture. The world quickly saw the inception of Generative AI and LLMs. Suddenly, there was an overwhelming demand for embedding AI into these applications. But again, practitioners realized the struggles of infusing Generative AI into large-scale enterprise applications.
xx During some very tough months, the AI practitioner community came up with techniques to use the best of ML and Generative AI in real-world, large-scale applications and data. This is when I realized the necessity to author a book to spread the word that MLOps, Gen AI, and data engineering can carefully coexist and create wonderful additions to non- intelligent large-scale enterprise applications. This book follows a structured approach, where a practitioner can relate to the struggle and an enthusiast who has perfected their AI skills is introduced to the real-world struggle of putting their skills into place for large applications and datasets. This book introduces AI and Generative AI concepts and explains ways to infuse them into applications to run them in production. I hope this book will be a revelation into the tricks and techniques needed to set you apart in the world of AI. InTroduCTIon
1© Arindam Ganguly 2025 A. Ganguly, Scaling Enterprise Solutions with Large Language Models, https://doi.org/10.1007/979-8-8688-1154-8_1 CHAPTER 1 Machine Learning Primer The world has gone through a lot of revolutions since the dawn of time. For example, during the stone age, humans invented powerful tools for basic survival (such as the wheel). With the advent of city life, humans started embedding structures into all forms of work, and this lead to the industrial revolution. One of the biggest revolutions taking place now is called the artificial intelligence (AI) revolution. Although it may seem intuitive from a 30,000-foot level, the current advancements are products of multiple waves of AI inventions, starting from the first Turing machine concept. To give you a better understanding of machine learning (ML), this chapter provides an overview of the subject. The Origins of Machine Learning According to Wikipedia, “The term machine learning was coined in 1959 by Arthur Samuel, an IBM employee and pioneer in the field of computer gaming and artificial intelligence. The synonym self-teaching computers was also used in this time period.” To understand this, consider the difference between traditional problem solving and machine learning (see Figure 1-1).
2 Problem solving has been an inherent human skill since the beginning of time. Problem-solving methods were put into structures in the form of algorithms. With the advent of computers, algorithms could be programmed into computers and solve complex problems flawlessly. These algorithms can be structured into computers as procedural or functional programming. Procedural programming takes each step one by one and uses iterations and conditions in a monolithic structure. On the other hand, functional programming uses the concept of mathematical functions to break a problem down into smaller problems and arrive at a final solution. With the advancement of time, more problem-solving skills have come to existence, but the basic theory remains the same. Given a set of data as the input and a set of rules, the outcome is a set of outputs. These problem-solving methods work best when there is a predefined set of rules or a known set or path to follow. When the ruleset is not known but the desired output is known, the input and output is used to create a set of rules. In other words, the machine tries to learn a rule (or pattern), in order to produce the known output as accurately as possible. Machine learning is an application of statistics. In other words, it is the study of statistical inference. Hence, all the machine learning algorithms are derived from various statistical inference techniques. The following sections look at some of the popular machine learning algorithms. Chapter 1 MaChine Learning priMer
3 Figure 1-1. Traditional computer programming vs machine learning Linear Regression Linear regression (see Figure 1-2) is one of the most popular and intuitive machine learning algorithms. It assumes that the dataset is linearly shaped and hence tries to use a linear line equation on the dataset. Y = WX+b Chapter 1 MaChine Learning priMer
4 Figure 1-2. Linear regression The task of the machine learning algorithm is to figure out W and b, given X is the set of inputs and Y is the set of outputs. Since there is no fixed rule, the algorithm iterates with some initial values of W and b, producing new values of Y (say Y’) and X (say X’) until the differences between Y and Y’ and X and X’ are so small that they can be ignored. Although this might seem easy, the detailed math behind it is very complex. But developers need not worry so much about the mathematical intricacies, because Python many packages to abstract all the complexities behind single lines of code. Python is the most popular language choice for data scientists and machine learning engineers. The most popular package for traditional machine learning algorithms is Scikit-Learn (see https://scikit- learn.org). Scikit-Learn can help you develop a simple linear regression with just four lines of code: Chapter 1 MaChine Learning priMer