Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Shivam R Solanki, Drupad K Khublani

Rating No ratings yet

This book explains the field of Generative Artificial Intelligence (AI), focusing on its potential and applications, and aims to provide you with an understanding of the underlying principles, techniques, and practical use cases of Generative AI models. The book begins with an introduction to the foundations of Generative AI, including an overview of the field, its evolution, and its significance in today’s AI landscape. It focuses on generative visual models, exploring the exciting field of transforming text into images and videos. A chapter covering text-to-video generation provides insights into synthesizing videos from textual descriptions, opening up new possibilities for creative content generation. A chapter covers generative audio models and prompt-to-audio synthesis using Text-to-Speech (TTS) techniques. Then the book switch gears to dive into generative text models, exploring the concepts of Large Language Models (LLMs), natural language generation (NLG), fine-tuning, prompt tuning, and reinforcement learning. The book explores techniques for fixing LLMs and making them grounded and indestructible, along with practical applications in enterprise-grade applications such as question answering, summarization, and knowledge-based generation. By the end of this book, you will understand Generative text, and audio and visual models, and have the knowledge and tools necessary to harness the creative and transformative capabilities of Generative AI. What You Will Learn What is Generative Artificial Intelligence? What are text-to-image synthesis techniques and conditional image generation? What is prompt-to-audio synthesis using Text-to-Speech (TTS) techniques? What are text-to-video models and how do you tune them? What are large language models, and how do you tune them? Who This Book Is For Those with intermediate to advanced technical knowledge in artifi

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, code-light tour of Generative AI—from diffusion models and GANs to text-to-image/audio/video pipelines and large language models—ideal for intermediate-to-advanced developers and data scientists who want a structured map of the field before diving into implementation. 【Book Arc】 - **Opening (~0%–7%)**: Introduces the book’s mission—explaining Generative AI’s principles, evolution, and real-world significance—and frames the field as a shift from static, preprogrammed responses to models that learn patterns and create new content. Includes author dedications and a nod to Alan Turing’s foundational role. - **Early (~14%–33%)**: Lays the technical groundwork: neural networks as the backbone, the generative-vs-discriminative distinction, and core model families (diffusion models, GANs, VAEs, Restricted Boltzmann Machines, Pixel RNNs). Also covers the historical milestones—Hinton’s deep belief nets (2006), Goodfellow’s GANs (2014), the transformer (2017), GPT’s rise (2018+), and mainstream breakthroughs like DALL-E 2 and ChatGPT (2022). - **Middle (~38%–52%)**: Moves into generative visual models, covering text-to-image synthesis and conditional image generation, then transitions to text-to-video generation—how to synthesize video from textual descriptions and how to tune these models. The section closes by pivoting to generative audio, with prompt-to-audio synthesis via Text-to-Speech (TTS). - **Middle (~52%–62%)**: Dives into Large Language Models (LLMs): training and adoption phases, transformer variants (encoder, decoder-only, encoder-decoder), and fine-tuning BERT as a concrete example. Ends with a forward-looking glimpse at the LLM horizon. - **Late (~67%–86%)**: Focuses on generative text in practice: NLP tasks (sentiment analysis, entity extraction, topic modeling), NLG tasks (creative writing, summarization, dialogue), advanced prompting (few-shot, chain-of-thought), and the prompting-vs-fine-tuning trade-off. Includes case studies on fine-tuning for sentiment analysis and question answering, plus parameter-efficient fine-tuning. - **Ending (~88%–100%)**: Returns to the big picture—Generative AI’s versatility across industries (design, music, education, science)—and reinforces key neural network architectures (CNNs for images, RNNs for sequences) and the generative-vs-discriminative distinction, closing with a summary table. 【Key Takeaways】 - **Generative AI is about creation, not just analysis** (Opening): Unlike discriminative models that classify or predict, generative models learn underlying patterns to produce new, original content—text, images, audio, video. This reframing matters because it shifts expectations from “tools that respond” to “tools that invent.” (Early) - **A handful of model families dominate the field** (Early): Diffusion models, GANs, VAEs, Restricted Boltzmann Machines, and Pixel RNNs each bring different strengths—realistic image generation, unsupervised learning, sequential data handling. Knowing which family fits your task is the first design decision. (Early) - **The transformer architecture is the backbone of modern LLMs** (Middle): Introduced in 2017 via “Attention Is All You Need,” self-attention mechanisms enabled more efficient, coherent text generation than earlier RNN-based approaches. This is the conceptual bridge from older NLP to today’s GPT-style models. (Middle) - **LLMs come in three architectural flavors** (Middle): Encoder models (like BERT) excel at understanding tasks, decoder-only models (like GPT) generate text, and encoder-decoder models handle sequence-to-sequence tasks. Choosing the right architecture depends on whether your goal is comprehension, generation, or translation-style tasks. (Middle) - **Prompting and fine-tuning are complementary, not competing** (Late): Few-shot prompting and chain-of-thought techniques can extract strong performance without retraining, while fine-tuning (including parameter-efficient methods) adapts models to specific domains or tasks. The trade-off is cost and control—prompting is cheap and fast, fine-tuning is deeper but heavier. (Late) - **Practical LLM applications span understanding and generation** (Late): From sentiment analysis and entity extraction to summarization, dialogue, and creative writing, the same underlying model can serve very different enterprise needs. Case studies on fine-tuning for sentiment analysis and question answering show how to operationalize these tasks. (Late) - **CNNs and RNNs remain relevant for visual and sequential generation** (Ending): CNNs handle image data with spatial hierarchies (e.g., DeepDream), while RNNs process sequences for text and music. Even in the transformer era, these architectures underpin many generative pipelines, especially for non-text modalities. (Ending) 【Reading Tips】 - **Skim the early history if you’re already familiar with AI basics** (~14%–33%): The milestones (Hinton, Goodfellow, Vaswani) are useful context, but the real value is in the model-family comparisons—focus there if you’re short on time. - **Deep-read the LLM chapters** (~52%–67%): The transformer variants and fine-tuning case studies are the book’s most actionable content. Pay special attention to the prompting-vs-fine-tuning discussion and the parameter-efficient fine-tuning section—these are directly applicable to real projects. - **Treat the visual and audio chapters as overviews** (~38%–52%): Text-to-image, text-to-video, and TTS are covered conceptually rather than with deep implementation detail. Skim for terminology and use cases, not for code-level mastery. - **Watch for the generative-vs-discriminative distinction** (Early and Ending): It recurs throughout the book and is the single most important conceptual anchor. If you grasp this early, the rest of the material slots into place more easily. - **Use the case studies as templates** (Late): The sentiment analysis and question-answering fine-tuning examples are practical blueprints. Even if you don’t replicate them, they show the workflow—data prep, model selection, tuning strategy—that you’ll need for your own tasks. 【Coverage Limits】 The excerpts provide strong coverage of the book’s structure, model taxonomy, and LLM-focused content, but do not include detailed code samples, specific implementation walkthroughs, or the full text of the visual/audio chapters. Practical tuning steps for text-to-video and TTS models are referenced but not elaborated in the available material.
Excerpt 1
What are text-to-video models and how do you tune them? What are large language models, and how do you tune them? Who This Book Is For Those with intermediat...
View in text
Page 4
ved the way for artificial intelligence. The authors
View in text
Page 5
22 Table of Contents
View in text
Page 6
111 Table of ConTenTs
View in text
Page 7
231 Table of ConTenTs
View in text
Page 8
348 Table of ConTenTs
View in text
Page 12
techniques, and practical use cases of Generative AI models. The book begins with an introduction to the foundations of Generative AI, including an overview...
View in text
Page 20
n understanding the differences between different languages. It could then identify which language a text belongs to but couldn’t compose its own creative te...
View in text
Tags
AI categories
Artificial IntelligenceAIProgramming Language
ISBN: 8868804026
Publisher: Apress
Publish Year: 2024
Language: English
Pages: 458
File Format: PDF
File Size: 4.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…