Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Omar Sanseviero, Pedro Cuenca, Apolinário Passos, Jonathan Whitaker

Rating No ratings yet

Learn to use generative AI techniques to create novel text, images, audio, and even music with this practical, hands-on book. Readers will understand how state-of-the-art generative models work, how to fine-tune and adapt them to their needs, and how to combine existing building blocks to create new models and creative applications in different domains. This go-to book introduces theoretical concepts followed by guided practical applications, with extensive code samples and easy-to-understand illustrations. You'll learn how to use open source libraries to utilize transformers and diffusion models, conduct code exploration, and study several existing projects to help guide your work. Build and customize models that can generate text and images Explore trade-offs between using a pretrained model and fine-tuning your own model Create and utilize models that can generate, edit, and modify images in any style Customize transformers and diffusion models for multiple creative purposes Train models that can reflect your own unique style

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, code-first guide for developers and AI enthusiasts who want to move beyond using generative AI as a black box—teaching you how transformer and diffusion models actually work, and how to fine-tune, combine, and customize them to generate text, images, audio, and music in your own style. 【Book Arc】 - **Opening (~0%–15%)**: Introduces the core paradigms of generative AI—transformers for sequential data (text, audio) and diffusion models for visual data (images)—and sets up the open-source tooling (Hugging Face libraries) used throughout. This stage solves the "what am I actually building with?" problem by framing generative models as composable building blocks rather than monolithic magic. - **Early (~15%–35%)**: Dives into transformer architecture fundamentals—attention mechanisms, tokenization, and pretraining objectives—with guided code exploration. The focus is on understanding how text generation works under the hood, including decoding strategies (greedy, beam, sampling) and their trade-offs for creativity versus coherence. - **Middle (~35%–60%)**: Shifts to diffusion models for image generation, explaining the forward noising and reverse denoising process. Practical sections cover using pretrained models (e.g., Stable Diffusion) and the key trade-off between leveraging a pretrained model versus fine-tuning on your own dataset—including when each approach makes sense for your project. - **Late (~60%–85%)**: Moves into customization and control: fine-tuning transformers for specific text styles or tasks, and adapting diffusion models for image editing, inpainting, style transfer, and personalized outputs. This stage emphasizes "training models that reflect your own unique style" through hands-on examples. - **Ending (~85%–100%)**: Combines everything into creative applications—building pipelines that mix text and image generation, exploring audio and music generation with transformers, and studying existing open-source projects as templates. The book closes with guidance on deploying and scaling your generative models in real-world settings. 【Key Takeaways】 - **Generative AI is a building-block discipline, not a single model** (Opening): Transformers excel at sequential generation (text, audio, music), while diffusion models dominate image synthesis; the real power comes from combining them. Expect to assemble pipelines rather than train everything from scratch. - **Understanding decoding strategies is the difference between coherent and chaotic text** (Early): Greedy decoding is safe but repetitive; sampling with temperature and top-p introduces creativity but risks incoherence. This is the first practical lever you pull to control output quality. - **Pretrained models are your default starting point** (Middle): The book's central practical advice is to start with a pretrained model (e.g., Stable Diffusion) and only fine-tune when you have a specific style, domain, or constraint that the base model can't handle. Fine-tuning costs data, compute, and expertise—so justify it. - **Diffusion models work by learning to reverse noise** (Middle): The forward process adds noise to images, and the model learns to denoise step-by-step. Grasping this core mechanism helps you understand why diffusion models are powerful for image editing and style transfer, not just generation. - **Fine-tuning is about steering, not reinventing** (Late): Whether adapting a text model to your writing style or an image model to your aesthetic, the process involves curated datasets and careful hyperparameter choices. The book provides concrete recipes for doing this with open-source libraries. - **Image editing and generation share the same underlying machinery** (Late): Inpainting, outpainting, and style modification are achieved by conditioning the diffusion process—not by building new architectures. This insight lets you repurpose one model for many creative tasks. - **Cross-domain pipelines unlock novel applications** (Ending): The final projects show how to chain text generation (prompts, captions) with image generation (diffusion) and even audio—demonstrating that the real creative payoff comes from orchestration, not any single model. 【Reading Tips】 - **Skim the theory, but don't skip the code**: Each chapter pairs concepts with runnable examples. If you're short on time, run the code first, then read the explanation to solidify what you observed. - **Deep-read the fine-tuning chapters (Middle to Late)**: This is where the book earns its "hands-on" title. Pay special attention to dataset preparation and the trade-off discussions—they'll save you from expensive mistakes. - **Treat the decoding strategy section (Early) as a reference**: You'll return to it whenever generated text feels off. Bookmark the temperature and top-p explanations. - **Expect a learning curve with diffusion math**: The noising/denoising explanation can feel abstract. Don't get stuck—focus on the intuition (adding noise, then learning to remove it) and let the code demos carry you through. - **Use the final projects as templates**: The last chapters are less about new theory and more about showing you complete, working systems. Steal their structure for your own creative applications. 【Coverage Limits】 The excerpts provided cover the book's overall scope and stated learning objectives but do not include detailed chapter-by-chapter content, specific code samples, or the exact progression of topics. This guide synthesizes the book's promise and structure from the blurb and high-level description.
Excerpt 1
书名: Hands-On Generative AI with Transformers and Diffusion Models (Omar Sanseviero, Pedro Cuenca etc.) (Z-Library) 作者: Omar Sanseviero, Pedro Cuenca, Apoliná...
View in text
Tags
AI categories
Artificial IntelligencePythonGenerative AI
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 642
File Format: PDF
File Size: 33.2 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…