Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Sebastian Raschka

Rating No ratings yet

Learn how to create, train, and tweak large language models (LLMs) by building one from the ground up! Bestselling author Sebastian Raschka guides you step by step through creating your own LLM. Each stage is explained with clear text, diagrams, and examples. You’ll go from the initial design and creation, to pretraining on a general corpus, and on to fine-tuning for specific tasks. • Plan and code all the parts of an LLM comparable to GPT-2 • Prepare a dataset suitable for LLM training • Construct a complete training pipeline • Fine-tune LLMs for text classification and with your own data • Use human feedback to ensure your LLM follows instructions • Load pretrained weights into an LLM • Develop LLMs that follow human instructions This book takes you inside the AI black box to tinker with the internal systems that power generative AI. As you work through each key stage of LLM creation, you’ll develop an in-depth understanding of how LLMs work, their limitations, and their customization methods. Your LLM can be developed on an ordinary laptop, and used as your own personal assistant.This book is a practical and eminently-satisfying hands-on journey into the foundations of generative AI. Without relying on any existing LLM libraries, you’ll code a base model, evolve it into a text classifier, and ultimately create a chatbot that can follow your conversational instructions. And you’ll really understand it because you built it yourself! Readers need intermediate Python skills and some knowledge of machine learning. The LLM you create will run on any modern laptop and can optionally utilize GPUs.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Build a Large Language Model (From Scratch) — Reading Guide ## 【One-Line Pitch】 A hands-on, code-first journey that teaches you to build, pretrain, and fine-tune your own GPT-2–scale LLM entirely from scratch—ideal for intermediate Python developers and ML practitioners who want to truly understand what's inside generative AI rather than just use it. ## 【Book Arc】 - **Opening (~0%–10%)**: Sets the stage with the book's core promise—building an LLM comparable to GPT-2 without relying on existing LLM libraries—and outlines the three-stage roadmap: architecture and data prep, pretraining, and fine-tuning. - **Early (~10%–33%)**: Covers the foundational building blocks: data preparation and sampling strategies, the attention mechanism, and the overall LLM architecture that will serve as the foundation model. - **Middle (~33%–67%)**: Moves into pretraining the foundation model on unlabeled data, including the training loop, model evaluation, and loading pretrained weights for further work. - **Late (~67%–90%)**: Transitions to fine-tuning—first for text classification using labeled datasets, then for creating a personal assistant or chat model using instruction datasets and human feedback. - **Ending (~90%–100%)**: Wraps up with the complete pipeline for developing LLMs that follow human instructions, consolidating everything into a functional personal assistant you built yourself. ## 【Key Takeaways】 - **Three-stage architecture is the backbone** (Early): The book's entire structure—(1) architecture and data prep, (2) pretraining, (3) fine-tuning—mirrors how real LLMs are built, giving you a mental model that applies beyond this project. - **Data preparation is half the battle** (Early): Before any model code, you must implement data sampling and understand how tokenized text is batched and fed into the model—a step often glossed over in other tutorials. - **Attention mechanisms are the core innovation** (Early): Understanding how attention works is essential to grasping why LLMs can handle long-range dependencies in text; the book implements this from scratch rather than importing it. - **Pretraining on unlabeled data creates the foundation** (Middle): The training loop and evaluation metrics are where you see your model actually learn language patterns, transforming random weights into a coherent text generator. - **Loading pretrained weights is a practical superpower** (Middle): The ability to load existing weights into your from-scratch architecture means you can skip expensive pretraining and jump straight to fine-tuning for your own tasks. - **Fine-tuning adapts the foundation to specific jobs** (Late): The same pretrained model can become either a text classifier (with labeled data) or a personal assistant (with instruction data)—showing the versatility of the foundation model approach. - **Human feedback makes models follow instructions** (Late): The final stage uses human feedback to align the model with user intent, turning a raw text generator into something that actually does what you ask. ## 【Reading Tips】 - **Skim the opening chapter** (~0%–10%): It's mostly roadmap and prerequisites; if you already know what an LLM is, jump straight to the data preparation section. - **Deep-read the attention mechanism and architecture chapters** (~10%–33%): These are the conceptual heart of the book; take time to trace through the diagrams and code examples carefully. - **Treat the pretraining chapters as a checkpoint** (~33%–67%): If you can get your model training and evaluating correctly here, the fine-tuning stages will feel like a natural extension rather than a struggle. - **Expect code-heavy sections throughout**: The book's value is in the implementation, so have your Python environment ready and type along rather than just reading. - **Use the fine-tuning chapters as a template** (~67%–100%): Even if you don't build the final chatbot, the classification fine-tuning workflow is directly reusable for your own projects. ## 【Coverage Limits】 The excerpts primarily cover the book's overall structure, stage breakdown, and high-level chapter topics. Detailed technical content—specific code implementations, mathematical formulations, and dataset preparation steps—is not covered in this guide. The excerpt at ~0% also contains an unrelated OCR fragment from a different book (iOS 16 For Beginners), which was disregarded. ##
Excerpt 1
书名: Build a Large Language Model (From Scratch) (Sebastian Raschka) (Z-Library) 作者: Sebastian Raschka Learn how to create, train, and tweak large language mo...
View in text
Excerpt 2
书名: iOS_16_For_Beginners (unknown) (Z-Library) 作者: unknown
View in text
Excerpt 3
GPUs. M A N N I N G Sebastian Raschka FROMSCRATCH BUILD A 1) Data preparation & sampling 2) Attention mechanism Building an LLM STAGE 1 Foundation model STA...
View in text
Page 6
ers to distinguish their products are claimed as trademarks. Where those designations appear in the book, and Manning Publications was aware of a trademark c...
View in text
Tags
AI categories
Artificial IntelligenceAIPython
ISBN: 9781633437166
Publish Year: 2024
Language: English
Pages: 370
File Format: PDF
File Size: 17.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…