本书旨在为对大语言模型感兴趣的读者系统地讲解相关基础知识、介绍前沿技术。作者团队将认真听取开源社区以及广大专家学者的建议,持续进行月度更新,致力打造易读、严谨、有深度的大模型教材。并且,本书还将针对每章内容配备相关的Paper List,以跟踪相关技术的最新进展。 本书第一版包括传统语言模型、大语言模型架构演化、Prompt工程、参数高效微调、模型编辑、检索增强生成等六章内容。为增加本书的易读性,每章分别以一种动物为背景,对具体技术进行举例说明,故此本书以六种动物作为封面。当前版本所含内容均来源于作者团队对相关方向的探索与理解,如有谬误,恳请大家多提issue,多多赐教。后续,作者团队还将继续探索大模型推理加速、大模型智能体等方向。相关内容也将陆续补充到本书的后续版本中,期待封面上的动物越来越多。
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
【One-Line Pitch】
A systematic, open-source textbook on large language models (LLMs) that walks readers from classical language modeling through modern architectures, prompt engineering, efficient fine-tuning, model editing, and retrieval-augmented generation—ideal for students, researchers, and practitioners who want a rigorous yet accessible foundation in LLM technology.
【Book Arc】
- **Opening (~0%–15%)**: Introduces the book's mission, author contributions, and chapter roadmap, then lays the groundwork with classical language models—covering statistical n-grams, RNN-based models, and Transformer-based models—along with sampling methods and evaluation metrics (intrinsic and extrinsic).
- **Early (~15%–40%)**: Explores LLM architectures in depth, starting with the "big data + big model → new intelligence" paradigm, then comparing encoder-only (BERT), encoder-decoder (T5, BART), and decoder-only (GPT, LLaMA) designs, plus non-Transformer alternatives like state space models (SSM) and test-time training (TTT).
- **Middle (~40%–67%)**: Dives into prompt engineering—defining prompts, in-context learning, chain-of-thought (CoT) reasoning, and practical prompting techniques—then applies these to real-world use cases such as LLM agents, data synthesis, Text-to-SQL, and GPTs.
- **Late (~67%–90%)**: Covers parameter-efficient fine-tuning (PEFT), explaining why it matters for downstream task adaptation, and detailing methods like parameter addition (input/model/output), parameter selection, and low-rank adaptation (LoRA) with its variants and task generalization.
- **Ending (~90%–100%)**: Focuses on model editing—its philosophy, definitions, properties, and datasets—then compares classic approaches (external extension vs. internal modification) and walks through specific techniques like T-Patcher (additional parameters) and ROME (locate-and-edit).
【Key Takeaways】
- **Language models evolved from statistics to deep learning** (Opening): The book traces n-gram models, RNNs, and Transformers, showing how each step addressed prior limitations—essential for understanding why LLMs work the way they do.
- **Architecture choice determines model capabilities** (Early): Encoder-only, encoder-decoder, and decoder-only designs serve different tasks; the book compares them functionally and historically, helping readers pick the right backbone for their use case.
- **Non-Transformer architectures are emerging alternatives** (Early): SSMs and TTT models challenge the dominance of Transformers, offering new efficiency and adaptability trade-offs worth tracking for future-proofing skills.
- **Prompt engineering is a skill with structure, not magic** (Middle): From in-context learning to chain-of-thought, the book breaks down how to select demonstrations, influence performance, and apply techniques like step-by-step reasoning and psychological cues.
- **Prompts unlock practical applications** (Middle): Real-world uses like LLM agents, data synthesis, and Text-to-SQL show how prompt design translates into tangible products, bridging theory and deployment.
- **Parameter-efficient fine-tuning saves resources without sacrificing quality** (Late): PEFT methods—especially LoRA—let you adapt large models to downstream tasks with minimal trainable parameters, making customization feasible on limited hardware.
- **Model editing offers surgical updates without full retraining** (Ending): Techniques like T-Patcher and ROME modify specific knowledge or behaviors, balancing precision, generalization, and robustness—critical for maintaining models post-deployment.
【Reading Tips】
- **Skim the opening chapter** if you're already familiar with n-grams, RNNs, and Transformers; focus instead on the sampling and evaluation sections, which are practical and often glossed over elsewhere.
- **Deep-read Chapter 2** (architecture) if you're choosing a model for a project—the functional comparison and historical evolution will help you justify your choice.
- **Treat Chapter 3 (prompt engineering) as a hands-on playbook**: Try the techniques (e.g., CoT, demonstration selection) on your own tasks, and revisit the application section for agent and Text-to-SQL patterns.
- **For PEFT and model editing (Chapters 4–5), prioritize the method taxonomies** (e.g., parameter addition vs. selection; external vs. internal editing) over individual papers; the book's structure helps you categorize new research you'll encounter later.
- **Use the Paper List per chapter** as a reading list to go deeper—the book is designed as a living document, so pair it with recent literature for cutting-edge updates.
【Coverage Limits】
This guide synthesizes the table of contents and chapter structure; detailed technical content (e.g., specific formulas, code examples, or full method walkthroughs) is not covered here, and the excerpts do not include the book's final chapters on retrieval-augmented generation or future topics like inference acceleration and agents.
Excerpt 1
书名: 大模型基础 (毛玉仁,高云君) (Z Library) 作者: 毛玉仁,高云君 本书旨在为对大语言模型感兴趣的读者系统地讲解相关基础知识、介绍前沿技术。作者团队将认真听取开源社区以及广大专家学者的建议,持续进行月度更新,致力打造易读、严谨、有深度的大模型教材。并且,本书还将针对每章内容配备相关的Paper...
View in text
Page 4
. 35 2.1.2 大数据 +大模型→能力扩展 . . . . . . . . . . . . . . . . . . 38 2.2 大语言模型架构概览 . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 2.2.1 主流模型架构的类别 . . . ...
View in text
Page 5
. . . . . . . . . . . . . . . . . 110 3.2.3 性能影响因素 . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 3.3 思维链 . . . . . . . . . . . . . . . . . . . . ...
View in text
Page 6
. . . . . . . . . . . . . . . . . . . . 165 4.3.2 基于学习的方法 . . . . . . . . . . . . . . . . . . . . . . . . . . 165 4.4 低秩适配方法 . . . . . . . . . . . . . . . . ...
View in text
Tags
AI categories
AIProgramming LanguageTechnology
Text Preview (First 20 pages)
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Generating text preview…
Loading comments...
Reply to Comment
Edit Comment