AI has acquired startling new language capabilities in just the past few years. Driven by the rapid advances in deep learning, language AI systems are able to write and understand text better than ever before. This trend enables the rise of new features, products, and entire industries. With this book, Python developers will learn the practical tools and concepts they need to use these capabilities today.
You'll learn how to use the power of pre-trained large language models for use cases like copywriting and summarization; create semantic search systems that go beyond keyword matching; build systems that classify and cluster text to enable scalable understanding of large amounts of text documents; and use existing libraries and pre-trained models for text classification, search, and clusterings.
This book also shows you how to:
Build advanced LLM pipelines to cluster text documents and explore the topics they belong to
Build semantic search engines...
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Hands-On Large Language Models
## 【One-Line Pitch】
A practical, code-first guide for Python developers who want to move beyond ChatGPT chat windows and build real applications—semantic search, text classification, topic modeling, and LLM pipelines—using open-source models from Hugging Face. If you learn best by running code and want to understand what's actually happening inside these models, this is your book.
## 【Book Arc】
- **Opening (~0%–11%)**: Part I establishes the conceptual foundation—what language models are, how tokenization splits text into manageable pieces, and how embeddings capture semantic meaning. The famous "Illustrated Transformer" is updated and expanded here, giving readers the architectural vocabulary used throughout the rest of the book.
- **Early (~11%–25%)**: Hands-on work begins with loading models and tokenizers from Hugging Face, then dives deep into the Transformer architecture itself—attention mechanisms, positional encodings (including modern rotary embeddings), and recent architectural improvements that distinguish today's LLMs from the original paper.
- **Early–Middle (~25%–39%)**: The focus shifts to practical NLP tasks: text classification with both encoder-only models (BERT and its variants) and generative models (Flan-T5, ChatGPT), plus zero-shot classification when labeled data is scarce. This section introduces BERTopic for clustering and topic modeling, showing how to build modular pipelines.
- **Middle (~39%–50%)**: Prompting becomes the central theme—how to design effective prompts for different tasks, use prompt templates for reusability, and build chains with memory using LangChain. The ReAct pattern (Thought → Action → Observation) is introduced for connecting LLMs to external tools like calculators and search engines.
- **Late (~50%–end)**: The book continues into advanced applications—semantic search engines, RAG (retrieval-augmented generation), and fine-tuning approaches—though the excerpts primarily cover the earlier pipeline-building stages.
## 【Key Takeaways】
- **Tokenization is the hidden gatekeeper of LLM performance** (Early): The tokenizer determines how text is split before the model ever sees it—BPE is the dominant method, but vocabulary size, initialization parameters, and training data domain all shape tokenizer behavior. Understanding this matters because tokenization directly affects model efficiency and language coverage.
- **Embeddings are the semantic currency of modern NLP** (Early): Word and sentence embeddings place similar meanings close together in vector space, enabling similarity measurement. High-quality embedding models are trained specifically for this task—averaging token embeddings is a common but inferior shortcut.
- **Encoder-only models (BERT family) remain the workhorses for task-specific NLP** (Early): While generative models like GPT get the attention, BERT and its variants (RoBERTa, DistilBERT, ALBERT, DeBERTa) excel at classification and embedding tasks while being significantly smaller. Model selection requires weighing language compatibility, architecture, size, and performance across 60,000+ available options.
- **Zero-shot classification is a viable first step when labeled data is scarce** (Early): Labeled data is expensive to produce, but zero-shot approaches let you test task feasibility using only label names. Flan-T5 achieves an F1 score of 0.84 on sentiment classification without any fine-tuning—a strong baseline before investing in annotation.
- **BERTopic brings modularity to topic modeling** (Middle): The framework separates representation, clustering, and representation refinement into interchangeable blocks. Generative models can be injected as a reranking step—not to label millions of documents, but to generate short, human-readable topic labels from keywords and representative documents.
- **Prompt templates and chains turn single LLM calls into reusable workflows** (Middle): Using LangChain, you can parameterize prompts (e.g., "Create a funny name for a business that sells {product}") and add modular components like memory. Memory is finite—with a window of two conversations, earlier context is forgotten, which is a critical design constraint.
- **The ReAct pattern connects LLMs to the external world** (Middle): Instead of relying on the model's internal knowledge, you can structure interactions as Thought → Action → Observation loops, where actions invoke external tools (calculators, search engines) and observations feed results back into the model's reasoning.
## 【Reading Tips】
- **Skim the Transformer architecture chapter if you're already familiar with attention**—the key new material is on rotary embeddings and recent architectural improvements, which are worth a careful read.
- **Run the code examples as you go**—the book is explicitly hands-on, and the Hugging Face ecosystem (transformers, sentence-transformers, BERTopic) is best learned by executing. The Phi-3-mini example in Chapter 2 is a good first checkpoint.
- **Pay special attention to the model selection discussions**—with tens of thousands of models on Hugging Face Hub, the book's guidance on choosing between encoder-only and generative architectures for different tasks is practical wisdom you'll reuse constantly.
- **The BERTopic chapter rewards careful reading**—the modular pipeline design (representation → clustering → reranking) is a template you can apply beyond topic modeling to other text analysis problems.
- **If you're primarily interested in building production systems**, focus on the prompting and chains chapters—they show how to move from single prompts to reusable, memory-aware workflows with external tool integration.
## 【Coverage Limits】
This guide covers the book's conceptual foundations, model architecture, text classification, topic modeling, and prompting/chain-building content visible in the excerpts. Fine-tuning, RAG, and advanced semantic search details appear in later chapters not fully represented in the source material.
##
Excerpt 1
ess support and encouragement. 感谢我的家人和朋友们的坚定不移的支持,尤其是我的父母。爸 爸,尽管你面临了挑战,但你总是能在我最需要的时候找到办法陪 伴我,谢谢你。妈妈,我们作为有抱负的作家所进行的对话非常美 妙,比你想象的更能激励我。感谢你们两位无止境的支持和鼓励。 Finally...
quite reasonable list! Now that we know it works, let’s see how to build such a system. 另一份相当合理的列表!既然我们知道它可行,那就让我们看看如 何构建这样一个系统。 Training a Song Embedding Mo...
t involved an important component, namely preference tuning. As mpnet-base-v2” model we used in the previous chapter, we use the “thenlper/gte-small” model i...
oost performance from 36.5 to 62.8, measured as Figure 8-25. Generative search formulates answers and summaries at the end of a search pipeline while citing ...
ocused models, The domain of the data 以文本为中心与以代码为中心的模型,数据领域 token embeddings, Token Embeddings-Creating Contextualized Word Embeddings with Language Models, ...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Hands-On Large Language Models (Jay Alammar, Maarten Grootendorst)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Hands-On Large Language Models (Jay Alammar, Maarten Grootendorst)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment