LLM reasoning models have the power to tackle truly challenging problems that require finding the right path through multiple steps. In this book you’ll learn how to build a working reasoning model from the ground up. You will start with an existing pre-trained LLM and then implement reasoning-focused improvements from scratch.
Sebastian Raschka, the bestselling author of Build a Large Language Model (From Scratch), is your guide on this exciting journey. Sebastian mentors you every step of the way with clear explanations, practical code, and a keen focus on what really matters. Understand LLM reasoning by creating your own reasoning model–from scratch!
In Build A Reasoning Model (From Scratch) you’ll learn how
• Implement core reasoning improvements for LLMs
• Evaluate models using judgment-based and benchmark-based methods
• Improve reasoning without updating model weights
• Use reinforcement learning to integrate external tools like calculators
• Apply distillation techniques to learn from larger reasoning models
• Understand the full reasoning model development pipeline
Reasoning models break problems into steps, producing more reliable answers in math, logic, and code. These improvements aren’t just a curiosity–they’re already integrated into top models like Grok 4 and GPT-5. Build A Reasoning Model (From Scratch) demystifies these complex models with a simple the best way to learn how something works is to build it yourself! You’ll begin with a pre-trained LLM, adding and improving its reasoning capabilities in ways you can see, test, and understand.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A hands-on guide to turning a pretrained LLM into a reasoning model, teaching you the full pipeline—evaluation, inference-time scaling, reinforcement learning, and distillation—by building each piece in code. Best for developers and ML practitioners who already know basic LLM mechanics and want to understand reasoning models by constructing one.
【Book Arc】
- **Opening (~0%–10%)**: Frames what reasoning models are, how they differ from standard instruction-tuned LLMs, and why building one from scratch is the clearest path to understanding. Introduces the token/pretraining/post-training vocabulary the rest of the book relies on.
- **Early (~10%–35%)**: Sets up the coding environment and walks through generating text with a pretrained LLM—tokenization, input preparation, streaming generation, and performance touches like `torch.compile`.
- **Middle (~35%–50%)**: Shifts to evaluation, building a math verifier that extracts, normalizes, and grades answers against reference solutions, and introducing verifiable rewards as groundwork for later RL.
- **Late (~50%–85%)**: Covers reasoning improvements without weight updates (inference-time scaling, voting, self-refinement) and then reasoning training, including reinforcement learning with external tools like calculators.
- **Ending (~85%–100%)**: Applies distillation to learn from larger reasoning models and ties the stages together into the full reasoning-model development pipeline.
【Key Takeaways】
- **Reasoning models are built on top of a pretrained LLM, not from zero** (Opening): the book starts from existing weights and layers reasoning-focused improvements on top, which keeps the scope practical and testable.
- **Pretraining, instruction tuning, and preference tuning are distinct stages** (Opening): understanding this separation clarifies where reasoning improvements fit and why an instruction-tuned model is still not a chatbot.
- **Evaluation is a first-class engineering problem** (Middle): the book implements a math verifier that extracts, normalizes, and grades answers, handling fractions, LaTeX, decimals, and percentages—robust grading is what makes later training signals trustworthy.
- **Verifiable rewards are the bridge from evaluation to reinforcement learning** (Middle): the same verifier logic used to score math answers becomes the reward signal for RL in later chapters.
- **Reasoning can improve without updating model weights** (Late): inference-time scaling techniques such as advanced text generation, voting, and self-refinement offer gains before any training is involved.
- **Reinforcement learning can integrate external tools** (Late): the book shows how RL can teach a model to use tools like calculators, extending reasoning beyond what the model alone can do.
- **Distillation transfers reasoning ability from larger models** (Ending): smaller models can learn reasoning behavior by learning from larger reasoning models rather than training from scratch.
- **The full pipeline is the point** (Ending): evaluation, inference-time techniques, RL, and distillation are presented as connected stages, not isolated tricks.
【Reading Tips】
- Deep-read the evaluation chapter even if you care mainly about training—the verifier and grading logic underpin the RL rewards later, so skimming it will make chapter 6 feel arbitrary.
- Treat the environment setup and text-generation chapters as a working baseline: get the code running end-to-end before moving on, since later chapters reuse the wrapper and generation utilities.
- Skim the tokenizer internals (BPE, token ID printing) if you already know them; the conceptual payoff is small compared to the verifier and RL sections.
- Pay attention to the stage diagram (evaluation → inference-time → training) as a mental map; it is the book's organizing spine and helps you place each technique.
- When reading the RL and distillation chapters, keep asking "what signal is the model learning from?"—that question connects verifiable rewards, tool use, and distillation into one story.
【Coverage Limits】
The excerpts cover the book's framing, environment setup, tokenization, generation, and the math verifier in detail, but the later chapters on inference-time scaling, reinforcement learning, and distillation are represented mainly by chapter listings and brief mentions. Specific implementation details, code, and results for those later stages are not covered here.
Page 8
57 4 ■ Improving reasoning with inference-time scaling 94 5 ■ Inference-time scaling via self-refinement 136 6 ■ Training reasoning models with reinforcement...
esolution more reliably than pip. It also creates isolated environments automatically and comes with its own Python executable (but will use the system Pytho...
If a token such as "Berlin" stands out more strongly than the alternatives, it is more likely to be selected later. If the differences shrink, lower- ranked...
is loaded correctly, let’s use it with the temperature and top-p sampling code from the previous chapter on a MATH-500 prompt. Listing 5.2 Generating text wi...
int log-probability: tensor(-29.8750, dtype=torch.bfloat16) As you can see, the difference between the "Berlin" and "Bridge" is now much more pronounced (-16...
ceived quality. The idea is that the reward model can auto- matically score new model outputs, eliminating the need for human annotation at every training st...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Build a Reasoning Model (From Scratch) (Sebastian Raschka) (z-library.sk, 1lib.sk, z-lib.sk)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Build a Reasoning Model (From Scratch) (Sebastian Raschka) (z-library.sk, 1lib.sk, z-lib.sk)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment