The era of generic, one-size-fits-all AI is ending. While foundational models like GPT, Claude, and Gemini excel at general knowledge, the next wave of innovation belongs to fine-tuned, domain-expert models that know your business, your code, and your data. This book is a practical guide to moving from the fragility of prompt engineering to the robustness, security, and control of model ownership. Written by award-winning researcher and bestselling author Laurence Moroney, Fine-Tuning AI walks you through the complete process of taking data and using it to train expert models that you own and run locally and privately on your laptop, phone, or data center. Aimed at Python-proficient developers and data owners who have hit the limits of APIs, high per-token costs, latency, or privacy concerns, the book provides a clear blueprint for building efficient, cost-effective systems that can outperform generalist models on specialized tasks.
Understand the value of data and prepare it for fine-tuning
Build workflows for fine-tuning language models
Apply advanced adaptation techniques such as PEFT and LoRA
Move beyond text with multimodal image understanding and generation
Optimize, quantize, and serve private expert models in production
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical blueprint for Python developers who have hit the limits of prompt engineering and hosted APIs, showing how to turn your own data into private, locally-run expert models that beat generalists on your specific task. Read it if you care about cost, latency, privacy, and actually owning your AI.
【Book Arc】
- **Opening (~0%–10%)**: Frames the shift from generalist models to owned specialists, diagnosing the "API Wall" — crushing scale costs, latency, and the CISO's discomfort with sharing data. Establishes why fine-tuning beats renting intelligence.
- **Early (~10%–35%)**: Quantifies the hidden economics: the "instruction tax" of re-sending system prompts, context-window pollution, Time to First Token, and metadata/compliance risks. Introduces the customization spectrum (prompting → RAG → fine-tuning) and argues hardware is no longer the blocker.
- **Middle (~35%–55%)**: Moves into fundamentals — transformer architecture, embeddings, positional encoding, and attention mechanisms (Query/Key/Value) — so you understand what you are actually customizing before touching it.
- **Late (~55%–85%)**: (Excerpts do not cover this range in detail.) The table of contents points to data excavation, synthetic data, formatting, adaptation mechanics, the training loop, alignment, a case study, and vision-language models.
- **Ending (~85%–100%)**: (Excerpts do not cover this range.) Listed chapters suggest evaluation, optimization and quantization, and serving the expert model in production.
【Key Takeaways】
- **The "API Wall" is the real problem, not model quality** (Opening): cost of scale, unpredictable latency, and data-sharing risk separate a demo from a business.
- **The instruction tax is a hidden, recurring cost** (Early): re-sending rules and formatting in every system prompt wastes tokens and degrades performance via context-window pollution; fine-tuning bakes instructions into model weights permanently.
- **Latency is a UX problem, not just a cost problem** (Early): Time to First Token and Tokens Per Second directly shape whether users will pay.
- **Data ownership is a compliance and competitive issue** (Early): third-party APIs leak metadata (who, when, how much), and GDPR-style deletion guarantees are hard when data left your infrastructure.
- **Customization is a spectrum** (Early): prompting and RAG are smarter prompting; only fine-tuning burns new patterns into the model — hardest, but most valuable long-term.
- **You must understand the architecture you're modifying** (Middle): embeddings, positional encoding, and attention (Q/K/V) are the "DNA" a fine-tuner needs before rebuilding anything.
- **Hardware is no longer the blocker** (Early): unified memory means a standard laptop can fine-tune models that once required server farms.
【Reading Tips】
- Deep-read Chapter 1's economics and privacy arguments — they are the book's persuasive core and justify every later technique.
- Treat the architecture chapter as prerequisite, not optional: skim if you already know transformers, but don't skip the attention intuition.
- Keep the table of contents as a roadmap; several later chapters (data, training loop, alignment, serving) are the practical payoff and deserve hands-on reading.
- Come with a concrete domain task and dataset in mind — the book is framed around applying techniques, not abstract theory.
【Coverage Limits】
This guide is based on stratified excerpts covering roughly the first half of the book (Chapters 1–2 and the table of contents); the data preparation, training, alignment, multimodal, and deployment chapters are named but not detailed here.
Excerpt 1
are also available for most titles ( https://oreilly.com ). For more information, contact our corporate/institutional sales department: 800-998-9938 or corpo...
nstruction tax There’s also a more subtle inefficiency here. The Generalist model doesn’t understand anything about your specific task until you tell it. To...
essary. You can retrain from a clean dataset if you need to. Ultimately—you have control. Model Drift I’ve personally been very surprised by how many times I...
ical rules, and not semantics, one might argue that it does! The attention mechanism calculates attention scores for the word ‘it’, to see what it refers to....
urons, and each could be thought of as a ‘pattern detector’. When the input matches a pattern that the neuron has learned to recognize, it will activate stro...
ouped into 8 groups of 4, and each shares 8 key-value heads. This reduces the KV cache by 75% compared to standard multi-head attention, dramatically improvi...
ations like: Rare words would require enormous vocabularies. Is it worth having a token dedicated to the word ‘antidisestablishmentarianism’? Non-English Lan...
then a card with the GPU on it as a separate plug-in device. As the GPU needs to work with numbers, those numbers have to be in memory—and transferring to an...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Fine-Tuning AI (Laurence Moroney)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Fine-Tuning AI (Laurence Moroney)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment