Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Alexander Thomas

Rating No ratings yet

Large language models have reshaped what's possible in AI—but their size, cost, and complexity can make them difficult to use in real-world production. Small language models (SLMs) offer a more practical alternative: they're efficient, focused, and built for applications that demand agility, privacy, and control. Hands-On Small Language Models provides a hands-on guide to understanding, building, and deploying these compact models to power specialized agentic applications. Author Alex Thomas, principal data scientist at John Snow Labs, draws on years of experience in natural language processing and applied AI to show how SLMs are democratizing the generative AI landscape. Through clear explanations and guided projects, including the development of a multi-functional movie chatbot, you'll learn how to combine, deploy, and monitor SLMs both locally and in the cloud.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Hands-On Small Language Models: A Practical Guide to Building Efficient AI Applications ## 【One-Line Pitch】 A hands-on guide for developers and data scientists who want to build practical, efficient AI applications using small language models instead of expensive large ones—complete with a real-world movie chatbot project that walks you through setup, model selection, and deployment. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces the core concept of small language models (SLMs) versus large language models (LLMs), explaining why SLMs are more practical for production applications that need agility, privacy, and control. Sets up the book's throughline project: Theoros, an agentic movie-search and information system built with publicly available data. - **Early (~9%–23%)**: Walks through the complete environment setup—installing the Python Data Science Stack (NumPy, SciPy, pandas, Matplotlib, Scikit-learn, Jupyter), plus NLP libraries like NLTK and NetworkX. Introduces the Model Context Protocol (MCP) from Anthropic as the standard for connecting AI applications to external systems, along with LiteLLM for model access abstraction. - **Early (~23%–32%)**: Covers the remaining infrastructure: OpenRouter for accessing hosted models with a single API key, Docker for running containers like Ollama and vector databases, LibreChat as a model-agnostic chat interface for realistic testing, and Ollama for running models locally with an OpenAI-compatible API. - **Middle (~32%–50%)**: Focuses on data acquisition from multiple sources—Wikipedia for RAG support, Wikidata for structured queries, Hugging Face datasets, and Kaggle. Includes a practical first experiment: building a function that queries a local model to generate SPARQL queries for finding movie directors, demonstrating the real-world challenges of working with SLMs. - **Middle (~50%–55%)**: Concludes the setup phase with a candid assessment of SLM limitations—the SPARQL generation experiment fails because models trained on general internet content struggle with niche query languages. Transitions into the next major topic: how to properly measure and compare models from different families. ## 【Key Takeaways】 - **SLMs are specialized tools, not general-purpose chatbots** (Opening): Small language models trade broad capability for efficiency, privacy, and control—making them ideal for focused production applications where cost and speed matter more than encyclopedic knowledge. - **MCP standardizes agentic application architecture** (Early): The Model Context Protocol provides a standardized way to organize prompts, tools, and data, eliminating boilerplate and making applications portable across different model families and deployment targets. - **LiteLLM abstracts away model differences** (Early): By using LiteLLM, you can switch between locally hosted models (via Ollama) and hosted services (via OpenRouter) with minimal configuration changes—a critical flexibility for production environments with varying security requirements. - **Docker simplifies local AI infrastructure** (Early): Running Ollama and other services in containers avoids system-level installation complexity, making it feasible to run SLMs on commodity hardware or even edge devices. - **Data diversity matters for agentic applications** (Middle): The Theoros project deliberately uses mixed data types—structured relational data, free-text articles, and graph-like relationships—to demonstrate how SLMs handle different information formats in a single application. - **SLMs struggle with niche technical languages** (Middle): The SPARQL generation experiment reveals a key limitation: models trained on general internet content often fail at specialized query languages, highlighting the importance of task-specific evaluation before committing to a model. - **Task-specific metrics beat general language benchmarks** (Middle): For SLMs, measuring quality requires looking beyond general language ability toward functionality-specific metrics, which means designing evaluation datasets for your particular use case. ## 【Reading Tips】 - **Skim the installation sections if you're experienced**: Chapters covering environment setup (conda, Docker, Ollama) are straightforward—skip ahead if you've already worked with these tools, but don't miss the MCP and LiteLLM explanations, which are conceptually important. - **Deep-read the model selection chapter**: The discussion of measurement and evaluation frameworks (starting around 55%) is the intellectual core of the book—this is where you'll learn how to think about SLM quality beyond surface-level benchmarks. - **Follow the Theoros project as your guide**: Rather than treating each chapter as isolated, track how the movie chatbot evolves—it's designed to demonstrate patterns you can apply to your own domain. - **Pay attention to the failure examples**: The SPARQL experiment failure is more instructive than many successes—it shows the real-world limits of SLMs and why task-specific testing matters before deployment. - **Note the deployment flexibility**: The book emphasizes multiple deployment paths (local, cloud, hosted services), so consider which scenario matches your constraints as you read. ## 【Coverage Limits】 The excerpts cover the book's opening through the beginning of model selection (roughly the first 55%), including full environment setup and initial experiments. Chapters on multi-SLM agentic applications, testing and compliance, deployment, and monitoring are listed in the table of contents but not covered in the available material. ##
Page 3
go is a registered trademark of O’Reilly Media, Inc. Hands- on Small Language Models, the cover image, and related trade dress are trademarks of O’Reilly Med...
View in text
Page 9
install Python Data Science Stack, also known as the Python Data Stack and the Pydata Stack. It is a collection of popular open source Python libraries used ...
View in text
Page 12
particular model family, and it does have MCP integration. Even if you are using the Google Colab environment, you will want to set up Libre Here are the ins...
View in text
Page 17
ponse = completion( model="ollama_chat/llama3", messages=[{ "content": "Briefly answer the following question, please. What would be the best movie to show ...
View in text
Excerpt 5
have this curve, we can calculate the area under that curve. This will give a number between 0 and 1. Furthermore, we know that by deciding by flipping a coi...
View in text
Excerpt 6
cation. This will allow users to filter using finer-grained criteria when looking for a movie. Let’s consider what are the requirements for this task. First ...
View in text
Excerpt 7
r subgenre prediction use-case. Subgenre assignment prompt. Please identify if the given movie belongs to one of the given subgenre. Genre: {genre} Possible ...
View in text
Excerpt 8
upervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Context length:...
View in text
Tags
AI categories
AIProgramming LanguageBackend
Publish Year: 2026
Language: English
File Format: PDF
File Size: 4.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…