Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Thársis T. P. Souza, Jonathan K. Regenstein Jr.

Large language models (LLMs) have transformed natural language processing, but deploying them in applications introduces numerous technical challenges. [This book] offers a clear, practical examination of the limitations developers and AI engineers face when building LLM-based applications. With a focus on implementation pitfalls (not just capabilities), this book provides actionable strategies supported by reproducible Python code and open source tools. Readers will learn how to navigate key obstacles in application evaluation, input management, testing, and safety. Designed for builders and technical product leads, this guide emphasizes practical solutions to real-world problems and promotes a grounded understanding of LLM constraints and trade-offs. - Design testing and evaluation strategies for nondeterministic systems - Manage context, RAG, and long-context retrieval - Address output inconsistency and structural unreliability - Implement safety and content moderation frameworks - Explore alignment challenges and mitigation techniques - Leverage open source models locally

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Large Language Models: The Hard Parts — Reading Guide ## 【One-Line Pitch】 A practical, open-source-focused field guide for engineers, data scientists, and technical leaders who need to move beyond "playing with LLMs" to building reliable, production-grade LLM applications—covering evaluation, safety, context management, and alignment with reproducible Python code. ## 【Book Arc】 - **Opening (~0%–9%)**: Establishes the book's core premise—that the hard parts of LLM work are evaluation, safety, context, alignment, structured outputs, RAG, and fine-tuning—and defines the target audience (data scientists, ML engineers, product managers, domain experts, and anyone tasked with "figuring out how we use AI"). The authors commit to open source tools for reproducibility and transparency. - **Early (~9%–25%)**: Covers first principles: strategic questions about value, data, stakeholders, and compliance before building; model types (base, instruction-tuned, domain-adapted); licensing considerations for enterprise use (Apache 2.0, MIT, custom licenses like Llama 3 and Qwen2.5); and the nondeterministic nature of LLMs, including temperature's effect on output distribution. - **Early–Middle (~25%–38%)**: Introduces the LLMBA (LLM-Based Application) evaluation framework—a conceptual design for comparing multiple configurations (models, prompts, data sources) against a task, culminating in a leaderboard. Covers metrics-based evaluation (BLEU, ROUGE variants including ROUGE-Lsum) with worked Python examples, then transitions to LLM-as-a-judge methodology, including its limitations (positional bias, length bias, prompt sensitivity, domain expertise gaps). - **Middle (~38%–47%)**: Explores the broader evaluation landscape: how to evaluate the evaluators themselves (comparing LLM judges against golden datasets or human scores), public leaderboards (Chatbot Arena, AlpacaEval, MT-Bench, LiveBench), and the ARC-AGI benchmark for measuring fluid intelligence. Introduces LangSmith as a production-grade evaluation tool with human-in-the-loop feedback and monitoring capabilities, demonstrated through a 10-K summary use case. ## 【Key Takeaways】 - **Strategic clarity precedes technical excellence** (Early): Before building, teams must answer questions about value, data readiness, stakeholder alignment, and compliance—technical skill alone doesn't determine success. This frames the entire book's approach. - **Model selection involves licensing trade-offs** (Early): Apache 2.0 (Mistral) and MIT (Phi-3) offer maximum commercial freedom; Llama 3 and Qwen2.5 have user thresholds (700M and 100M) and prohibit using outputs to train competing models. Choose based on scale and future AI development plans. - **LLMs are text completion engines, not conversational assistants** (Early): Base models continue text statistically rather than "answer" questions—understanding this shapes expectations for what fine-tuning and prompting can achieve. - **Temperature controls output randomness** (Early): At temperature = 1, the learned distribution is used as-is; higher values flatten probabilities, making surprising tokens more likely. This is a key lever for balancing creativity vs. consistency. - **Single metrics and single examples are unreliable** (Early): BLEU and ROUGE scores vary significantly even on simple sentences—evaluation requires representative datasets and aggregate statistics, not one-off tests. - **LLM-as-a-judge is scalable but biased** (Early–Middle): Positional bias, egocentric bias, length bias, and prompt sensitivity are real limitations; in professional domains (finance, law, medicine), judge LLMs may lack expertise and need fine-tuning or human oversight. - **Evaluate your evaluators** (Middle): Judge model quality is measured by correlation with golden datasets or human scores—this validation step is essential before trusting automated evaluation at scale. - **Production monitoring requires dedicated tooling** (Middle): LangSmith-style platforms provide continuous evaluation, alerting, human feedback collection, and trend visualization—infrastructure that teams would otherwise need to build themselves. ## 【Reading Tips】 - **Skim the licensing tables and model comparisons** (Early): The specific license terms will change over time; focus on the decision framework (freedom vs. restrictions, scale thresholds) rather than memorizing current details. - **Deep-read the evaluation chapters** (Early–Middle): The BLEU/ROUGE examples and LLM-as-a-judge discussion are the book's core value—work through the Python code to understand the mechanics, not just the concepts. - **Pay attention to the LLMBA framework** (Early): The conceptual design for comparing multiple application configurations is a mental model you can reuse across projects, even if you don't adopt the exact implementation. - **Treat the leaderboard discussion as context, not prescription** (Middle): Public benchmarks like Chatbot Arena and LiveBench are useful for model selection but shouldn't replace task-specific evaluation of your own application. - **If you're a manager or product lead** (Opening): Focus on the first principles chapter and the evaluation strategy discussions; the code-heavy sections can be skimmed for vocabulary and decision frameworks. ## 【Coverage Limits】 Excerpts cover roughly the first half of the book (through ~47%), focusing on evaluation, model selection, and tooling. Later chapters on RAG, structured outputs, safety frameworks, alignment, and fine-tuning are mentioned in the table of contents but not covered in this guide's source material. ##
Page 13
252 Alignment Evaluation: LLM-as-a-Judge 259 DPO Dataset Composition 267 Our Choice of Base Model 268 The Evaluation Methodology 269 The Fine-Tuning Process...
View in text
Excerpt 2
mercially with minimal restrictions, making them the safest choice for enterprises that want full flexibility without legal complexity. Custom commercial lic...
View in text
Excerpt 3
s both the reference and candidate summaries for comparison • Requests scores on a standardized 1–10 scale 4. Finally, we use OpenAI’s API with the client.be...
View in text
Excerpt 4
5- UNITED Apple Inc. filed its Apple Inc.’s 10-K 0.386076 0.704104 turbo STATES\nSECURITIES annual Form 10-K filing for the fiscal AND EXCHANGE for the year...
View in text
Excerpt 5
is what comes together with a prompt and a model to form an LLMBA, and it’s what allows our application to answer a question or accomplish a task while being...
View in text
Excerpt 6
the document that are related to or similar to that query. These are the chunks that let the LLM answer the query: def retrieve_and_summarize_risks(vectorsto...
View in text
Excerpt 7
5.1 billion shares of common stock outstanding." } } 140 | Chapter 5: Structured Data Output prompt example “Is Enzo a good name for a baby?” by running inpu...
View in text
Excerpt 8
ber of parameters or insert LLM-Specific Safety Risks | 163 Tasking the preparedness team A dedicated team drives the technical work of the Prep
View in text
Tags
AI categories
Artificial IntelligenceAIProgramming
ISBN: 8341622521
Publisher: O'Reilly Media
Publish Year: 2026
Language: English
Pages: 341
File Format: PDF
File Size: 16.6 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…