Engineering Generative-AI Based Software discusses both the process of developing this kind of AI-based software and its architectures, combining theory with practice. Sections review the most relevant models and technologies, detail software engineering practices for such systems, e.g., eliciting functional and non-functional requirements specific to generative AI, explore various architectural styles and tactics for such systems, including different programming platforms, and show how to create robust licensing models. Finally, readers learn how to manage data, both during training and when generating new data, and how to use generated data and user feedback to constantly evolve generative AI-based software.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Engineering Generative AI-Based Software — Reading Guide
## 【One-Line Pitch】
A practical engineering handbook for software developers and architects who want to move beyond using generative AI as a toy and instead build production-grade systems around large language models, covering everything from model training to architecture, quality, and data management.
## 【Book Arc】
- **Opening (~0%–10%)**: Introduces the technological context—how mobile computing and the transformer architecture (Vaswani, 2017) reshaped software development—and sets up the book's dual focus on both the process and architecture of generative AI software.
- **Early (~10%–23%)**: Walks through hands-on model training with HuggingFace tooling, including masked language modeling with data collators, fine-tuning approaches, and embedding visualization with t-SNE; also outlines the book's chapter structure covering instruct models, software components, quality, architecture, and development processes.
- **Early (~23%–32%)**: Explains instruct models—how they're built from pre-trained language models fine-tuned on prompt-response pairs—and demonstrates prompt engineering techniques including zero-shot, one-shot, and few-shot learning, with concrete Python examples.
- **Middle (~32%–42%)**: Discusses when to use pure model generation versus hybrid development combining models with traditional software, covering RAG (retrieval-augmented generation) patterns and the practical trade-offs of each approach.
- **Middle (~42%–48%)**: Shifts to software engineering concerns: architectural decisions (microservices vs. local deployment), testing strategies that move from oracle-based to ML-based testing, and the principle of focusing on customer value over model perfection.
## 【Key Takeaways】
- **Transformers enabled the generative AI era** (Early): The combination of multi-head attention with encoder-decoder architecture allowed models to capture context beyond adjacent words and separate pre-training from task-specific fine-tuning, making self-supervised learning at scale practical.
- **Self-supervised training uses masking** (Early): Data collators with masked language modeling (e.g., 15% token masking probability) prepare training data without human labels, and the HuggingFace Trainer API makes this accessible with just a few lines of Python.
- **Instruct models are fine-tuned conversational agents** (Early): Starting from a pre-trained language model, you fine-tune on prompt-response datasets to create models that generalize to new questions rather than memorizing training examples; the prompt can be abridged or modified by the UI before reaching the model.
- **Prompt engineering has a spectrum of strategies** (Middle): Zero-shot and one-shot learning add context to steer attention heads, while few-shot prompting with examples enables reasoning; each strategy trades off prompt complexity against model capability.
- **Hybrid development is often necessary** (Middle): Pure model generation works for simple tasks, but complex problems requiring abstract thinking or strict template adherence (forms, test case specifications) benefit from combining models with traditional programs that orchestrate and synchronize outputs.
- **Focus on customer value, not model perfection** (Middle): OpenAI's ChatGPT launch illustrates the lean start-up principle—shipping a minimum viable product and iterating beats waiting for the perfect model; technology evolves too quickly to delay.
- **Testing shifts from oracles to ML-based approaches** (Middle): Traditional oracle-based testing (comparing against expected outputs) transitions to machine-learning-based testing, which requires new evaluation metrics and strategies for validating generative outputs.
## 【Reading Tips】
- **Skim the historical context** (Opening): The iPhone/transformer backstory is useful framing but not essential; move quickly to the technical content.
- **Deep-read the code listings**: The HuggingFace examples (data collators, Trainer setup, feature extraction, t-SNE visualization) are the book's practical core—reproduce them to build muscle memory.
- **Pay attention to the prompt-response dataset format**: The JSON example showing instruction/input/output/prompt structure is the key to understanding how instruct models are fine-tuned; study it carefully.
- **Note the BLEU metric discussion**: The book shows how BLEU ignores word order (giving perfect scores to nonsense sentences)—a cautionary tale for anyone evaluating generative output naively.
- **Use the chapter roadmap as your guide**: The book explicitly outlines its structure in Chapter 1; if you're only interested in architecture or quality, jump to those chapters directly.
## 【Coverage Limits】
This guide covers the opening through roughly the middle of the book (model training, instruct models, prompt engineering, and early architecture/testing discussions). The excerpts do not cover the later chapters on quality models, architectural styles in depth, licensing models, or data management during training and generation—those sections are beyond this guide's scope.
##
Page 5
products does not constitute endorsement or sponsorship by The MathWorks of a particular pedagogical approach or particular use of the MATLAB® software. No ...
11 d a t a _ c o l l a t o r= d a t a _ c o l l a t o r, 12 t r a i n _ d a t a s e t= t o k e n i z e d _ d a t a s e t[ ’ t r a i n’ ] ) 13 14 t r ...
d similar technologies. Chapter 2 • Generative AI basics 19 The first part of the model – the language model – is pre-trained using text relevant to the dom...
(Abrahão et al., 2024). In traditional testing, we use pre- defined oracles to check the results, but in AI testing, we must be more flexible due to the pro...
P r o g r a m = j s o n _ d a t a[ ’ r e s p o n s e’ ] 12 r e t u r n s t r P r o g r a m Listing 3.7: Python function that uses Ollama framework to gen...
prompts, communication channels, and storage of data; all to ensure that the customers’ data is safe with them. • Metadata – The extent of documentation ava...
s monolithic architecture. In the monolithic architecture, all modules of the software form one large component; we do not care much about the connections ...
e f g e n e r a t e _ t e x t( s e l f, p r o m p t) : 9 # C h e c k i f t h e p r o m p t i s l e s s t h a n 5 1 2 w o r d s 10 i f l e n( p r o m p ...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Loading comments...
Reply to Comment
Edit Comment