机器翻译:基础与模型 (niutrans.com) 这是一个教程,目的是对机器翻译的统计建模和深度学习方法进行较为系统的介绍(前身为《机器翻译:统计建模与深度学习方法》)。其内容被编纂成书,可以供计算机相关专业高年级本科生及研究生学习之用,亦可作为自然语言处理,特别是机器翻译相关研究人员的参考资料。本书用tex编写,所有源代码均已开放
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
# Machine Translation: Foundations and Models
## 【One-Line Pitch】
A comprehensive, open-source textbook that systematically covers machine translation from statistical modeling to modern deep learning approaches, written by the team behind the NiuTrans open-source MT system. Ideal for senior undergraduates, graduate students, and NLP researchers who want both theoretical depth and practical implementation insights.
## 【Book Arc】
- **Opening (~0%–10%)**: Introduces the field's history, from early rule-based systems through statistical methods to the deep learning era, and lays out the book's four-part structure covering foundations, statistical MT, neural MT, and frontier topics.
- **Early (~10%–29%)**: Builds essential groundwork—statistical language modeling, search algorithms, Chinese word segmentation, syntactic parsing, and translation quality evaluation methods (both human and automatic, with and without references).
- **Middle (~29%–43%)**: Dives into statistical machine translation, starting with word-based models (IBM Model 1), then distortion and fertility models (IBM Models 2–5), phrase-based models with log-linear frameworks and stack decoding, and syntax-based approaches using synchronous context-free grammars.
- **Middle (~43%–57%)**: Transitions to neural methods, beginning with neural network fundamentals—tensors, forward propagation, backpropagation, gradient optimization, and neural language models—before introducing the encoder-decoder framework and recurrent neural network-based translation models.
- **Late (~57%–100%)**: Covers advanced neural architectures (CNN-based and self-attention/Transformer models), then explores training, inference, structure optimization, low-resource translation, multimodal/multi-level translation, and practical application techniques drawn from the authors' competition and product development experience.
## 【Key Takeaways】
- **Statistical modeling is the backbone of MT** (Early): The book builds from language modeling fundamentals—n-gram models, smoothing, and perplexity—to show how probability theory underpins translation. This foundation is essential for understanding both classical and neural approaches.
- **Word alignment is the core problem in statistical MT** (Early): IBM Models 1–5 progressively refine alignment modeling, from simple lexical translation to distortion (reordering) and fertility (one-to-many mappings). Understanding these models reveals why phrase-based methods emerged as a practical improvement.
- **Phrase-based models shifted translation units from words to phrases** (Middle): By extracting phrase pairs consistent with word alignments and using log-linear models with multiple features, these systems solved many word-level translation problems. The book covers phrase extraction, reordering models, and minimum error rate training in detail.
- **Syntax-based models add structural constraints** (Middle): Hierarchical phrase models using synchronous context-free grammars and linguistically-informed syntax models both attempt to capture reordering patterns that pure phrase models miss, with CKY decoding and cube pruning as key implementation techniques.
- **Neural MT fundamentally changed the paradigm** (Middle): The encoder-decoder framework with attention mechanisms replaced explicit alignment with learned representations. The book traces this evolution from RNN-based models (LSTM, GRU, bidirectional, multi-layer) to the attention mechanism that became the foundation for modern architectures.
- **Training neural MT requires mastering practical optimization** (Late): Loss functions, gradient-based optimization, parallelization strategies, handling vanishing/exploding gradients, and overfitting are covered in depth—critical knowledge for anyone actually building NMT systems.
- **Low-resource and multimodal MT are active frontiers** (Late): The book addresses unsupervised translation, leveraging monolingual data, and extending MT to speech and image inputs, plus document-level translation—areas where the field continues to evolve beyond sentence-level paradigms.
## 【Reading Tips】
- **Skim Chapter 1** if you already know MT history; the real value starts with Chapter 2's statistical foundations, which are prerequisites for everything that follows.
- **Deep-read Chapters 5–8** (statistical MT) even if you focus on neural methods—the concepts of alignment, decoding, and reordering directly inform how you understand attention mechanisms and Transformer architectures.
- **Treat Chapter 9 as a self-contained ML refresher**: if you're comfortable with neural networks, you can move quickly; if not, this chapter provides the necessary background on tensors, backpropagation, and optimization.
- **Use the "Summary and Further Reading" sections** at each chapter's end to decide what to explore deeper—the book is designed so each chapter can stand alone as a topic reference.
- **Pair the book with the open-source NiuTrans codebase** (available on GitHub) to see the algorithms implemented in practice, especially for decoding and training chapters.
## 【Coverage Limits】
This guide covers the book's structure and key technical themes based on the table of contents and chapter overviews. Detailed mathematical derivations, specific algorithm pseudocode, and the full content of later chapters (particularly 13–18) are not covered in the available excerpts.
##
Page 4
是,朋 友、同事们一直鼓励将内容正式出版。虽然担心书的内容不够精致,无法给同行作 为参考,但是最终还是下定决心重构内容。所幸,得到电子工业出版社的支持,形成 新版,共十八章。 写作中,每当笔者翻起以前的资料,都会想起当年的一些故事。与其说这部书 是写给读者,还不如说这本书是写给笔者自己,写给所有同笔者一样,经历过...
View in text
Page 10
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 2.3.3 语言模型的评价 . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 9
. . 161 5.5.3 训练 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162 5.6 小结及拓展阅读 . . . . . . . . ....
View in text
Page 13
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 238 8.3.2 基于树结构的文法 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 15
. . . . . . . . . . . . . . . . . . 357 10.5 训练及推断 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 359 10.5.1 训...
View in text
Page 17
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 435 13.5.1 什么是知识蒸馏 . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 18
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 517 16.1.1 数据增强 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 19
18.3 交互式机器翻译 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 580 18.4 翻译结果的可干预性. . . . . . . . . . . . . . . . . . . ....
View in text
Tags
AI categories
Artificial IntelligenceProgramming LanguageTechnology
Text Preview (First 20 pages)
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Generating text preview…
Loading comments...
Reply to Comment
Edit Comment