Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving state-of-the-art results on a variety of natural language processing tasks. If you're a data scientist or coder, this practical book -now revised in full color- shows you how to train and scale these large models using Hugging Face Transformers, a Python-based deep learning library. Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes. In this guide, authors Lewis Tunstall, Leandro von Werra, and Thomas Wolf, among the creators of Hugging Face Transformers, use a hands-on approach to teach you how transformers work and how to integrate them in your applications. You'll quickly learn a variety of tasks they can help you solve. Build, debug, and optimize transformer models for core NLP tasks, such as text classification, named entity recognition, and question answering Learn how transformers can be used for cross-lingual transfer learning Apply transformers in real-world scenarios where labeled data is scarce Make transformer models efficient for deployment using techniques such as distillation, pruning, and quantization Train transformers from scratch and learn how to scale to multiple GPUs and distributed environments
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
Brief outline
【One-Line Pitch】
Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving state-of-t…
【Book Arc】
- **Opening (~0%–12%)**: Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes.; Multilingual Named Entity Recognition.
- **Early (~12%–35%)**: For example, Figure 1-5 visual‐ izes the attention weights for an English to French translation model, where each pixel denotes a weight.; but sometimes we would like to ask more targeted questions.
- **Middle (~35%–65%)**: ed(model_ckpt, num_labels num_labels) .to(device)) You will see a warning that some parts of the model are randomly initialized.; 60 Chapter 3: Transformer Anatomy And that’s it—we’ve gone through all the steps to implement a simplified form of self- attention!
- **Late (~65%–88%)**: Now, we can load the model weights as usual with the from_pretrained() function with the additional config argument.; he new errors until we were satisfied with the performance.
- **Ending (~88%–100%)**: r this reason, it is the preferred metric for benchmarking.; Amanda: Lemme check Hannah: file_gif Amanda: Sorry, can't find it.
【Key Takeaways】
- **Transformers have been…** (Opening): Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes.
- **Multilingual Named Ent…** (Opening): Multilingual Named Entity Recognition.
- **It covers everything f…** (Opening): It covers everything from the Transformer architecture itself, to the Transformers library and the entire ecosystem around it.
- **For example** (Early): For example, Figure 1-5 visual‐ izes the attention weights for an English to French translation model, where each pixel denotes a weight.
- **but sometimes we would…** (Early): but sometimes we would like to ask more targeted questions.
- **character in the prece…** (Early): character in the preceding shell command, that’s because we’re running the commands in a Jupyter notebook.
【Reading Tips】
- Use Passage locations below to jump into the text and set reading anchors
- If this is a brief outline, click Regenerate (top right) for a synthesized guide
【Coverage Limits】
Compressed outline without the model (~33 index chunks). Full structured guide needs AI available.
Excerpt 1
28 From Text to Tokens 29 Character Tokenization 29 Word Tokenization 31 Subword Tokenization 33 Tokenizing the Whole Dataset 35 Training a Text Classifier 3...
but sometimes we would like to ask more targeted questions. This is where we can use question answering. Question Answering In question answering, we provide...
ome of the input structure? There is: subword tokenization. 4 GPT-2 is the successor of GPT, and it captivated the public’s attention with its impressive abi...
ojections, each one representing a so-called attention head. The resulting multi-head attention layer is illustrated in Figure 3-5. But why do we need more t...
Name Tagging and Linking for 282 Languages,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics 1 (July 2017): 1946–1958...
erman model fares on the whole French test set by writing a simple function that encodes a dataset and generates the classification report on it: def evaluat...
ons for which no answer is possible, like the empty answers.answer_start examples in SubjQA. In these cases the model will assign a high start and end score...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Natural Language Processing with Transformers, Revised Edition (Lewis Tunstall, Leandro von Werra, Thomas Wolf)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Natural Language Processing with Transformers, Revised Edition (Lewis Tunstall, Leandro von Werra, Thomas Wolf)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment