Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Lewis Tunstall, Leandro von Werra, Thomas Wolf

Rating No ratings yet

Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving state-of-the-art results on a variety of natural language processing tasks. If you're a data scientist or coder, this practical book -now revised in full color- shows you how to train and scale these large models using Hugging Face Transformers, a Python-based deep learning library. Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes. In this guide, authors Lewis Tunstall, Leandro von Werra, and Thomas Wolf, among the creators of Hugging Face Transformers, use a hands-on approach to teach you how transformers work and how to integrate them in your applications. You'll quickly learn a variety of tasks they can help you solve. Build, debug, and optimize transformer models for core NLP tasks, such as text classification, named entity recognition, and question answering Learn how transformers can be used for cross-lingual transfer learning Apply transformers in real-world scenarios where labeled data is scarce Make transformer models efficient for deployment using techniques such as distillation, pruning, and quantization Train transformers from scratch and learn how to scale to multiple GPUs and distributed environments

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

Brief outline
【One-Line Pitch】 Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving state-of-t… 【Book Arc】 - **Opening (~0%–12%)**: Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes.; Multilingual Named Entity Recognition. - **Early (~12%–35%)**: For example, Figure 1-5 visual‐ izes the attention weights for an English to French translation model, where each pixel denotes a weight.; but sometimes we would like to ask more targeted questions. - **Middle (~35%–65%)**: ed(model_ckpt, num_labels num_labels) .to(device)) You will see a warning that some parts of the model are randomly initialized.; 60 Chapter 3: Transformer Anatomy And that’s it—we’ve gone through all the steps to implement a simplified form of self- attention! - **Late (~65%–88%)**: Now, we can load the model weights as usual with the from_pretrained() function with the additional config argument.; he new errors until we were satisfied with the performance. - **Ending (~88%–100%)**: r this reason, it is the preferred metric for benchmarking.; Amanda: Lemme check Hannah: file_gif Amanda: Sorry, can't find it. 【Key Takeaways】 - **Transformers have been…** (Opening): Transformers have been used to write realistic news stories, improve Google Search queries, and even create chatbots that tell corny jokes. - **Multilingual Named Ent…** (Opening): Multilingual Named Entity Recognition. - **It covers everything f…** (Opening): It covers everything from the Transformer architecture itself, to the Transformers library and the entire ecosystem around it. - **For example** (Early): For example, Figure 1-5 visual‐ izes the attention weights for an English to French translation model, where each pixel denotes a weight. - **but sometimes we would…** (Early): but sometimes we would like to ask more targeted questions. - **character in the prece…** (Early): character in the preceding shell command, that’s because we’re running the commands in a Jupyter notebook. 【Reading Tips】 - Use Passage locations below to jump into the text and set reading anchors - If this is a brief outline, click Regenerate (top right) for a synthesized guide 【Coverage Limits】 Compressed outline without the model (~33 index chunks). Full structured guide needs AI available.
Excerpt 1
28 From Text to Tokens 29 Character Tokenization 29 Word Tokenization 31 Subword Tokenization 33 Tokenizing the Whole Dataset 35 Training a Text Classifier 3...
View in text
Excerpt 2
but sometimes we would like to ask more targeted questions. This is where we can use question answering. Question Answering In question answering, we provide...
View in text
Excerpt 3
ome of the input structure? There is: subword tokenization. 4 GPT-2 is the successor of GPT, and it captivated the public’s attention with its impressive abi...
View in text
Excerpt 4
ojections, each one representing a so-called attention head. The resulting multi-head attention layer is illustrated in Figure 3-5. But why do we need more t...
View in text
Excerpt 5
Name Tagging and Linking for 282 Languages,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics 1 (July 2017): 1946–1958...
View in text
Excerpt 6
erman model fares on the whole French test set by writing a simple function that encodes a dataset and generates the classification report on it: def evaluat...
View in text
Excerpt 7
ummary(text): return "\n".join(sent_tokenize(text)[:3]) summaries["baseline"] = three_sentence_summary(sample_text) Text Summarization Pipelines | 143 By tak...
View in text
Excerpt 8
ons for which no answer is possible, like the empty answers.answer_start examples in SubjQA. In these cases the model will assign a high start and end score...
View in text
Tags
AI categories
Python
ISBN: 1098136799
Publisher: O'Reilly Media
Publish Year: 2022
Language: English
Pages: 409
File Format: PDF
File Size: 17.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…