Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Hobson Lane, Maria Dyshel

Rating No ratings yet

No description

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to building real NLP pipelines in Python, taking you from tokenization and vectorization through topic modeling and semantic search, with a candid eye on the ethics and incentives shaping today's language technologies. Best for developers and data practitioners who want working code and conceptual grounding rather than pure theory. 【Book Arc】 - **Opening (~0%–10%)**: Frames why NLP matters and why it is "magic" — machines handling human language rather than computerese — while confronting the AI ethics and "enshittification" problem of search and recommendation systems that optimize for investors, not users. - **Early (~10%–32%)**: Builds the foundational pipeline: tokenization (Python split, NLTK, spaCy, CoreNLP, Hugging Face WordPiece), one-hot vectors, BPE and character n-grams, stop words, stemming/normalization trade-offs, and the move from Counter dictionaries to Pandas vector representations with cosine similarity. - **Middle (~32%–48%)**: Shifts from counting to meaning: TF-IDF with log-scaled frequencies (Zipf's Law), topic vectors and semantic search, LDA topic modeling with train/test splits and confusion matrices, and dimensionality reduction via TruncatedSVD vs. PCA. - **Late (~48%+ excerpt coverage)**: Introduces steering and learned distance metrics — using feedback to shape embeddings and clustering so vectors minimize a cost function and focus on the information content you care about. (Excerpts do not cover chapters beyond this point in detail.) 【Key Takeaways】 - **NLP's real-world stakes are ethical, not just technical** (Opening): The book opens by arguing that search and recommendation NLP often serves investor incentives over user needs — a framing that recurs as a design consideration. - **Tokenization choices cascade through the whole pipeline** (Early): Whitespace splitting leaves punctuation attached and misses key terms; BPE and WordPiece handle languages without whitespace and even ciphers, so tokenizer selection is a foundational decision. - **Stop words carry relational meaning** (Early): Removing them can collapse distinctions like "reported to the CEO" vs. "reported as the CEO," forcing longer n-grams to preserve hierarchy — a concrete trade-off between efficiency and information. - **Normalization should be a fallback, not a default** (Early): Searching unstemmed tokens first and falling back to normalized matches (with transparent caveats) produces humbler, more useful chatbots; synonym substitution has four legitimate uses including data augmentation and adversarial testing. - **Vectors beat dictionaries for text** (Early): Pandas coerces Counter dictionaries into DataFrames with NaN handling, and cosine similarity (normalized dot product, range −1 to 1) measures how much documents point in the same semantic direction. - **TF-IDF needs log scaling because of Zipf's Law** (Middle): Raw frequencies make common words dominate exponentially; log-transforming term and document frequencies yields more uniform, useful scores. - **Topic vectors enable semantic search and keyword extraction** (Middle): Linear combinations of words represent meaning, letting you find documents by meaning even when users can't articulate the right query, and summarize documents with their most meaningful terms. - **Steering with feedback is the frontier** (Late): Learned distance metrics let you adjust similarity scores to clustering and embedding algorithms so vectors minimize a cost function — moving beyond unsupervised topic extraction toward goal-directed representations. 【Reading Tips】 - **Deep-read the tokenization and vectorization chapters** (Early): These are the load-bearing walls of every later pipeline; the code examples reward running in a notebook rather than skimming. - **Skim the ethics framing on first pass, return to it later**: The opening's AI ethics discussion is provocative but not prerequisite to the technical material — revisit it once you've built pipelines and can judge the claims. - **Watch the trade-off discussions, not just the recipes**: Stop words, stemming, and normalization are presented as judgment calls with consequences; the reasoning transfers to problems beyond the book's examples. - **Treat the MEAP status as a caveat**: As an early-access version, some chapters may be incomplete or reordered; verify against the final edition for anything you depend on. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book (through steering and learned distance metrics); later chapters on deep learning, transformers, and production deployment are not represented in the source material and are not summarized here.
Page 17
learn computerese. When software can process languages not designed for machines to understand, it is magic — something we thought only humans could do. More...
View in text
Excerpt 2
o the math or processing of NL*P* on natural language text. >>> import pandas as pd >>> onehot_vectors = np.zeros( ... (len(tokens), vocab_size), int) # #1 >...
View in text
Excerpt 3
ing up tokens in text. But vectors are where it’s really at. And it turns out that dictionaries can be coerced into a DataFrame or Series just by calling the...
View in text
Excerpt 4
’s how you do it with an sklearn function: >>> from sklearn.metrics import confusion_matrix >>> confusion_matrix(y_test, lda.predict(X_test)) array([[1261, 9...
View in text
Excerpt 5
within your software. For example, word embeddings make it possible to respond with flexibility by giving words fuzziness and nuance that previous representa...
View in text
Excerpt 6
ch you want to match. Your adjective-noun stencil has holes in the first row and the first column for the adjective at the beginning of a 2- gram. You will n...
View in text
Excerpt 7
r learned parameters there are, the greater the capacity of your model to learn more things about the data. But the whole point of all the clever ideas, like...
View in text
Excerpt 8
sequence always set as a special [CLS] token. Sentences are distinguished by a trailing separator token, [SEP]. Tokens in a sequence are further distinguishe...
View in text
Tags
AI categories
Artificial IntelligencePythonData
Publish Year: 2024
Language: English
Pages: 942
File Format: PDF
File Size: 12.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…