Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ryan Mitchell

Rating No ratings yet

If programming is magic, then web scraping is surely a form of wizardry. By writing a simple automated program, you can query web servers, request data, and parse it to extract the information you need. This thoroughly updated third edition not only introduces you to web scraping but also serves as a comprehensive guide to scraping almost every type of data from the modern web. Part I focuses on web scraping mechanics: using Python to request information from a web server, performing basic handling of the server's response, and interacting with sites in an automated fashion. Part II explores a variety of more specific tools and applications to fit any web scraping scenario you're likely to encounter. Parse complicated HTML pages Develop crawlers with the Scrapy framework Learn methods to store the data you scrape Read and extract data from documents Clean and normalize badly formatted data Read and write natural languages Crawl through forms and logins Scrape JavaScript and crawl through APIs Use and write image-to-text software Avoid scraping traps and bot blockers Use scrapers to test your website

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, end-to-end field guide to extracting data from the modern web with Python—covering everything from your first HTTP request to crawling JavaScript-heavy sites and dodging bot blockers. Best for developers, analysts, and data-minded tinkerers who know some Python and want to turn messy web pages into usable datasets. 【Book Arc】 - **Opening (~0%–10%)**: Frames web scraping as "wizardry" and sets the conceptual foundation—how the web's standards and browser behavior shape what a scraper must do, plus the legal and ethical terrain (robots.txt, trespass to chattels, terms of service). - **Early (~10%–32%)**: The mechanics of getting data—requesting pages with Python, handling HTTP errors gracefully, and parsing HTML with BeautifulSoup, including navigating tags, attributes, parents, and siblings, then building simple recursive crawlers to map or gather from a site. - **Middle (~32%–48%)**: Scaling up and structuring the work—designing scraper projects (broad vs. targeted), modeling scraped data into classes, using the Scrapy framework with link extractors and crawl rules, and storing results in databases like MySQL. - **Late (~48%–70%)**: Handling harder content—reading and extracting from documents, cleaning and normalizing badly formatted data, and working with natural-language text. - **Ending (~70%–100%)**: Advanced scenarios—crawling through forms and logins, scraping JavaScript and APIs, image-to-text software, avoiding scraping traps and bot blockers, and using scrapers to test your own site. (Excerpts do not cover the final chapters in detail.) 【Key Takeaways】 - **Scraping is applied web literacy** (Opening): You're substituting your own program for a browser, so understanding HTTP, HTML, and browser standards is the real prerequisite—not just Python syntax. - **Legal and ethical awareness is part of the craft** (Opening): robots.txt, terms of service, and laws like the CFAA and DMCA shape what's permissible; the book treats this as foundational, not an afterthought. - **Robust error handling separates toy scrapers from real ones** (Early): Anticipating HTTP errors, missing pages, and unexpected data formats prevents silent failures mid-run. - **BeautifulSoup's find_all is the workhorse of parsing** (Early): Mastering tag, attribute, text, and recursive parameters lets you drill from a page's top layer down to the exact data you need. - **Crawlers are built from link discovery** (Early): Distinguishing internal vs. external links and recursively traversing a site enables site maps, content gathering, and structural analysis. - **Plan the project before writing code** (Middle): Classifying scrapers as broad vs. targeted and modeling data into classes (or subclasses like Product/Article) determines your architecture and storage schema. - **Scrapy scales crawling with rules and extractors** (Middle): CrawlSpider, LinkExtractor, and CSS/XPath selectors handle complex, multi-page extraction and structured output (CSV, XML). - **The modern web demands more than static HTML** (Late/Ending): JavaScript rendering, APIs, logins, forms, and anti-bot defenses require progressively more advanced techniques—and the book points to external resources where topics outgrow a single chapter. 【Reading Tips】 - **Deep-read Part I, skim Part II selectively.** The first part is a reusable reference for core libraries and techniques; the second is a menu—jump to the chapters matching your actual scraping scenario. - **Type the code, don't just read it.** The book is explicitly hands-on; error handling and parsing behavior only click when you hit real exceptions. - **Treat the legal chapter as required, not optional.** It's short but shapes every project decision afterward. - **Use the jump-around structure deliberately.** Chapters build on each other, but the book is designed for targeted lookup—note cross-references as you go. - **Expect pointers, not exhaustive coverage.** For databases, image processing, and NLP, the book gets you started and defers to other resources. 【Coverage Limits】 This guide is synthesized from stratified excerpts covering roughly the first half of the book in detail; the later chapters on documents, NLP, JavaScript, APIs, and anti-bot techniques are represented mainly by the table of contents and brief mentions, so specifics there are not fully captured.
Page 10
d in enough detail to get you started writing web scrapers! Part I covers the subject of web scraping and web crawling in depth, with a strong focus on a sma...
View in text
Excerpt 2
and aren’t sure how to access its content via Python, rest assured that there’s probably another travel site with the exact same data that you can try. Sales...
View in text
Excerpt 3
the target site can be used here—it doesn’t need to be the exact URL of the BeautifulSoup object passed in. This function creates a set called internalLinks...
View in text
Excerpt 4
d> 7 February 2023, at 01:14.</lastUpdated> </item> In the JSON format, lists are preserved as lists. Of course, you can use the Item objects yourself and wr...
View in text
Excerpt 5
thon 2.5,” or code examples like: Cleaning Data with Pandas This section is not about the endearing bears native to China, but the Python data analysis packa...
View in text
Excerpt 6
llow other parts of speech. Whenever an ambiguous word such as “dust” is encountered, the rules of the context-free contents are being submitted to https://p...
View in text
Excerpt 7
.frame(frame) try:   Wait for preview reader to load     WebDriverWait(driver, 600).until(   EC.presence_of_element_located((By.ID, 'kr-renderer')) except Ti...
View in text
Excerpt 8
run. The %reset line is useful because it resets the memory and destroys all user-created variables in the Jupyter The demo page located at http://pythonscra...
View in text
Tags
AI categories
PythonWeb TechnologyData
ISBN: 1098145356
Publisher: O'Reilly Media
Publish Year: 2024
Language: English
Pages: 336
File Format: PDF
File Size: 11.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…