If programming is magic, then web scraping is surely a form of wizardry. By writing a simple automated program, you can query web servers, request data, and parse it to extract the information you need. This thoroughly updated third edition not only introduces you to web scraping but also serves as a comprehensive guide to scraping almost every type of data from the modern web.
Part I focuses on web scraping mechanics: using Python to request information from a web server, performing basic handling of the server's response, and interacting with sites in an automated fashion. Part II explores a variety of more specific tools and applications to fit any web scraping scenario you're likely to encounter.
Parse complicated HTML pages
Develop crawlers with the Scrapy framework
Learn methods to store the data you scrape
Read and extract data from documents
Clean and normalize badly formatted data
Read and write natural languages
Crawl through forms and logins
Scrape JavaScript and crawl through APIs
Use and write image-to-text software
Avoid scraping traps and bot blockers
Use scrapers to test your website
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical, end-to-end field guide to extracting data from the modern web with Python—covering everything from your first HTTP request to crawling JavaScript-heavy sites and dodging bot blockers. Best for developers, analysts, and data-minded tinkerers who know some Python and want to turn messy web pages into usable datasets.
【Book Arc】
- **Opening (~0%–10%)**: Frames web scraping as "wizardry" and sets the conceptual foundation—how the web's standards and browser behavior shape what a scraper must do, plus the legal and ethical terrain (robots.txt, trespass to chattels, terms of service).
- **Early (~10%–32%)**: The mechanics of getting data—requesting pages with Python, handling HTTP errors gracefully, and parsing HTML with BeautifulSoup, including navigating tags, attributes, parents, and siblings, then building simple recursive crawlers to map or gather from a site.
- **Middle (~32%–48%)**: Scaling up and structuring the work—designing scraper projects (broad vs. targeted), modeling scraped data into classes, using the Scrapy framework with link extractors and crawl rules, and storing results in databases like MySQL.
- **Late (~48%–70%)**: Handling harder content—reading and extracting from documents, cleaning and normalizing badly formatted data, and working with natural-language text.
- **Ending (~70%–100%)**: Advanced scenarios—crawling through forms and logins, scraping JavaScript and APIs, image-to-text software, avoiding scraping traps and bot blockers, and using scrapers to test your own site. (Excerpts do not cover the final chapters in detail.)
【Key Takeaways】
- **Scraping is applied web literacy** (Opening): You're substituting your own program for a browser, so understanding HTTP, HTML, and browser standards is the real prerequisite—not just Python syntax.
- **Legal and ethical awareness is part of the craft** (Opening): robots.txt, terms of service, and laws like the CFAA and DMCA shape what's permissible; the book treats this as foundational, not an afterthought.
- **Robust error handling separates toy scrapers from real ones** (Early): Anticipating HTTP errors, missing pages, and unexpected data formats prevents silent failures mid-run.
- **BeautifulSoup's find_all is the workhorse of parsing** (Early): Mastering tag, attribute, text, and recursive parameters lets you drill from a page's top layer down to the exact data you need.
- **Crawlers are built from link discovery** (Early): Distinguishing internal vs. external links and recursively traversing a site enables site maps, content gathering, and structural analysis.
- **Plan the project before writing code** (Middle): Classifying scrapers as broad vs. targeted and modeling data into classes (or subclasses like Product/Article) determines your architecture and storage schema.
- **Scrapy scales crawling with rules and extractors** (Middle): CrawlSpider, LinkExtractor, and CSS/XPath selectors handle complex, multi-page extraction and structured output (CSV, XML).
- **The modern web demands more than static HTML** (Late/Ending): JavaScript rendering, APIs, logins, forms, and anti-bot defenses require progressively more advanced techniques—and the book points to external resources where topics outgrow a single chapter.
【Reading Tips】
- **Deep-read Part I, skim Part II selectively.** The first part is a reusable reference for core libraries and techniques; the second is a menu—jump to the chapters matching your actual scraping scenario.
- **Type the code, don't just read it.** The book is explicitly hands-on; error handling and parsing behavior only click when you hit real exceptions.
- **Treat the legal chapter as required, not optional.** It's short but shapes every project decision afterward.
- **Use the jump-around structure deliberately.** Chapters build on each other, but the book is designed for targeted lookup—note cross-references as you go.
- **Expect pointers, not exhaustive coverage.** For databases, image processing, and NLP, the book gets you started and defers to other resources.
【Coverage Limits】
This guide is synthesized from stratified excerpts covering roughly the first half of the book in detail; the later chapters on documents, NLP, JavaScript, APIs, and anti-bot techniques are represented mainly by the table of contents and brief mentions, so specifics there are not fully captured.
Page 10
d in enough detail to get you started writing web scrapers! Part I covers the subject of web scraping and web crawling in depth, with a strong focus on a sma...
and aren’t sure how to access its content via Python, rest assured that there’s probably another travel site with the exact same data that you can try. Sales...
the target site can be used here—it doesn’t need to be the exact URL of the BeautifulSoup object passed in. This function creates a set called internalLinks...
d> 7 February 2023, at 01:14.</lastUpdated> </item> In the JSON format, lists are preserved as lists. Of course, you can use the Item objects yourself and wr...
thon 2.5,” or code examples like: Cleaning Data with Pandas This section is not about the endearing bears native to China, but the Python data analysis packa...
llow other parts of speech. Whenever an ambiguous word such as “dust” is encountered, the rules of the context-free contents are being submitted to https://p...
run. The %reset line is useful because it resets the memory and destroys all user-created variables in the Jupyter The demo page located at http://pythonscra...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Web Scraping with Python Data Extraction from the Modern Web (Ryan Mitchell)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Web Scraping with Python Data Extraction from the Modern Web (Ryan Mitchell)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment