Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Valliappa Lakshmanan

Rating No ratings yet

No description

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to building complete, real-time data pipelines on Google Cloud—from ingesting raw data to training and serving machine learning models—told through one running example. Best for data scientists and engineers who already know some Python and want to see how GCP's serverless and big-data tools fit together end to end. 【Book Arc】 - **Opening (~0%–10%)**: Frames the core premise—data science is about making the same decision systematically and repeatedly, at scale, rather than once. Introduces why the public cloud's elasticity and serverless model change the economics of data processing. - **Early (~10%–32%)**: Gets you operational: creating a GCP project, setting budgets, using Cloud Shell, and ingesting a real dataset (US flight data) into Cloud Storage and BigQuery. Covers the "why" behind Google's East-West data center design and why presharding isn't needed. - **Middle (~32%–52%)**: Moves into pipeline mechanics—scheduling monthly downloads with Cloud Run and Cloud Scheduler, loading data into BigQuery, and the principles of honest, interactive, explainable data visualization. - **Late (~52%–75%)**: Applies statistical and machine learning methods to the running decision problem, including Bayesian classification, logistic regression with Spark ML, feature engineering, and model evaluation. - **Ending (~75%–100%)**: Scales the workflow—dynamically resizing clusters, orchestration with Cloud Composer and workflow templates, serverless Spark, and autoscaling. (Excerpts thin out here; later chapters are only partially covered.) 【Key Takeaways】 - **The real value is systematic, repeatable decisions, not one-off analysis** (Early): The book's central thesis is that data science should produce a service that makes the same decision many times, not a single report. This shapes every architectural choice that follows. - **Serverless and elasticity change the cost calculus** (Early): Running 100 machines for one hour costs the same as 10 machines for ten hours—so you should get answers faster. BigQuery and Cloud Dataflow let you skip VM management entirely. - **Ingestion is a pipeline, not a download** (Early–Middle): The flight-data example walks through downloading, unzipping, gzipping, uploading to Cloud Storage, and loading into BigQuery—then automating it with Cloud Run and Cloud Scheduler. - **Bucket names are globally visible** (Early): A subtle but practical security note—choosing a bucket name can leak information about unreleased products to competitors. - **Explainability is a design requirement, not an afterthought** (Middle): Displays and models must be accurate, honest, interactive, and able to explain why a recommendation was made—especially to catch amplified bias. - **Spark ML provides a structured path from features to evaluation** (Late): Logistic regression, training datasets, prediction, and an experimental framework are covered as a coherent workflow rather than isolated techniques. - **Orchestration and autoscaling are what make it production-grade** (Ending): Workflow templates, Cloud Composer, and serverless Spark turn a working pipeline into something that runs reliably at scale. 【Reading Tips】 - **Deep-read the ingestion chapters (Early–Middle)**: The download-to-BigQuery pipeline is the book's backbone; understanding it makes later chapters much easier. - **Skim the GCP console setup details** if you're already familiar with projects, billing, and Cloud Shell—the conceptual points matter more than the click-by-click. - **Treat the visualization chapter as a design checklist**, not just a tools tour; the questions it poses about honesty and explainability apply to any dashboard you build. - **Follow the code, but adapt it**: The running flight-data example is meant to be executed; typing it out (or running the repo) will teach more than reading. - **Don't expect deep ML theory**: The book uses ML as a means to complete the pipeline, not as a statistics textbook. 【Coverage Limits】 The excerpts cover the book's structure and early-to-middle chapters well, but later chapters (orchestration, serverless Spark, and the final ML sections) are only partially represented. Specific code details and chapter titles beyond those shown are not fully covered here.
Excerpt 1
239 Autoscaling 239 Serverless Spark 240 Summary 242 Suggested Resources 243 7. Logistic Regression Using Spark ML. . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 2
is a one-off is the primary difference between data analyt‐ ics and data science. Data analytics is about manually analyzing data to make a single decision o...
View in text
Excerpt 3
of data analysis.16 Uploading Data to Google Cloud Storage For durability of this raw dataset, let’s upload it to Google Cloud Storage. To do that, you first...
View in text
Excerpt 4
f the Cloud Run deployment command in the previous section: echo {\"bucket\":\"${BUCKET}\"\} > /tmp/message cat /tmp/message gcloud scheduler jobs create htt...
View in text
Excerpt 5
ges over the dataset. Figure 3-17 shows the specifications. Note the Sort column at the end—it is important to have a reliable sort order in dash‐ boards so...
View in text
Excerpt 6
15-05-01 {"FL_DATE": "2015-04-30", "UNIQUE_CARRIER": "DL", 00:00:00 UTC "ORIGIN_AIRPORT_SEQ_ID": "1295302", "ORIGIN": "LGA", "DEST_AIRPORT_SEQ_ID": "1320402"...
View in text
Excerpt 7
ot;16 that is, it is an estimate of the probability distri‐ bution function (PDF).17 We see that, even though the distribution peaks around 10 minutes (which...
View in text
Excerpt 8
YTER --project $PROJECT \ --scopes https://www.googleapis.com/auth/cloud-platform A minute or so later, the Cloud Dataproc cluster is created, all ready to g...
View in text
Tags
AI categories
DataBig DataCloud Native
google cloud
ISBN: 1098118952
Publisher: O'Reilly Media
Publish Year: 2022
Language: English
Pages: 462
File Format: PDF
File Size: 17.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…