Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Hadley Wickham, Mine Çetinkaya-Rundel, Garrett Grolemund

Rating No ratings yet

Use R to turn data into insight, knowledge, and understanding. With this practical book, aspiring data scientists will learn how to do data science with R and RStudio, along with the tidyverseâ??a collection of R packages designed to work together to make data science fast, fluent, and fun. Even if you have no programming experience, this updated edition will have you doing data science quickly. You'll learn how to import, transform, and visualize your data and communicate the results. And you'll get a complete, big-picture understanding of the data science cycle and the basic tools you need to manage the details. Updated for the latest tidyverse features and best practices, new chapters show you how to get data from spreadsheets, databases, and websites. Exercises help you practice what you've learned along the way. You'll understand how to: • Visualize: Create plots for data exploration and communication of results • Transform: Discover variable types and the tools to work with them • Import: Get data into R and in a form convenient for analysis • Program: Learn R tools for solving data problems with greater clarity and ease • Communicate: Integrate prose, code, and results with Quarto

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to doing data science in R with the tidyverse, taking you from your first plot to a complete, reproducible analysis workflow. Best for aspiring data scientists and analysts who want practical skills, even with no prior programming experience. 【Book Arc】 - **Opening (~0%–10%)**: Sets up the whole data-science cycle and your toolkit—installing R/RStudio, loading the tidyverse, and reading the book's code conventions—so you can run code immediately. - **Early (~10%–30%)**: Builds the core loop of visualizing and transforming data: ggplot2 grammar of graphics, dplyr verbs, the pipe, groups, and tidy data pivoting with tidyr. - **Middle (~30%–50%)**: Moves into workflow craft—project organization, importing data from files, getting help, and layering plots with geoms, stats, and coordinate systems. - **Late (~50%–75%)**: Deepens the toolkit with programming tools (functions, iteration) and communication via Quarto, plus expanded import sources like spreadsheets, databases, and websites. - **Ending (~75%–100%)**: Rounds out the cycle with modeling and exploratory data analysis, then ties everything together into a reproducible, communicable result. 【Key Takeaways】 - **The data-science cycle is the book's spine** (Opening): import → tidy → transform → visualize → model → communicate, giving you a mental map before you touch syntax. - **Visualization is a grammar, not a gallery** (Early): mapping variables to aesthetics versus setting fixed values, and choosing geoms, is the core skill for exploration and communication. - **Tidy data is the prerequisite for everything downstream** (Early): pivot_longer() and pivot_wider() reshape messy real-world data into variables-in-columns, observations-in-rows form. - **dplyr plus the pipe makes transformation fluent** (Early): filter, group_by, summarize, and the native |> pipe let you express multi-step data manipulation readably. - **Workflow discipline separates beginners from experts** (Middle): RStudio Projects, working directories, and reproducible examples (reprex) prevent the chaos of scattered files and unrepeatable analyses. - **Import is broader than CSV** (Middle): read_csv() covers most files, but the book extends to spreadsheets, databases, and websites so real data sources are in reach. - **Programming tools reduce repetition** (Late): writing functions and iterating clarify data problems that copy-paste code obscures. - **Communication is part of the analysis** (Late): Quarto integrates prose, code, and results so findings are shareable, not just computed. 【Reading Tips】 - Deep-read the visualization and transformation chapters early; they are the foundation every later chapter assumes. - Skim the workflow chapters (code style, getting help, projects) on a first pass, then return when your own projects get messy. - Type and run the exercises rather than reading passively—the book is explicitly practice-driven, and errors teach the most. - Treat pivoting and the pipe as the two hardest early hurdles; slow down there before moving to modeling. - Use the book as a reference after finishing: jump back to import or Quarto chapters when a real project demands them. 【Coverage Limits】 The excerpts cover the book's structure, early visualization/transformation material, workflow, and pivoting, but do not detail the modeling, programming, or Quarto chapters in depth. Specific chapter titles and later examples are therefore summarized at a high level.
Excerpt 1
500 Summary 501 Part VI. Communicate 28. Quarto. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 2
aces. Comments R will ignore any text after # for that line. This allows you to write comments, text that is ignored by R but read by humans. We’ll sometimes...
View in text
Excerpt 3
ptonite 2000-04-08 81 70 68 67 66 #> 4 3 Doors Down Loser 2000-10-21 76 76 72 69 67 #> 5 504 Boyz Wobble Wobble 2000-04-15 57 34 25 17 17 #> 6 98^0 Give Me J...
View in text
Excerpt 4
= displ, y = hwy, shape = drv)) + geom_smooth() # Right ggplot(mpg, aes(x = displ, y = hwy, linetype = drv)) + geom_smooth() 122 | Chapter 9: Layers On the x...
View in text
Excerpt 5
geom_point() + x_scale + y_scale + col_scale # Right ggplot(compact, aes(x = displ, y = hwy, color = drv)) + geom_point() + x_scale + y_scale + col_scale Sca...
View in text
Excerpt 6
ff with some details of grouping components of the pattern. The terms we use here are the technical names for each component. They’re not always the most evo...
View in text
Excerpt 7
ymd_hms("2024-06-02 04:00:00", tz = "Pacific/Auckland") x3 #> [1] "2024-06-02 04:00:00 NZST" You can verify that they’re the same time using subtraction: x1...
View in text
Excerpt 8
talled automatically when you install the tidyverse package. Later, we’ll also use the writexl package, which allows us to create Excel spreadsheets. library...
View in text
Tags
AI categories
DataProgramming LanguageEducation
ISBN: 1492097403
Publisher: O'Reilly Media
Publish Year: 2023
Language: English
Pages: 579
File Format: PDF
File Size: 19.8 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…