Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Marco Cremonini

Rating No ratings yet

Communicate the data that is powering our changing world with this essential text The advent of machine learning and neural networks in recent years, along with other technologies under the broader umbrella of ‘artificial intelligence,’ has produced an explosion in Data Science research and applications. Data Visualization, which combines the technical knowledge of how to work with data and the visual and communication skills required to present it, is an integral part of this subject. The expansion of Data Science is already leading to greater demand for new approaches to Data Visualization, a process that promises only to grow. Data Visualization in R and Python offers a thorough overview of the key dimensions of this subject. Beginning with the fundamentals of data visualization with Python and R, two key environments for data science, the book proceeds to lay out a range of tools for data visualization and their applications in web dashboards, data science environments, graphics, maps, and more. With an eye towards remarkable recent progress in open-source systems and tools, this book offers a cutting-edge introduction to this rapidly growing area of research and technological development. Data Visualization in R and Python readers will also find: Coverage suitable for anyone with a foundational knowledge of R and Python Detailed treatment of tools including the Ggplot2, Seaborn, and Altair libraries, Plotly/Dash, Shiny, and others Case studies accompanying each chapter, with full explanations for data operations and logic for each, based on Open Data from many different sources and of different formats Data Visualization in R and Python is ideal for any student or professional looking to understand the working principles of this key field.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Visualization in R and Python ## 【One-Line Pitch】 A practical, dual-language guide to modern data visualization that teaches you how to create compelling graphics using both R's ggplot2 and Python's Seaborn/Altair ecosystems, with real-world Open Data case studies. Ideal for students and professionals who want to move beyond basic plotting and master the full spectrum of visualization techniques. ## 【Book Arc】 - **Opening (~0%–9%)**: Establishes the book's philosophy—data visualization is a transversal discipline accessible to anyone working with data, not just statisticians or designers. Sets expectations: readers need fundamentals of data wrangling in R or Python, but the book provides complete code for all examples. - **Early (~9%–28%)**: Covers foundational plot types starting with scatterplots and line plots (Chapter 1), bar plots (Chapter 2), and facets (Chapter 3)—the grid-of-plots technique for comparing subgroups. Each chapter presents R (ggplot) and Python (Seaborn) implementations side-by-side using real datasets like World Bank inflation data and U.S. city temperatures. - **Early-to-Middle (~28%–38%)**: Advances to distribution-focused visualizations: histograms and kernel density estimates for univariate and bivariate analysis, diverging bar plots and lollipop plots for positive/negative data, and violin plots that combine boxplot statistics with density information. - **Middle (~38%–47%)**: Tackles the practical problem of overplotting—when too many data points obscure patterns. Presents jittering, sina plots, and beeswarm plots as solutions, then moves to raincloud plots (half-violin combinations) and ridgeline plots for comparing distributions across many categories. - **Late (~47% onward)**: Continues with advanced ordering techniques for categorical data (using factor levels to sort by descriptive statistics rather than alphabetically) and progresses toward more complex visualization types, including maps, web dashboards, and interactive applications using Plotly/Dash and Shiny. ## 【Key Takeaways】 - **Data visualization is a transversal skill** (Opening): You don't need a computer science or design background—anyone working with data in fields from economics to molecular biology can benefit. The book assumes only basic data wrangling knowledge in either R or Python. - **Learning one language transfers to the other** (Opening): If you know data operations in R or Python, understanding the other language's logic requires minimal effort—mostly syntax details. Knowing both is increasingly valuable in modern data science. - **Facets solve the multi-group comparison problem** (Early): Instead of manually subsetting data frames and managing multiple plots, facet visualization creates a grid of related plots in a single execution—each showing one unique value of a grouping variable. Caution needed: excessive facets become unreadable. - **Bivariate histograms reveal density patterns** (Early): Using functions like `histplot()` with two variables (e.g., number of reviews vs. price) shows where points cluster in two-dimensional space. The `discrete` attribute controls binning for categorical vs. continuous variables. - **Violin plots combine boxplot statistics with density** (Early): The length of tails in a violin plot corresponds to outlier distance in a boxplot, while the shape shows data density. Comparing violin and density plots side-by-side confirms they convey the same information differently. - **Overplotting has no universal solution** (Middle): Jittering, sina plots, and beeswarm plots each handle overlapping points differently. Sina and beeswarm plots convey additional distribution information beyond traditional jitter, but the choice depends on your data and message. - **Ordering categories by data, not alphabetically** (Middle): To create meaningful ridgeline plots, transform country names into factors and reorder levels using descriptive statistics (e.g., mean). This three-step technique—sort, factor, reorder—makes visual comparisons immediately intuitive. ## 【Reading Tips】 - **Skim the opening chapters if you know either R or Python**: The book's premise is that language transfer is easy—focus on the syntax differences rather than reading every example in both languages. - **Deep-read the overplotting chapter (Chapter 8)**: This is where practical data visualization gets tricky. The comparison between jitter, sina, and beeswarm approaches is essential for real-world messy data. - **Pay attention to dataset sources**: Each chapter uses Open Data (World Bank, OECD, U.S. city weather data) with clear citations—useful if you want to replicate examples or practice variations. - **Watch for ggplot-specific functions that lack Seaborn equivalents**: Lollipop plots, for instance, have efficient ggplot implementations but require matplotlib workarounds in Python—knowing these gaps helps you choose your primary tool. - **Study the ordering technique for categorical data**: The factor-level reordering method (Chapter 10) appears repeatedly in advanced visualizations and is a transferable skill beyond ridgeline plots. ## 【Coverage Limits】 Excerpts cover roughly the first half of the book (through ridgeline plots and categorical ordering). Later chapters on maps, web dashboards, Plotly/Dash, and Shiny applications are mentioned in the table of contents but not detailed in the available material. ##
Page 11
.3 Third Version: Tabs, Widgets, and Advanced Themes 286 15.4 Observe and Reactive 289 16 Advanced Shiny Dashboards 295 16.1 First Version: Sidebar, Widgets,...
View in text
Excerpt 2
the total quantity for each pollutant, we can start with a simple bar plot using function sns.barplot(), to which we add a few options: attribute order to or...
View in text
Excerpt 3
025 by John Wiley & Sons, Inc. Companion website: www.wiley.com/go/Cremonini/DataVisualization1e World: OECD-FAO Agricultural Outlook, year 2000 India 5081 R...
View in text
Excerpt 4
untries should be ordered based on the metric chosen (e.g., a descriptive statistic) and (b) the list of ordered countries should be created. 2. (a) Ordered...
View in text
Excerpt 5
alue circle could be seen, corresponding to Altair function mark_circle(), followed by local attributes opacity and size, then encoding and so on. It is the...
View in text
Excerpt 6
000 40,000 60,000 80,000 100,000 120,000 Asian Black Hawaii Ind/Nat Lat/Hisp Multiple Non Lat/Hisp Overall homeless White Year_Year 2022 (b) Figure 14.26 (Co...
View in text
Excerpt 7
that appears in server <- function(input, output, session). Session management has mostly to do with the man- agement of concurrent accesses from multiple us...
View in text
Excerpt 8
). Columns long and lat will be associated to the Cartesian axes x and y, while attribute group will be assigned to column group. Function geom_polygon() sup...
View in text
Tags
AI categories
Data VisualizationPython
ISBN: 1394289480
Publisher: John Wiley & Sons
Publish Year: 2025
Language: English
Pages: 578
File Format: PDF
File Size: 24.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…