Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Rick J. Scavetta, Boyan Angelov

Rating No ratings yet

Success in data science depends on the flexible and appropriate use of tools. That includes Python and R, two of the foundational programming languages in the field. This book guides data scientists from the Python and R communities along the path to becoming bilingual. By recognizing the strengths of both languages, you'll discover new ways to accomplish data science tasks and expand your skill set. Authors Rick Scavetta and Boyan Angelov explain the parallel structures of these languages and highlight where each one excels, whether it's their linguistic features or the powers of their open source ecosystems. You'll learn how to use Python and R together in real-world settings and broaden your job opportunities as a bilingual data scientist. • Learn Python and R from the perspective of your current language • Understand the strengths and weaknesses of each language • Identify use cases where one language is better suited than the other • Understand the modern open source ecosystem available for both, including packages, frameworks, and workflows • Learn how to integrate R and Python in a single workflow • Follow a case study that demonstrates ways to use these languages together

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical guide for data scientists who already know one of Python or R and want to become genuinely bilingual, learning each language on its own terms and then combining them in real workflows. Best for working analysts, researchers, and students who feel limited by a single-language toolkit. 【Book Arc】 - **Opening (~0%–10%)**: Frames the bilingual mindset and the book's structure, then traces the parallel histories of Python and R—from Bell Labs and S to the PyData stack and the "language wars"—so you understand why each ecosystem looks the way it does. - **Early (~10%–35%)**: Teaches each language from the other's perspective, starting with R for Pythonistas: core syntax, data structures, projects and packages, and the Tidyverse versus base R distinction. - **Middle (~35%–60%)**: Continues the cross-language mapping, covering R's naming quirks, lists and model output, indexing, and the practical mechanics of loading packages and importing data. - **Late (~60%–85%)**: Moves into the modern context—data formats (image, text, time series, spatial), the open source package ecosystem, and workflow-specific choices across EDA, visualization, machine learning, data engineering, and reporting. - **Ending (~85%–100%)**: Explores the interfaces that let Python and R run together in a single workflow, culminating in a case study that demonstrates bilingual integration in practice. 【Key Takeaways】 - **Bilingualism is a mindset, not just syntax** (Early): The book explicitly warns against two traps—declaring your native language "better" and translating word-for-word. Different languages enable different constructs, and that's the point. - **R's history explains its personality** (Early): R descends from S at Bell Labs and was built "For Us, By Us" by statisticians, which is why its idioms and community feel distinct from Python's general-purpose computing roots. - **Python's data stack was assembled in layers** (Early): NumPy (2005) and SciPy enabled pandas (2009), forming the PyData stack; IPython and later Jupyter made notebooks the dominant Python data-science interface. - **Package management differs meaningfully** (Middle): Pythonistas tend to use virtual environments, while useRs historically install system-wide; `renv` is the current R answer to project-specific libraries. - **R's naming and loading rules have real consequences** (Middle): Base R silently rewrites illegal column names (replacing characters with `.` and prefixing numbers with `X`), while Tidyverse functions preserve them—a subtle source of bugs when reading others' scripts. - **Lists are R's natural container for heterogeneous results** (Middle): Statistical test and model outputs (coefficients, residuals) are stored as named lists, which is why `$` access and indexing patterns matter so much. - **Choose languages by task, not loyalty** (Late): The book maps data formats and workflows—image, text, time series, spatial, ML, engineering, reporting—to the language and packages best suited to each. - **Integration is the payoff** (Ending): The final stage moves from separate scripts to single workflows that weave Python and R together, demonstrated through a case study. 【Reading Tips】 - If you're a Pythonista, deep-read the R chapters (Early–Middle) and skim the Python recap; reverse this if you're a useR. - Treat the history chapter as context, not trivia—it explains why the ecosystems diverge and helps you predict where each language will feel natural. - Pay close attention to the package-management and naming sections; these are the practical friction points that trip up newcomers in daily work. - Use the workflow and data-format chapters as a decision reference you'll return to when starting a new project. - Don't skip the integration case study—it's where the book's "best of both worlds" promise becomes concrete. 【Coverage Limits】 The excerpts cover the book's framing, history, and early-to-middle R material in detail, but the later workflow, ecosystem, and integration chapters are represented mainly by tables of contents and brief mentions. Specific code examples, benchmarks, and the full case study are not covered here.
Page 4
. For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Michelle Smith Inde...
View in text
Page 14
ir ist heiß and ich bin heiß are both “I’m hot” in English, but a German speaker will distinguish hotness from the weather versus physique. Other words like...
View in text
Excerpt 3
y. Thus, it’s useful to mention what we’re not going to do. First, we’re not writing for the naïve data scientist. If you want to learn R from scratch, there...
View in text
Excerpt 4
tomic vectors, but we have some other tricks up our sleeve, of course—see indexing with [] in “How to Find…Stuff ” on page 32.) We didn’t get into the detail...
View in text
Excerpt 5
, you’re probably familiar with the python map() func‐ tion. An analogous function can be found in the Tidyverse purrr package. This is 4 If you want to read...
View in text
Excerpt 6
list. We actually already saw that when we defined the dict earlier: {'weight':['mean','std']} So both [] and {} alone are valid in Python and behave differe...
View in text
Excerpt 7
t knowing whether those methods exist. If the tools you use are designed well (as in the better design in Figure 4-3), often they will work as expected! Perf...
View in text
Excerpt 8
ow, torch Data engineeringb Flask, BentoML, FastAPI plumber Reporting Jupyter, Streamlit R Markdown, Shiny a Data munging (or wrangling) is such a fundamenta...
View in text
Tags
AI categories
PythonDataProgramming Language
ISBN: 1492093408
Publisher: O'Reilly Media
Publish Year: 2021
Language: English
Pages: 198
File Format: PDF
File Size: 18.2 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…