Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Immanuel Trummer

Rating No ratings yet

Speed up common data science tasks with AI assistants like ChatGPT and Large Language Models (LLMs) from Anthropic, Cohere, Open AI, Google, Hugging Face, and more! Data Analysis with LLMs teaches you to use the new generation of AI assistants and Large Language Models (LLMs) to aid and accelerate common data science tasks. Learn how to use LLMs to: Analyze text, tables, images, and audio files Extract information from multi-modal data lakes Classify, cluster, transform, and query multimodal data Build natural language query interfaces over structured data sources Use LangChain to build complex data analysis pipelines Prompt engineering and model configuration All practical, Data Analysis with LLMs takes you from your first prompts through advanced techniques like creating LLM-based agents for data analysis and fine-tuning existing models. You’ll learn how to extract data, build natural language query interfaces, and much more. About the book Data Analysis with LLMs shows you exactly how to integrate generative AI into your day-to-day work as a data scientist. In it, Cornell professor Immanuel Trummer guides you through a series of engaging projects that introduce OpenAI’s Python library, tools like LangChain and LlamaIndex, and LLMs from Anthropic, Cohere, and Hugging Face. As you go, you’ll use AI to query structured and unstructured data, analyze sound and images, and optimize the cost and quality of your data analysis process. What's inside Classify, cluster, transform, and query multimodal data Build natural language query interfaces over structured data sources Create LLM-based agents for autonomous data analysis Prompt engineering and model configuration About the reader For data scientists and data analysts who know the basics of Python.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Analysis with LLMs: Text, Tables, Images and Sound ## 【One-Line Pitch】 A practical, project-driven guide for data scientists and analysts who want to harness ChatGPT and other large language models to accelerate everything from text classification to natural language database queries. If you know basic Python and want to move beyond toy demos into real data workflows, this book shows you the how, the costs, and the pitfalls. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces what LLMs can do for data analysis, explains core concepts like zero-shot vs. few-shot learning, and walks through the anatomy of a good prompt—task description, context, output format, and optional examples. - **Early (~9%–19%)**: Covers prompt templates for text-to-SQL translation and introduces model configuration basics, including output restrictions and fine-tuning. Also discusses comparing models across providers using benchmarks like HELM. - **Early (~19%–34%)**: A hands-on chapter using ChatGPT via the web interface to classify, summarize, extract from, and respond to product reviews—demonstrating that one conversation can chain multiple text-processing tasks. - **Middle (~34%–44%)**: Transitions from manual ChatGPT use to programmatic access, showing how to translate natural language questions into SQL queries against a SQLite database (the "BananaDB" example) and how to handle query output quirks. - **Middle (~44%–47%)**: Introduces the OpenAI Python library—listing available models, understanding token usage and pricing (e.g., GPT-4o costs $5/M input tokens, $15/M output tokens), and configuring termination conditions like max_tokens. ## 【Key Takeaways】 - **Prompt design is the core skill** (Opening): A good prompt includes the task description, relevant context, and explicit output format—plus optional few-shot examples. Getting this right determines whether the model returns usable results or noise. - **Zero-shot vs. few-shot learning matters** (Opening): You can often get away with zero examples if the task is well-described, but adding a few labeled examples dramatically improves reliability for nuanced tasks like sentiment classification. - **LLMs can translate natural language to SQL** (Early): Using a prompt template with placeholders for the database schema and user question, you can build natural language query interfaces over relational databases—a pattern that extends to graph databases with Cypher later in the book. - **Model configuration controls cost and quality** (Early): Restricting output to a single token for classification tasks, limiting max_tokens, and choosing a cheaper model (e.g., GPT-3.5 Turbo instead of GPT-4o) can cut costs by 10x without sacrificing much quality. - **One conversation can chain multiple tasks** (Early): ChatGPT can classify, extract, summarize, and even draft responses to the same review within a single session—but you must avoid clicking "New Chat" or the context is lost. - **Token pricing is asymmetric and model-dependent** (Middle): Generating tokens costs roughly 3x more than reading them, and model choice dominates your bill. Always estimate token usage before processing large datasets. - **LLMs take liberties with instructions** (Middle): When generating SQL, ChatGPT may add extra columns or expand queries beyond what you asked. Always verify output against your requirements, especially if the query feeds into another tool. ## 【Reading Tips】 - **Skim the opening chapters** (~0%–9%) if you already know what prompts are; the key insight is the prompt template structure, which you'll reuse throughout. - **Deep-read the ChatGPT chapter** (~19%–34%): This is where the book's hands-on philosophy shines. Follow along with the BananaBook review example to internalize how classification, extraction, and summarization differ in practice. - **Pay attention to the SQL translation section** (~34%–44%): The "BananaDB" Colab notebook example is the template for all later structured data work. Note the triple-quote trick for multiline queries and the warning about query expansion. - **The OpenAI library chapter** (~44%–47%) is essential for scaling beyond the web interface—read it carefully to understand token accounting and pricing before you process hundreds of reviews. - **Watch for warnings**: The book flags when human oversight is needed (e.g., auto-replying to customers) and when models may hallucinate or expand output unexpectedly. ## 【Coverage Limits】 The excerpts cover roughly the first half of the book (through ~47%), focusing on text analysis, prompt design, SQL translation, and the OpenAI Python library. Later chapters on images, audio, LangChain, agents, and fine-tuning are mentioned but not covered in this guide. ##
Excerpt 1
of Python. Text, tables, images and sound Immanuel Trummer M A N N I N G To my beloved family CONTENTS vii 5.4 A natural language query interface for graph d...
View in text
Excerpt 2
abase The database contains the results of a survey, stored in a table called 'SurveyResults' with the following columns: ... Question: ② Question to transla...
View in text
Excerpt 3
of the different text-processing tasks we discussed in this section would have required a specialized language model. The latest generation of language model...
View in text
Excerpt 4
ore precisely, we find values for the following properties: completion_tokens—The number of generated tokens prompt_tokens—The number of tokens in the input...
View in text
Excerpt 5
For instance, consider the following prompt as an example: This movie is a piece of reality very well realized ... ① Review Is the sentiment positive or nega...
View in text
Excerpt 6
nr_clusters) df.to_csv('result.csv') 5.1 Chapter outline 77 Using this language, users can often perform a wide range of analysis operations on structured da...
View in text
Excerpt 7
sql(.*)‘‘‘', answer, re.DOTALL)[0] print(f'SQL: {query}') 5.3 A general natural language query interface 89 we get rid of unnecessary delimiters between rows...
View in text
Excerpt 8
atively multimodal model; we can use it for all these tasks. First, we will see how to use GPT-4o to answer free-form questions (in natural language) about i...
View in text
Tags
AI categories
Artificial IntelligencePythonData
ISBN: 1633437647
Publish Year: 2025
Language: English
Pages: 233
File Format: PDF
File Size: 23.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…