Speed up common data science tasks with AI assistants like ChatGPT and Large Language Models (LLMs) from Anthropic, Cohere, Open AI, Google, Hugging Face, and more!
Data Analysis with LLMs teaches you to use the new generation of AI assistants and Large Language Models (LLMs) to aid and accelerate common data science tasks.
Learn how to use LLMs to:
Analyze text, tables, images, and audio files
Extract information from multi-modal data lakes
Classify, cluster, transform, and query multimodal data
Build natural language query interfaces over structured data sources
Use LangChain to build complex data analysis pipelines
Prompt engineering and model configuration
All practical, Data Analysis with LLMs takes you from your first prompts through advanced techniques like creating LLM-based agents for data analysis and fine-tuning existing models. You’ll learn how to extract data, build natural language query interfaces, and much more.
About the book
Data Analysis with LLMs shows you exactly how to integrate generative AI into your day-to-day work as a data scientist. In it, Cornell professor Immanuel Trummer guides you through a series of engaging projects that introduce OpenAI’s Python library, tools like LangChain and LlamaIndex, and LLMs from Anthropic, Cohere, and Hugging Face. As you go, you’ll use AI to query structured and unstructured data, analyze sound and images, and optimize the cost and quality of your data analysis process.
What's inside
Classify, cluster, transform, and query multimodal data
Build natural language query interfaces over structured data sources
Create LLM-based agents for autonomous data analysis
Prompt engineering and model configuration
About the reader
For data scientists and data analysts who know the basics of Python.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Data Analysis with LLMs: Text, Tables, Images and Sound
## 【One-Line Pitch】
A practical, project-driven guide for data scientists and analysts who want to harness ChatGPT and other large language models to accelerate everything from text classification to natural language database queries. If you know basic Python and want to move beyond toy demos into real data workflows, this book shows you the how, the costs, and the pitfalls.
## 【Book Arc】
- **Opening (~0%–9%)**: Introduces what LLMs can do for data analysis, explains core concepts like zero-shot vs. few-shot learning, and walks through the anatomy of a good prompt—task description, context, output format, and optional examples.
- **Early (~9%–19%)**: Covers prompt templates for text-to-SQL translation and introduces model configuration basics, including output restrictions and fine-tuning. Also discusses comparing models across providers using benchmarks like HELM.
- **Early (~19%–34%)**: A hands-on chapter using ChatGPT via the web interface to classify, summarize, extract from, and respond to product reviews—demonstrating that one conversation can chain multiple text-processing tasks.
- **Middle (~34%–44%)**: Transitions from manual ChatGPT use to programmatic access, showing how to translate natural language questions into SQL queries against a SQLite database (the "BananaDB" example) and how to handle query output quirks.
- **Middle (~44%–47%)**: Introduces the OpenAI Python library—listing available models, understanding token usage and pricing (e.g., GPT-4o costs $5/M input tokens, $15/M output tokens), and configuring termination conditions like max_tokens.
## 【Key Takeaways】
- **Prompt design is the core skill** (Opening): A good prompt includes the task description, relevant context, and explicit output format—plus optional few-shot examples. Getting this right determines whether the model returns usable results or noise.
- **Zero-shot vs. few-shot learning matters** (Opening): You can often get away with zero examples if the task is well-described, but adding a few labeled examples dramatically improves reliability for nuanced tasks like sentiment classification.
- **LLMs can translate natural language to SQL** (Early): Using a prompt template with placeholders for the database schema and user question, you can build natural language query interfaces over relational databases—a pattern that extends to graph databases with Cypher later in the book.
- **Model configuration controls cost and quality** (Early): Restricting output to a single token for classification tasks, limiting max_tokens, and choosing a cheaper model (e.g., GPT-3.5 Turbo instead of GPT-4o) can cut costs by 10x without sacrificing much quality.
- **One conversation can chain multiple tasks** (Early): ChatGPT can classify, extract, summarize, and even draft responses to the same review within a single session—but you must avoid clicking "New Chat" or the context is lost.
- **Token pricing is asymmetric and model-dependent** (Middle): Generating tokens costs roughly 3x more than reading them, and model choice dominates your bill. Always estimate token usage before processing large datasets.
- **LLMs take liberties with instructions** (Middle): When generating SQL, ChatGPT may add extra columns or expand queries beyond what you asked. Always verify output against your requirements, especially if the query feeds into another tool.
## 【Reading Tips】
- **Skim the opening chapters** (~0%–9%) if you already know what prompts are; the key insight is the prompt template structure, which you'll reuse throughout.
- **Deep-read the ChatGPT chapter** (~19%–34%): This is where the book's hands-on philosophy shines. Follow along with the BananaBook review example to internalize how classification, extraction, and summarization differ in practice.
- **Pay attention to the SQL translation section** (~34%–44%): The "BananaDB" Colab notebook example is the template for all later structured data work. Note the triple-quote trick for multiline queries and the warning about query expansion.
- **The OpenAI library chapter** (~44%–47%) is essential for scaling beyond the web interface—read it carefully to understand token accounting and pricing before you process hundreds of reviews.
- **Watch for warnings**: The book flags when human oversight is needed (e.g., auto-replying to customers) and when models may hallucinate or expand output unexpectedly.
## 【Coverage Limits】
The excerpts cover roughly the first half of the book (through ~47%), focusing on text analysis, prompt design, SQL translation, and the OpenAI Python library. Later chapters on images, audio, LangChain, agents, and fine-tuning are mentioned but not covered in this guide.
##
Excerpt 1
of Python. Text, tables, images and sound Immanuel Trummer M A N N I N G To my beloved family CONTENTS vii 5.4 A natural language query interface for graph d...
abase The database contains the results of a survey, stored in a table called 'SurveyResults' with the following columns: ... Question: ② Question to transla...
of the different text-processing tasks we discussed in this section would have required a specialized language model. The latest generation of language model...
ore precisely, we find values for the following properties: completion_tokens—The number of generated tokens prompt_tokens—The number of tokens in the input...
For instance, consider the following prompt as an example: This movie is a piece of reality very well realized ... ① Review Is the sentiment positive or nega...
nr_clusters) df.to_csv('result.csv') 5.1 Chapter outline 77 Using this language, users can often perform a wide range of analysis operations on structured da...
sql(.*)‘‘‘', answer, re.DOTALL)[0] print(f'SQL: {query}') 5.3 A general natural language query interface 89 we get rid of unnecessary delimiters between rows...
atively multimodal model; we can use it for all these tasks. First, we will see how to use GPT-4o to answer free-form questions (in natural language) about i...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Data Analysis with LLMs Text, tables, images and sound (Immanuel Trummer)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Data Analysis with LLMs Text, tables, images and sound (Immanuel Trummer)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment