Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Mohamed Elgendy

Rating No ratings yet

Computer vision is central to many leading-edge innovations, including self-driving cars, drones, augmented reality, facial recognition, and much, much more. Amazing new computer vision applications are developed every day, thanks to rapid advances in AI and deep learning (DL). Deep Learning for Vision Systems teaches you the concepts and tools for building intelligent, scalable computer vision systems that can identify and react to objects in images, videos, and real life. With author Mohamed Elgendy's expert instruction and illustration of real-world projects, you’ll finally grok state-of-the-art deep learning techniques, so you can build, contribute to, and lead in the exciting realm of computer vision! About the technology How much has computer vision advanced? One ride in a Tesla is the only answer you’ll need. Deep learning techniques have led to exciting breakthroughs in facial recognition, interactive simulations, and medical imaging, but nothing beats seeing a car respond to real-world stimuli while speeding down the highway. About the book How does the computer learn to understand what it sees? Deep Learning for Vision Systems answers that by applying deep learning to computer vision. Using only high school algebra, this book illuminates the concepts behind visual intuition. You'll understand how to use deep learning architectures to build vision system applications for image generation and facial recognition.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, math-friendly guide that takes you from the pixel-level basics of computer vision to building and tuning deep learning models like CNNs and GANs, ideal for developers and data scientists who want to understand the "why" behind the algorithms rather than just applying them. 【Book Arc】 - **Opening (~0%–9%)**: Introduces computer vision (CV) as a field, contrasting human visual perception with machine vision. Establishes the core CV pipeline—image input, preprocessing, feature extraction, and model interpretation—and previews the major architectures (ANNs, CNNs, GANs, transfer learning) covered later. - **Early (~9%–25%)**: Dives into how computers "see" images as matrices of pixel values, covering grayscale vs. RGB color channels and the importance of feature selection. Uses a relatable example (distinguishing Greyhounds from Labradors) to show why raw pixels alone aren't enough for classification. - **Early (~25%–34%)**: Builds the foundation of neural networks, starting with single perceptrons and their limits on nonlinear problems. Explains why multilayer perceptrons (MLPs) are needed and walks through activation functions, emphasizing why linear functions fail and introducing nonlinear options like sigmoid. - **Middle (~34%–44%)**: Focuses on the learning process: defining error functions (why errors must be positive), comparing loss functions (mean squared error vs. cross-entropy), and the impracticality of brute-force optimization. Introduces gradient descent and its variants (batch, stochastic, mini-batch) as the solution. - **Middle (~44%–47%)**: Transitions to convolutional neural networks (CNNs), starting with the input layer and how images are preprocessed into tensors. Sets up the key difference from MLPs: convolutional layers that learn hierarchical features (edges → shapes → object parts) instead of fully connected layers. 【Key Takeaways】 - **Computer vision is about perception, not just pixels** (Early): A CV system needs both a sensing device and an interpreting device; the pipeline (input → preprocessing → feature extraction → model) is the backbone of all applications, from self-driving cars to facial recognition. - **Computers see images as numbers, not pictures** (Early): A grayscale image is a 2D matrix (0–255 per pixel), and a color image is a 3D matrix with RGB channels (e.g., 700×700×3). Understanding this representation is the first step to designing any vision model. - **Feature selection is a design choice, not an accident** (Early): Using the Greyhound vs. Labrador example, the book shows that features like height or eye color have varying discriminative power—some are noisy, some are useful—which motivates why deep learning learns features automatically. - **Single perceptrons fail on nonlinear problems** (Early): A single straight line can't separate complex data; you need multiple perceptrons (an MLP) to create decision boundaries. However, adding too many neurons risks overfitting, so architecture complexity must be balanced. - **Activation functions must be nonlinear** (Early): A linear transfer function (identity) makes deep networks useless—composing linear functions yields only linear results, and gradients stay constant, so no real learning occurs. Nonlinear functions like sigmoid are essential. - **Error functions must be positive to be meaningful** (Middle): If errors can be negative, they cancel out (e.g., +10 and –10 average to zero, falsely implying perfect predictions). Loss functions like MSE (regression) and cross-entropy (classification) ensure errors accumulate constructively. - **Brute-force optimization is impossible; gradient descent is the answer** (Middle): Even a simple network would take longer than the universe's age to train by trying all weight combinations. Gradient descent iteratively updates weights using the slope of the error curve, with variants (BGD, SGD, mini-batch) trading off speed vs. stability. - **CNNs learn features hierarchically** (Middle): Unlike MLPs that flatten images, CNNs use convolutional layers where the first layer detects edges/lines, the next detects shapes, and deeper layers detect complex parts (faces, wheels). This mirrors how the network "understands" images progressively. 【Reading Tips】 - **Skim the opening chapter (0–9%)** if you're already familiar with CV basics; the key takeaway is the four-step pipeline, which frames the rest of the book. Focus on the "computers see matrices" section (Early) if you're new to image data. - **Deep-read the perceptron and activation function sections (25–34%)**: These are the conceptual core. Work through the linear vs. nonlinear examples and the activation function cheat sheet (Table 2.1) carefully—they explain why deep learning works at all. - **Treat the optimization chapter (34–44%) as a reference**: The math on gradient descent variants can be dense. Skim the brute-force calculation (it's just motivation), but pay attention to the differences between BGD, SGD, and mini-batch—you'll choose among them in practice. - **The CNN section (44% onward) is where the book shifts to vision-specific content**: If you're here for CNNs specifically, start deep-reading from the 28×28 matrix example (Middle) and note how the architecture differs from MLPs; the hierarchical feature learning explanation is the "aha" moment. - **Don't skip the author's preface (0–6%)**: It explains the book's philosophy—learning to read research papers and think independently, not just memorizing facts. This mindset will help you when the math gets tough. 【Coverage Limits】 The excerpts cover roughly the first half of the book (up to ~47%), focusing on CV fundamentals, neural network basics, and the introduction to CNNs. Later chapters on GANs, transfer learning, and advanced applications (e.g., image generation, facial recognition) are mentioned but not detailed in this guide.
Page 9
te decay and adaptive learning 170 ■ Mini-batch size 171 4.7 Optimization algorithms 174 Gradient descent with momentum 174 ■ Adam 175 Number of epochs and e...
View in text
Excerpt 2
255 for white. The image in figure 1.16 is of size 24 × 24. This size indicates the width and height of the image: there are 24 pixels horizontally and 24 ve...
View in text
Excerpt 3
e–z function ity between 0 and 1, which reduces Ø(z) 1.5 extreme values or outliers in the data. Usually 0.0 used to classify –8 –6 –4 –2 0 2 4 6 8 two class...
View in text
Excerpt 4
0 0 0 0 0 0 0 0 0 14 200 254 134 0 0 0 0 0 0 0 0 0 28 × 28 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 11 246 253 59 0 0 0 0 0 0 0 0 0 = 784 pixels 0 0 0 17 11 18 0 0 0 0...
View in text
Excerpt 5
ron represents a certain feature that, when multiplied by a weight, is transformed into another feature. When we randomly turn off some of these nodes, we fo...
View in text
Excerpt 6
Train the model Evaluate the model Visualize the results Here are the steps: The scikit-learn library to generate sample data 1 Import the dependencies: Kera...
View in text
Excerpt 7
geNet dataset to com- pare algorithms. More on that later. NOTE The snippets in this chapter are not meant to be runnable. The goal is to show you how to imp...
View in text
Excerpt 8
uses a stack of a total of nine inception modules and a max pooling layer every several blocks to reduce dimensionality. To simplify this implementa- tion, w...
View in text
Tags
AI categories
deep learningArtificial Intelligence
ISBN: 1617296198
Publish Year: 2020
Language: English
Pages: 475
File Format: PDF
File Size: 16.8 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…