CS 7643 Deep Learning — Full
Practice Exam with questions with
answers in bold and rationale
updated 2026 graded A+
Section 1: Background and Fundamentals (Q1–Q15)
Q1. What distinguishes deep learning from traditional machine learning?
A. Deep learning requires labeled data exclusively
B. Deep learning automatically learns hierarchical feature representations from raw
data
C. Deep learning uses only linear models
D. Deep learning cannot be applied to image data
Correct Answer: B
Rationale: Deep learning's defining characteristic is its ability to learn hierarchical
feature representations directly from raw data, whereas traditional machine learning
often relies on manually engineered features. This automatic feature learning is what
enables deep networks to excel on unstructured data like images, text, and audio -7-
15 .
Q2. Which of the following is NOT a typical component of a deep learning
model?
A. Linear layers
B. Convolutional layers
C. Hard-coded decision trees
D. Pooling layers
Correct Answer: C
Rationale: Deep learning models are composed of differentiable modules like linear
layers, convolutional layers, and pooling layers that are trained end-to-end via
gradient descent. Hard-coded decision trees are a symbolic, non-differentiable
traditional ML component and are not part of neural network architectures -7 .
,Q3. What best defines a "parametric model"?
A. A model whose structure is fixed and whose parameters are learned from data
B. A model that stores all training examples for prediction
C. A model that does not require training
D. A model that can only classify two classes
Correct Answer: A
Rationale: Parametric models have a fixed number of parameters that are learned
from data during training. Neural networks are parametric models. In contrast, non-
parametric models like k-nearest neighbors grow with the training set size and store
training data -7-15 .
Q4. Which programming language is primarily used for hands-on assignments
in CS 7643?
A. Java
B. C++
C. Python
D. R
Correct Answer: C
Rationale: Python is used throughout the course, first with NumPy for foundational
implementations of neural network components, and later with PyTorch for building
and training deep learning models efficiently -7 .
Q5. What is the role of PyTorch in CS 7643?
A. It is used only for data visualization
B. It is a deep learning framework used to implement neural networks after initial
NumPy exercises
C. It replaces all mathematical calculations with symbolic algebra
D. It is not used in the course
Correct Answer: B
,Rationale: Students first implement core operations like forward and backward
passes in NumPy to understand the mathematics, then transition to PyTorch to build
and train neural networks more efficiently with GPU acceleration and automatic
differentiation -7 .
Q6. Which of the following is a prerequisite for CS 7643 at Georgia Tech?
A. Introduction to Psychology
B. Linear algebra, calculus (partial derivatives), probability/statistics, and an
introductory machine learning course
C. Web development experience
D. No prerequisites are required
Correct Answer: B
Rationale: The course requires a strong mathematical background including linear
algebra, multivariate calculus (especially partial derivatives), probability and statistics,
and at least an introductory machine learning course. Self-study does not satisfy the
ML prerequisite -6-7 .
Q7. What is the primary focus of CS 7643?
A. Theoretical foundations only
B. Practical implementations only
C. Both theoretical foundations and practical implementations
D. History of artificial intelligence
Correct Answer: C
Rationale: CS 7643 covers both theoretical foundations (mathematics of
backpropagation, optimization principles) and practical implementations through
programming assignments and a final project applying deep learning to real-world
problems -12 .
Q8. Which dataset is used in Homework 1 to train a ConvNet?
, A. ImageNet
B. CIFAR-10
C. MNIST
D. COCO
Correct Answer: B
Rationale: Homework 1 in CS 7643 involves training a Convolutional Neural Network
on the CIFAR-10 dataset, which consists of 60,000 32×32 color images in 10 classes -
12 .
Q9. What does a computation graph represent?
A. A visual representation of the training data
B. A directed graph where nodes represent operations and edges represent data flow
C. A graph showing the accuracy of the model over time
D. A graph of the network architecture only
Correct Answer: B
Rationale: A computation graph is a directed graph where nodes represent
operations (matrix multiplication, activation functions, etc.) and edges represent the
flow of tensors (data). It enables efficient gradient computation via backpropagation
by applying the chain rule in reverse topological order -12 .
Q10. What is automatic differentiation?
A. A method for automatically generating training data
B. A technique for automatically computing derivatives of functions defined by
computer programs
C. A method for automatically selecting hyperparameters
D. A technique for automatically labeling data
Correct Answer: B
Rationale: Automatic differentiation (autodiff) is a technique that computes exact
derivatives of functions defined by computer programs by systematically applying
the chain rule to elementary operations. It is the engine behind backpropagation in
modern frameworks like PyTorch -12 .
Practice Exam with questions with
answers in bold and rationale
updated 2026 graded A+
Section 1: Background and Fundamentals (Q1–Q15)
Q1. What distinguishes deep learning from traditional machine learning?
A. Deep learning requires labeled data exclusively
B. Deep learning automatically learns hierarchical feature representations from raw
data
C. Deep learning uses only linear models
D. Deep learning cannot be applied to image data
Correct Answer: B
Rationale: Deep learning's defining characteristic is its ability to learn hierarchical
feature representations directly from raw data, whereas traditional machine learning
often relies on manually engineered features. This automatic feature learning is what
enables deep networks to excel on unstructured data like images, text, and audio -7-
15 .
Q2. Which of the following is NOT a typical component of a deep learning
model?
A. Linear layers
B. Convolutional layers
C. Hard-coded decision trees
D. Pooling layers
Correct Answer: C
Rationale: Deep learning models are composed of differentiable modules like linear
layers, convolutional layers, and pooling layers that are trained end-to-end via
gradient descent. Hard-coded decision trees are a symbolic, non-differentiable
traditional ML component and are not part of neural network architectures -7 .
,Q3. What best defines a "parametric model"?
A. A model whose structure is fixed and whose parameters are learned from data
B. A model that stores all training examples for prediction
C. A model that does not require training
D. A model that can only classify two classes
Correct Answer: A
Rationale: Parametric models have a fixed number of parameters that are learned
from data during training. Neural networks are parametric models. In contrast, non-
parametric models like k-nearest neighbors grow with the training set size and store
training data -7-15 .
Q4. Which programming language is primarily used for hands-on assignments
in CS 7643?
A. Java
B. C++
C. Python
D. R
Correct Answer: C
Rationale: Python is used throughout the course, first with NumPy for foundational
implementations of neural network components, and later with PyTorch for building
and training deep learning models efficiently -7 .
Q5. What is the role of PyTorch in CS 7643?
A. It is used only for data visualization
B. It is a deep learning framework used to implement neural networks after initial
NumPy exercises
C. It replaces all mathematical calculations with symbolic algebra
D. It is not used in the course
Correct Answer: B
,Rationale: Students first implement core operations like forward and backward
passes in NumPy to understand the mathematics, then transition to PyTorch to build
and train neural networks more efficiently with GPU acceleration and automatic
differentiation -7 .
Q6. Which of the following is a prerequisite for CS 7643 at Georgia Tech?
A. Introduction to Psychology
B. Linear algebra, calculus (partial derivatives), probability/statistics, and an
introductory machine learning course
C. Web development experience
D. No prerequisites are required
Correct Answer: B
Rationale: The course requires a strong mathematical background including linear
algebra, multivariate calculus (especially partial derivatives), probability and statistics,
and at least an introductory machine learning course. Self-study does not satisfy the
ML prerequisite -6-7 .
Q7. What is the primary focus of CS 7643?
A. Theoretical foundations only
B. Practical implementations only
C. Both theoretical foundations and practical implementations
D. History of artificial intelligence
Correct Answer: C
Rationale: CS 7643 covers both theoretical foundations (mathematics of
backpropagation, optimization principles) and practical implementations through
programming assignments and a final project applying deep learning to real-world
problems -12 .
Q8. Which dataset is used in Homework 1 to train a ConvNet?
, A. ImageNet
B. CIFAR-10
C. MNIST
D. COCO
Correct Answer: B
Rationale: Homework 1 in CS 7643 involves training a Convolutional Neural Network
on the CIFAR-10 dataset, which consists of 60,000 32×32 color images in 10 classes -
12 .
Q9. What does a computation graph represent?
A. A visual representation of the training data
B. A directed graph where nodes represent operations and edges represent data flow
C. A graph showing the accuracy of the model over time
D. A graph of the network architecture only
Correct Answer: B
Rationale: A computation graph is a directed graph where nodes represent
operations (matrix multiplication, activation functions, etc.) and edges represent the
flow of tensors (data). It enables efficient gradient computation via backpropagation
by applying the chain rule in reverse topological order -12 .
Q10. What is automatic differentiation?
A. A method for automatically generating training data
B. A technique for automatically computing derivatives of functions defined by
computer programs
C. A method for automatically selecting hyperparameters
D. A technique for automatically labeling data
Correct Answer: B
Rationale: Automatic differentiation (autodiff) is a technique that computes exact
derivatives of functions defined by computer programs by systematically applying
the chain rule to elementary operations. It is the engine behind backpropagation in
modern frameworks like PyTorch -12 .