Georgia Tech CS 7643 Deep Learning
— Full Practice Exam (100
Questions with Answers and
Rationales) updated 2026 graded
A+
Course: CS 7643 Deep Learning, Georgia Institute of Technology
Format: Multiple Choice
Coverage: Modules 1–9 from the official course syllabus -1
Module 1: Background and Fundamentals (Q1–Q10)
1. What distinguishes deep learning from traditional machine learning?
A. Deep learning requires labeled data exclusively
B. Deep learning automatically learns hierarchical feature representations from raw
data
C. Deep learning uses only linear models
D. Deep learning cannot be applied to image data
Correct Answer: B
Rationale: Deep learning focuses on learning complex, hierarchical feature
representations directly from raw data, whereas traditional machine learning often
relies on manually engineered features -3 .
2. Which of the following is a parametric model?
A. k-Nearest Neighbors
B. Logistic Classifier
C. Decision Tree (fully grown)
D. Kernel SVM
Correct Answer: B
Rationale: A parametric model has a fixed number of parameters learned from data.
A logistic classifier has weights and biases whose number is fixed regardless of
dataset size -8 .
,3. Which of the following is FALSE about parametric models?
A. The number of parameters is associated with the number of data points, not the
dimension of data features
B. Parameters are learned from training data
C. The model structure is fixed before training
D. Prediction is fast at test time
Correct Answer: A
Rationale: The number of parameters in a parametric model is associated with the
dimension of the data features, not the number of data points -8 .
4. The gradient of a big dataset can be approximated by the gradient of a small
subset of data.
A. True
B. False
Correct Answer: A
Rationale: This is the foundation of stochastic gradient descent (SGD). Using mini-
batches provides an unbiased estimate of the full gradient -8 .
5. Which of the following is NOT a typical component of a deep learning
model?
A. Linear layers
B. Convolutional layers
C. Hard-coded decision trees
D. Pooling layers
Correct Answer: C
Rationale: Deep learning models are composed of learnable layers (linear,
convolutional, pooling, etc.). Hard-coded decision trees are not a component of deep
neural networks -3 .
6. What is the role of PyTorch in CS 7643?
A. It is used only for data visualization
B. It is a deep learning framework used to implement neural networks after initial
NumPy exercises
C. It replaces all mathematical calculations with symbolic algebra
D. It is not used in the course
Correct Answer: B
Rationale: After implementing core operations in NumPy, students use PyTorch to
build and train neural networks more efficiently, leveraging GPU acceleration -3 .
, 7. Which programming language is primarily used for hands-on assignments in
CS 7643?
A. Java
B. C++
C. Python
D. R
Correct Answer: C
Rationale: Python is used throughout the course, first with NumPy for foundational
implementations and later with PyTorch for deep learning -3 .
8. For an image classification model with raw scores: cat = -1, dog = 0.5, mouse
= 1.0, and the correct label is "mouse," what is the cross-entropy loss (to 3
decimal places)?
A. 0.584
B. 1.346
C. 0.752
D. 2.108
Correct Answer: A
Rationale: Using softmax and cross-entropy, the loss for the correct class (mouse)
with scores [-1, 0.5, 1.0] computes to approximately 0.584 -8 .
9. What is the primary objective of representation learning in deep learning?
A. To manually design features for each task
B. To automatically discover useful features/representations from raw data
C. To reduce the number of parameters to zero
D. To avoid using neural networks
Correct Answer: B
Rationale: Representation learning aims to automatically discover useful
features/representations for a task from raw data, enabling end-to-end learning -9 .
10. A model with input → W1 → ReLU1 → W2 → ReLU2 → output is underfitting.
Which change is absolutely required?
A. Change ReLU2 to another function
B. Add softmax before the output
C. Increase model capacity
D. Add dropout
Correct Answer: C
Rationale: Underfitting indicates the model lacks capacity to capture the underlying
patterns. Increasing depth or width (capacity) is the primary remedy -8 .
— Full Practice Exam (100
Questions with Answers and
Rationales) updated 2026 graded
A+
Course: CS 7643 Deep Learning, Georgia Institute of Technology
Format: Multiple Choice
Coverage: Modules 1–9 from the official course syllabus -1
Module 1: Background and Fundamentals (Q1–Q10)
1. What distinguishes deep learning from traditional machine learning?
A. Deep learning requires labeled data exclusively
B. Deep learning automatically learns hierarchical feature representations from raw
data
C. Deep learning uses only linear models
D. Deep learning cannot be applied to image data
Correct Answer: B
Rationale: Deep learning focuses on learning complex, hierarchical feature
representations directly from raw data, whereas traditional machine learning often
relies on manually engineered features -3 .
2. Which of the following is a parametric model?
A. k-Nearest Neighbors
B. Logistic Classifier
C. Decision Tree (fully grown)
D. Kernel SVM
Correct Answer: B
Rationale: A parametric model has a fixed number of parameters learned from data.
A logistic classifier has weights and biases whose number is fixed regardless of
dataset size -8 .
,3. Which of the following is FALSE about parametric models?
A. The number of parameters is associated with the number of data points, not the
dimension of data features
B. Parameters are learned from training data
C. The model structure is fixed before training
D. Prediction is fast at test time
Correct Answer: A
Rationale: The number of parameters in a parametric model is associated with the
dimension of the data features, not the number of data points -8 .
4. The gradient of a big dataset can be approximated by the gradient of a small
subset of data.
A. True
B. False
Correct Answer: A
Rationale: This is the foundation of stochastic gradient descent (SGD). Using mini-
batches provides an unbiased estimate of the full gradient -8 .
5. Which of the following is NOT a typical component of a deep learning
model?
A. Linear layers
B. Convolutional layers
C. Hard-coded decision trees
D. Pooling layers
Correct Answer: C
Rationale: Deep learning models are composed of learnable layers (linear,
convolutional, pooling, etc.). Hard-coded decision trees are not a component of deep
neural networks -3 .
6. What is the role of PyTorch in CS 7643?
A. It is used only for data visualization
B. It is a deep learning framework used to implement neural networks after initial
NumPy exercises
C. It replaces all mathematical calculations with symbolic algebra
D. It is not used in the course
Correct Answer: B
Rationale: After implementing core operations in NumPy, students use PyTorch to
build and train neural networks more efficiently, leveraging GPU acceleration -3 .
, 7. Which programming language is primarily used for hands-on assignments in
CS 7643?
A. Java
B. C++
C. Python
D. R
Correct Answer: C
Rationale: Python is used throughout the course, first with NumPy for foundational
implementations and later with PyTorch for deep learning -3 .
8. For an image classification model with raw scores: cat = -1, dog = 0.5, mouse
= 1.0, and the correct label is "mouse," what is the cross-entropy loss (to 3
decimal places)?
A. 0.584
B. 1.346
C. 0.752
D. 2.108
Correct Answer: A
Rationale: Using softmax and cross-entropy, the loss for the correct class (mouse)
with scores [-1, 0.5, 1.0] computes to approximately 0.584 -8 .
9. What is the primary objective of representation learning in deep learning?
A. To manually design features for each task
B. To automatically discover useful features/representations from raw data
C. To reduce the number of parameters to zero
D. To avoid using neural networks
Correct Answer: B
Rationale: Representation learning aims to automatically discover useful
features/representations for a task from raw data, enabling end-to-end learning -9 .
10. A model with input → W1 → ReLU1 → W2 → ReLU2 → output is underfitting.
Which change is absolutely required?
A. Change ReLU2 to another function
B. Add softmax before the output
C. Increase model capacity
D. Add dropout
Correct Answer: C
Rationale: Underfitting indicates the model lacks capacity to capture the underlying
patterns. Increasing depth or width (capacity) is the primary remedy -8 .