CS7643 DEEP LEARNING QUIZ 2 TEST QUESTIONS AND
CORRECT ANSWERS (VERIFIED ANSWERS) PLUS
RATIONALES 2026 Q&A | INSTANT DOWNLOAD PDF.
CORE DOMAINS *
1. Neural Network Fundamentals *
2. Optimization Techniques *
3. Convolutional Neural Networks *
4. Recurrent Neural Networks and Sequence Modeling *
5. Regularization and Training Techniques *
6. Advanced Architectures and Transfer Learning *
7. Computer Vision Applications *
8. Reinforcement Learning *
9. Generative Models *
, 10. Attention and Transformers *
INTRODUCTION *
The CS7643 Deep Learning Quiz 2 Assessment evaluates
graduate-level understanding of neural network
architectures, optimization algorithms, and modern deep
learning techniques at Georgia Institute of Technology.
This comprehensive assessment covers fundamental
concepts including backpropagation, convolutional
operations, recurrent architectures, attention mechanisms,
and regularization strategies. The examination employs
multiple-choice and scenario-based questions requiring
candidates to demonstrate mathematical reasoning,
architectural analysis, and practical implementation
knowledge. Emphasis is placed on understanding
parameter counting, gradient flow, receptive fields, and
model design decisions. Successful performance validates
mastery of deep learning theory and its application to
real-world problems. *
SECTION ONE: QUESTIONS 1–100
Question 1
A neural network architect is designing a multi-layer
perceptron for a binary classification task. The final layer
,must output a probability score between 0 and 1. Which
activation function is most appropriate for the output layer?
A. ReLU
B. Tanh
🟢 C. Sigmoid
D. Softmax
🔴 RATIONALE: Sigmoid maps any real-valued input to the
interval (0,1), making it suitable for binary classification
probability outputs. Softmax is used for multi-class
classification, while ReLU and Tanh do not produce
probability scores.
Question 2
A data scientist observes that a deep neural network with
ten hidden layers is performing worse on training data than
a shallow network with two hidden layers. The training loss
plateaus early and gradients in early layers are near zero.
Which phenomenon best explains this behavior?
A. Overfitting due to excessive model capacity
B. Underfitting due to insufficient model depth
🟢 C. Vanishing gradients caused by sigmoid activation
functions
D. Exploding gradients caused by high learning rate
, 🔴 RATIONALE: Sigmoid activations saturate for large
positive or negative inputs, causing gradients to become
extremely small during backpropagation. This prevents early
layers from learning effectively, a problem known as
vanishing gradients.
Question 3
A researcher initializes a convolutional neural network using
Xavier initialization but observes that the variance of
activations decreases dramatically across layers. Which
modification is most likely to resolve this issue?
A. Switch to a smaller learning rate
🟢 B. Replace Xavier initialization with He initialization
C. Add more convolutional layers
D. Increase the batch size
🔴 RATIONALE: Xavier initialization is designed for zero-
centered activations like tanh and sigmoid. For ReLU-based
networks, He initialization scales weights by 2/n, which
better maintains variance across layers.
Question 4
A machine learning engineer trains a convolutional neural
network and notices that many neurons output zero for all
inputs after a few epochs. The network uses ReLU
CORRECT ANSWERS (VERIFIED ANSWERS) PLUS
RATIONALES 2026 Q&A | INSTANT DOWNLOAD PDF.
CORE DOMAINS *
1. Neural Network Fundamentals *
2. Optimization Techniques *
3. Convolutional Neural Networks *
4. Recurrent Neural Networks and Sequence Modeling *
5. Regularization and Training Techniques *
6. Advanced Architectures and Transfer Learning *
7. Computer Vision Applications *
8. Reinforcement Learning *
9. Generative Models *
, 10. Attention and Transformers *
INTRODUCTION *
The CS7643 Deep Learning Quiz 2 Assessment evaluates
graduate-level understanding of neural network
architectures, optimization algorithms, and modern deep
learning techniques at Georgia Institute of Technology.
This comprehensive assessment covers fundamental
concepts including backpropagation, convolutional
operations, recurrent architectures, attention mechanisms,
and regularization strategies. The examination employs
multiple-choice and scenario-based questions requiring
candidates to demonstrate mathematical reasoning,
architectural analysis, and practical implementation
knowledge. Emphasis is placed on understanding
parameter counting, gradient flow, receptive fields, and
model design decisions. Successful performance validates
mastery of deep learning theory and its application to
real-world problems. *
SECTION ONE: QUESTIONS 1–100
Question 1
A neural network architect is designing a multi-layer
perceptron for a binary classification task. The final layer
,must output a probability score between 0 and 1. Which
activation function is most appropriate for the output layer?
A. ReLU
B. Tanh
🟢 C. Sigmoid
D. Softmax
🔴 RATIONALE: Sigmoid maps any real-valued input to the
interval (0,1), making it suitable for binary classification
probability outputs. Softmax is used for multi-class
classification, while ReLU and Tanh do not produce
probability scores.
Question 2
A data scientist observes that a deep neural network with
ten hidden layers is performing worse on training data than
a shallow network with two hidden layers. The training loss
plateaus early and gradients in early layers are near zero.
Which phenomenon best explains this behavior?
A. Overfitting due to excessive model capacity
B. Underfitting due to insufficient model depth
🟢 C. Vanishing gradients caused by sigmoid activation
functions
D. Exploding gradients caused by high learning rate
, 🔴 RATIONALE: Sigmoid activations saturate for large
positive or negative inputs, causing gradients to become
extremely small during backpropagation. This prevents early
layers from learning effectively, a problem known as
vanishing gradients.
Question 3
A researcher initializes a convolutional neural network using
Xavier initialization but observes that the variance of
activations decreases dramatically across layers. Which
modification is most likely to resolve this issue?
A. Switch to a smaller learning rate
🟢 B. Replace Xavier initialization with He initialization
C. Add more convolutional layers
D. Increase the batch size
🔴 RATIONALE: Xavier initialization is designed for zero-
centered activations like tanh and sigmoid. For ReLU-based
networks, He initialization scales weights by 2/n, which
better maintains variance across layers.
Question 4
A machine learning engineer trains a convolutional neural
network and notices that many neurons output zero for all
inputs after a few epochs. The network uses ReLU