CS 7643 EXAM 2 ACTUAL 2026/2027 HIGH YIELD PRACTICE
QUESTIONS AND STUDY GUIDE ACCURATE EXAM COMPLETE
Question 1
Which of the following best describes the operation performed by a convolutional layer?
A. Element-wise multiplication of the kernel with the image patch, summed across all channels
B. Matrix multiplication between the kernel and the flattened image patch
C. Averaging all pixels within the receptive field
D. Applying a non-linear activation to each pixel independently
Correct Answer: A
Rationale: Convolution computes the dot product between the kernel weights and the
image patch for each channel, then sums across channels to produce a single output value. This
preserves spatial structure while learning local features .
Question 2
What is the primary difference between convolution and cross-correlation in the context of
CNNs?
A. Convolution flips the kernel 180 degrees before computing the dot product; cross-correlation
does not.
B. Cross-correlation uses a larger kernel than convolution.
C. Convolution only works on grayscale images.
D. Cross-correlation applies padding while convolution does not.
Correct Answer: A
Rationale: In signal processing, convolution requires flipping the kernel (rotating 180°).
Most deep learning frameworks implement cross-correlation but call it “convolution.” This
distinction matters for gradient computation but not for the forward pass in practice .
Question 3
Given an input image of size 224×224 and a convolutional kernel of size 5×5 with stride 1 and no
padding, what is the spatial dimension of the output feature map?
A. 224×224
B. 220×220
,C. 219×219
D. 218×218
Correct Answer: B
Rationale: The output size formula for valid convolution is (H − K + 1) × (W − K + 1).
Substituting: (224 − 5 + 1) = 220. So the output is 220×220 .
Question 4
Which of the following is TRUE regarding weight sharing in convolutional layers?
A. Each output node has its own unique set of weights.
B. The same kernel weights are applied to every spatial location in the input.
C. Weights are shared only across different channels.
D. Weight sharing increases the total number of learnable parameters.
Correct Answer: B
Rationale: Weight sharing means the same kernel (set of weights) slides across the entire
input, detecting the same feature regardless of spatial location. This drastically reduces
parameters compared to fully connected layers .
Question 5
How does the number of input channels affect the output spatial size of a convolutional layer?
A. It increases the output spatial size proportionally.
B. It decreases the output spatial size proportionally.
C. It has no effect on the output spatial size.
D. It doubles the output spatial size.
Correct Answer: C
Rationale: The output spatial dimensions are determined by the input height/width, kernel
size, stride, and padding — not by the number of channels. The dot product is computed for
each channel and summed to produce one output value per spatial location .
Question 6
How many learnable parameters does a convolutional layer have, given the following: input
channels = 3, number of kernels = 64, kernel size = 3×3, and bias is used?
,A. 64 × (3 × 3 × 3) = 1,728
B. 64 × (3 × 3 × 3 + 1) = 1,792
C. 3 × (64 × 3 × 3 + 1) = 1,731
D. 64 × 3 × 3 = 576
Correct Answer: B
Rationale: Each kernel has K1 × K2 × Channels weights plus 1 bias. Total = M × (K1 × K2 × Ch
+ 1) = 64 × (3 × 3 × 3 + 1) = 64 × 28 = 1,792 .
Question 7
What is the primary purpose of a pooling layer in a CNN?
A. To introduce non-linearity
B. To increase the spatial resolution of feature maps
C. To reduce spatial dimensionality and provide translation invariance
D. To learn new features from the input
Correct Answer: C
Rationale: Pooling reduces the spatial dimensions of feature maps (downsampling), which
decreases computation and provides a degree of translation invariance. Pooling layers have no
learnable parameters .
Question 8
How many learned parameters does a 2×2 max pooling layer have?
A. 4
B. 2
C. 1
D. 0
Correct Answer: D
Rationale: Max pooling (and average pooling) are fixed operations with no learnable
weights or biases. The operation is deterministic and parameter-free .
Question 9
What is the receptive field of a neuron in a convolutional network?
, A. The total number of neurons in the layer
B. The region of the input image that affects that neuron’s output
C. The kernel size divided by the stride
D. The number of channels in the input
Correct Answer: B
Rationale: The receptive field is the specific region of the input (image patch) from which a
neuron receives input. In deeper layers, the effective receptive field grows as features are
composed hierarchically .
Question 10
Which statement best distinguishes invariance from equivariance in CNNs?
A. Invariance means output changes with input translation; equivariance means output stays
the same.
B. Invariance means output stays the same under input transformation; equivariance means
output transforms similarly.
C. Both terms mean the same thing.
D. Invariance applies only to pooling layers; equivariance applies only to convolution layers.
Correct Answer: B
Rationale: Invariance: if the feature moves slightly, the output value remains unchanged
(e.g., classification). Equivariance: if the feature translates, the output values shift by the same
translation and can be detected at the new location .
Question 11
What happens to the output size of a convolutional layer when you add padding of P=1 to an
input of size 32×32 with a 3×3 kernel and stride 1?
A. Output remains 30×30
B. Output becomes 32×32
C. Output becomes 34×34
D. Output becomes 28×28
Correct Answer: B
QUESTIONS AND STUDY GUIDE ACCURATE EXAM COMPLETE
Question 1
Which of the following best describes the operation performed by a convolutional layer?
A. Element-wise multiplication of the kernel with the image patch, summed across all channels
B. Matrix multiplication between the kernel and the flattened image patch
C. Averaging all pixels within the receptive field
D. Applying a non-linear activation to each pixel independently
Correct Answer: A
Rationale: Convolution computes the dot product between the kernel weights and the
image patch for each channel, then sums across channels to produce a single output value. This
preserves spatial structure while learning local features .
Question 2
What is the primary difference between convolution and cross-correlation in the context of
CNNs?
A. Convolution flips the kernel 180 degrees before computing the dot product; cross-correlation
does not.
B. Cross-correlation uses a larger kernel than convolution.
C. Convolution only works on grayscale images.
D. Cross-correlation applies padding while convolution does not.
Correct Answer: A
Rationale: In signal processing, convolution requires flipping the kernel (rotating 180°).
Most deep learning frameworks implement cross-correlation but call it “convolution.” This
distinction matters for gradient computation but not for the forward pass in practice .
Question 3
Given an input image of size 224×224 and a convolutional kernel of size 5×5 with stride 1 and no
padding, what is the spatial dimension of the output feature map?
A. 224×224
B. 220×220
,C. 219×219
D. 218×218
Correct Answer: B
Rationale: The output size formula for valid convolution is (H − K + 1) × (W − K + 1).
Substituting: (224 − 5 + 1) = 220. So the output is 220×220 .
Question 4
Which of the following is TRUE regarding weight sharing in convolutional layers?
A. Each output node has its own unique set of weights.
B. The same kernel weights are applied to every spatial location in the input.
C. Weights are shared only across different channels.
D. Weight sharing increases the total number of learnable parameters.
Correct Answer: B
Rationale: Weight sharing means the same kernel (set of weights) slides across the entire
input, detecting the same feature regardless of spatial location. This drastically reduces
parameters compared to fully connected layers .
Question 5
How does the number of input channels affect the output spatial size of a convolutional layer?
A. It increases the output spatial size proportionally.
B. It decreases the output spatial size proportionally.
C. It has no effect on the output spatial size.
D. It doubles the output spatial size.
Correct Answer: C
Rationale: The output spatial dimensions are determined by the input height/width, kernel
size, stride, and padding — not by the number of channels. The dot product is computed for
each channel and summed to produce one output value per spatial location .
Question 6
How many learnable parameters does a convolutional layer have, given the following: input
channels = 3, number of kernels = 64, kernel size = 3×3, and bias is used?
,A. 64 × (3 × 3 × 3) = 1,728
B. 64 × (3 × 3 × 3 + 1) = 1,792
C. 3 × (64 × 3 × 3 + 1) = 1,731
D. 64 × 3 × 3 = 576
Correct Answer: B
Rationale: Each kernel has K1 × K2 × Channels weights plus 1 bias. Total = M × (K1 × K2 × Ch
+ 1) = 64 × (3 × 3 × 3 + 1) = 64 × 28 = 1,792 .
Question 7
What is the primary purpose of a pooling layer in a CNN?
A. To introduce non-linearity
B. To increase the spatial resolution of feature maps
C. To reduce spatial dimensionality and provide translation invariance
D. To learn new features from the input
Correct Answer: C
Rationale: Pooling reduces the spatial dimensions of feature maps (downsampling), which
decreases computation and provides a degree of translation invariance. Pooling layers have no
learnable parameters .
Question 8
How many learned parameters does a 2×2 max pooling layer have?
A. 4
B. 2
C. 1
D. 0
Correct Answer: D
Rationale: Max pooling (and average pooling) are fixed operations with no learnable
weights or biases. The operation is deterministic and parameter-free .
Question 9
What is the receptive field of a neuron in a convolutional network?
, A. The total number of neurons in the layer
B. The region of the input image that affects that neuron’s output
C. The kernel size divided by the stride
D. The number of channels in the input
Correct Answer: B
Rationale: The receptive field is the specific region of the input (image patch) from which a
neuron receives input. In deeper layers, the effective receptive field grows as features are
composed hierarchically .
Question 10
Which statement best distinguishes invariance from equivariance in CNNs?
A. Invariance means output changes with input translation; equivariance means output stays
the same.
B. Invariance means output stays the same under input transformation; equivariance means
output transforms similarly.
C. Both terms mean the same thing.
D. Invariance applies only to pooling layers; equivariance applies only to convolution layers.
Correct Answer: B
Rationale: Invariance: if the feature moves slightly, the output value remains unchanged
(e.g., classification). Equivariance: if the feature translates, the output values shift by the same
translation and can be detected at the new location .
Question 11
What happens to the output size of a convolutional layer when you add padding of P=1 to an
input of size 32×32 with a 3×3 kernel and stride 1?
A. Output remains 30×30
B. Output becomes 32×32
C. Output becomes 34×34
D. Output becomes 28×28
Correct Answer: B