CS 7641 OL Unit Quiz| Questions and Answers Latest
Updated 2026/2027 | Georgia Institute of Technology
OL Unit Quiz
CS 7641: Machine Learning
1 Section 1 - Randomized Optimization
Part 1. Randomized Hill Climbing (MCMA)
Which statements about randomized hill climbing (RHC) are true?
a) At each step, RHC samples a neighbor at random and moves only if that neighbor improves
the objective.
b) Without random restarts or sideways moves, RHC is guaranteed to eventually reach a global
optimum in a finite, discrete landscape.
c) If you evaluate many random neighbors per step and choose the best among those samples,
RHC approaches greedy hill climbing.
d) RHC requires gradient information to choose ascent directions.
e) On flat plateaus (all neighbors equal under a strict “improving-only” rule), RHC can stall.
f) With uniform neighbor sampling, the expected improvement per step generally increases as
you approach a local optimum.
Part 2. Simulated Annealing (MCMA)
Which statements about SA are true?
a) At temperature T , a worse move with cost increase ∆E > 0 is accepted with probability
exp(−∆E/T ).
b) With logarithmic cooling and infinite time, SA converges to a global optimum with probability
1.
c) Raising T increases the penalty for uphill moves.
d) SA’s acceptance tests depend on absolute energies, not differences.
e) SA requires gradient and Hessian information.
f) Beneficial moves are accepted with probability exp(+∆E/T ).
1
, Part 3. Genetic Algorithms (Numerical)
A GA operates on fixed-length binary strings of length l = 4. Population size is N = 8. Selec-
tion is fitness-proportionate (roulette wheel) with replacement. Single-point crossover occurs with
probability pc = 0.7. Bitwise mutation occurs independently per bit with probability pm = 0.01.
(a) : Fitness-Proportionate Selection A population of 4 individuals has fitnesses [1, 2, 3, 4]
with total fitness 10. Let x denote the individual with fitness 3. Question: What is the expected
number of copies of x selected into the next mating pool of size N = 8 (selection only)?
(b) : Crossover Survival of a Schema Consider the schema H = 1∗0∗ (positions 1 and 3 fixed).
Two parents both match H. With single-point crossover and pc = 0.7, what is the probability that
an offspring still matches H after crossover (ignoring mutation)?
(c) : Mutation Survival of a Schema Using the same schema H = 1∗0∗ and mutation rate
pm = 0.01, what is the probability the schema survives mutation (i.e., its fixed bits are unchanged)?
2 Section 2 - Deconstructing AdamW
Part 1. AdaGrad & RMSProp (MCMA)
Which statements about per-parameter adaptivity are true?
√
a) AdaGrad accumulates squared gradients and uses ( Gt + ε)−1, shrinking steps on frequently
large-gradient coordinates.
b) AdaGrad’s cumulative memory can make effective learning rates vanish late in training.
c) RMSProp replaces AdaGrad’s cumulative sum with an EMA of squared gradients to track
nonstationary curvature.
d) Typical RMSProp second-moment decay ρ lies around 0.9–0.99.
e) AdaGrad increases learning rates for frequently updated features.
f) RMSProp requires increasing ρ during training to remain stable.
Part 2. L2 vs. Weight Decay & AdamW (MCMA)
Which statements reflect the equivalence/inequivalence results and the AdamW fix?
a) Under SGD with a scalar preconditioner, adding an L2 penalty is equivalent to multiplicative
weight decay each step.
b) Under adaptive preconditioning, L2 shrinkage is coordinate-wise and coupled to gradient
history, breaking equivalence to uniform decay.
c) Equivalence to weight decay holds iff the preconditioner Pt is a positive scalar multiple of the
identity.
2
Updated 2026/2027 | Georgia Institute of Technology
OL Unit Quiz
CS 7641: Machine Learning
1 Section 1 - Randomized Optimization
Part 1. Randomized Hill Climbing (MCMA)
Which statements about randomized hill climbing (RHC) are true?
a) At each step, RHC samples a neighbor at random and moves only if that neighbor improves
the objective.
b) Without random restarts or sideways moves, RHC is guaranteed to eventually reach a global
optimum in a finite, discrete landscape.
c) If you evaluate many random neighbors per step and choose the best among those samples,
RHC approaches greedy hill climbing.
d) RHC requires gradient information to choose ascent directions.
e) On flat plateaus (all neighbors equal under a strict “improving-only” rule), RHC can stall.
f) With uniform neighbor sampling, the expected improvement per step generally increases as
you approach a local optimum.
Part 2. Simulated Annealing (MCMA)
Which statements about SA are true?
a) At temperature T , a worse move with cost increase ∆E > 0 is accepted with probability
exp(−∆E/T ).
b) With logarithmic cooling and infinite time, SA converges to a global optimum with probability
1.
c) Raising T increases the penalty for uphill moves.
d) SA’s acceptance tests depend on absolute energies, not differences.
e) SA requires gradient and Hessian information.
f) Beneficial moves are accepted with probability exp(+∆E/T ).
1
, Part 3. Genetic Algorithms (Numerical)
A GA operates on fixed-length binary strings of length l = 4. Population size is N = 8. Selec-
tion is fitness-proportionate (roulette wheel) with replacement. Single-point crossover occurs with
probability pc = 0.7. Bitwise mutation occurs independently per bit with probability pm = 0.01.
(a) : Fitness-Proportionate Selection A population of 4 individuals has fitnesses [1, 2, 3, 4]
with total fitness 10. Let x denote the individual with fitness 3. Question: What is the expected
number of copies of x selected into the next mating pool of size N = 8 (selection only)?
(b) : Crossover Survival of a Schema Consider the schema H = 1∗0∗ (positions 1 and 3 fixed).
Two parents both match H. With single-point crossover and pc = 0.7, what is the probability that
an offspring still matches H after crossover (ignoring mutation)?
(c) : Mutation Survival of a Schema Using the same schema H = 1∗0∗ and mutation rate
pm = 0.01, what is the probability the schema survives mutation (i.e., its fixed bits are unchanged)?
2 Section 2 - Deconstructing AdamW
Part 1. AdaGrad & RMSProp (MCMA)
Which statements about per-parameter adaptivity are true?
√
a) AdaGrad accumulates squared gradients and uses ( Gt + ε)−1, shrinking steps on frequently
large-gradient coordinates.
b) AdaGrad’s cumulative memory can make effective learning rates vanish late in training.
c) RMSProp replaces AdaGrad’s cumulative sum with an EMA of squared gradients to track
nonstationary curvature.
d) Typical RMSProp second-moment decay ρ lies around 0.9–0.99.
e) AdaGrad increases learning rates for frequently updated features.
f) RMSProp requires increasing ρ during training to remain stable.
Part 2. L2 vs. Weight Decay & AdamW (MCMA)
Which statements reflect the equivalence/inequivalence results and the AdamW fix?
a) Under SGD with a scalar preconditioner, adding an L2 penalty is equivalent to multiplicative
weight decay each step.
b) Under adaptive preconditioning, L2 shrinkage is coordinate-wise and coupled to gradient
history, breaking equivalence to uniform decay.
c) Equivalence to weight decay holds iff the preconditioner Pt is a positive scalar multiple of the
identity.
2