Page |1
CS 7641 Machine Learning Unit 4 Exam Practice Questions
(2026-2027) Verified Update!!!!!
Instructions: This exam consists of multiple-choice questions.
Choose the best answer for each question. The questions are
organized by the major topic areas within Unit 4: Reinforcement
Learning Foundations, Markov Decision Processes, Value
Iteration, Policy Iteration, Q-Learning, and Advanced RL
Concepts.
Section 1: Reinforcement Learning Foundations (Questions)
1. What is Reinforcement Learning?
A) Supervised learning with labeled data
B) Unsupervised clustering
C) Agent must learn behavior through trial-and-error
interactions in a dynamic environment
D) Dimensionality reduction
Answer: C
Rationale: Reinforcement learning involves an agent that learns
optimal behavior through trial-and-error interactions with a
dynamic environment, receiving rewards or penalties.
, Page |2
2. What is the primary goal of reinforcement learning?
A) To minimize classification error
B) To maximize cumulative reward over time
C) To cluster similar data points
D) To reduce dimensionality
Answer: B
Rationale: The goal of reinforcement learning is to learn a
policy that maximizes the cumulative reward the agent receives
over time.
3. What is a reward in reinforcement learning?
A) A scalar feedback signal indicating the desirability of a state
or action
B) A label for classification
C) A cluster centroid
D) A principal component
Answer: A
Rationale: A reward is a scalar feedback signal that indicates
how good or bad a state or action is, guiding the agent's
learning.
4. What is a policy in reinforcement learning?
A) A mapping from states to actions that defines the agent's
behavior
B) A reward function
, Page |3
C) A transition probability
D) A value function
Answer: A
Rationale: A policy π(s) specifies the action to take in each
state, defining the agent's behavior.
5. What is the value function in reinforcement learning?
A) The expected return (cumulative discounted reward) starting
from a state under a policy
B) The immediate reward
C) The number of states
D) The action taken
Answer: A
Rationale: The value function V^π(s) is the expected return
starting from state s and following policy π thereafter.
6. What is the Q-function in reinforcement learning?
A) The expected return starting from a state, taking an action,
and then following a policy
B) The immediate reward
C) The number of actions
D) The transition probability
Answer: A
Rationale: The Q-function Q^π(s,a) is the expected return
starting from state s, taking action a, and thereafter following
policy π.
, Page |4
7. What is the difference between a model-based and a
model-free RL method?
A) Model-based methods learn or have a model of the
environment; model-free methods do not
B) Model-based methods are always better
C) Model-free methods require a model
D) They are the same
Answer: A
Rationale: Model-based RL methods use or learn a model of
the environment's dynamics (transition and reward functions).
Model-free methods learn values or policies directly from
experience without a model.
8. What is the exploration-exploitation tradeoff?
A) The tradeoff between exploring new actions to learn more
and exploiting known good actions to maximize reward
B) The tradeoff between training and testing
C) The tradeoff between precision and recall
D) The tradeoff between bias and variance
Answer: A
Rationale: Exploration involves trying new actions to discover
better strategies; exploitation involves using known actions that
yield high rewards.
9. What is a deterministic policy?
A) A policy that maps each state to a single action
CS 7641 Machine Learning Unit 4 Exam Practice Questions
(2026-2027) Verified Update!!!!!
Instructions: This exam consists of multiple-choice questions.
Choose the best answer for each question. The questions are
organized by the major topic areas within Unit 4: Reinforcement
Learning Foundations, Markov Decision Processes, Value
Iteration, Policy Iteration, Q-Learning, and Advanced RL
Concepts.
Section 1: Reinforcement Learning Foundations (Questions)
1. What is Reinforcement Learning?
A) Supervised learning with labeled data
B) Unsupervised clustering
C) Agent must learn behavior through trial-and-error
interactions in a dynamic environment
D) Dimensionality reduction
Answer: C
Rationale: Reinforcement learning involves an agent that learns
optimal behavior through trial-and-error interactions with a
dynamic environment, receiving rewards or penalties.
, Page |2
2. What is the primary goal of reinforcement learning?
A) To minimize classification error
B) To maximize cumulative reward over time
C) To cluster similar data points
D) To reduce dimensionality
Answer: B
Rationale: The goal of reinforcement learning is to learn a
policy that maximizes the cumulative reward the agent receives
over time.
3. What is a reward in reinforcement learning?
A) A scalar feedback signal indicating the desirability of a state
or action
B) A label for classification
C) A cluster centroid
D) A principal component
Answer: A
Rationale: A reward is a scalar feedback signal that indicates
how good or bad a state or action is, guiding the agent's
learning.
4. What is a policy in reinforcement learning?
A) A mapping from states to actions that defines the agent's
behavior
B) A reward function
, Page |3
C) A transition probability
D) A value function
Answer: A
Rationale: A policy π(s) specifies the action to take in each
state, defining the agent's behavior.
5. What is the value function in reinforcement learning?
A) The expected return (cumulative discounted reward) starting
from a state under a policy
B) The immediate reward
C) The number of states
D) The action taken
Answer: A
Rationale: The value function V^π(s) is the expected return
starting from state s and following policy π thereafter.
6. What is the Q-function in reinforcement learning?
A) The expected return starting from a state, taking an action,
and then following a policy
B) The immediate reward
C) The number of actions
D) The transition probability
Answer: A
Rationale: The Q-function Q^π(s,a) is the expected return
starting from state s, taking action a, and thereafter following
policy π.
, Page |4
7. What is the difference between a model-based and a
model-free RL method?
A) Model-based methods learn or have a model of the
environment; model-free methods do not
B) Model-based methods are always better
C) Model-free methods require a model
D) They are the same
Answer: A
Rationale: Model-based RL methods use or learn a model of
the environment's dynamics (transition and reward functions).
Model-free methods learn values or policies directly from
experience without a model.
8. What is the exploration-exploitation tradeoff?
A) The tradeoff between exploring new actions to learn more
and exploiting known good actions to maximize reward
B) The tradeoff between training and testing
C) The tradeoff between precision and recall
D) The tradeoff between bias and variance
Answer: A
Rationale: Exploration involves trying new actions to discover
better strategies; exploitation involves using known actions that
yield high rewards.
9. What is a deterministic policy?
A) A policy that maps each state to a single action