CS 7641: Machine Learning
1 Question 1 - Markov Decision Processes
Part 1. Policies in Reinforcement Learning (MCMA)
You are training an agent to navigate a maze with multiple goal states. The agent uses a stochastic
policy that maps states to probability distributions over actions.
Which of the following statements about policies are true? Select all that apply.
A. A deterministic policy always selects the same action for a given state.
B. A stochastic policy can assign non-zero probability to multiple actions in the same state.
C. The optimal policy is always unique for any MDP.
D. A policy can be evaluated by computing its expected return starting from each possible state.
E. The policy improvement theorem guarantees that a better policy can always be found after
each iteration.
F. A random policy that chooses all actions uniformly is generally optimal for large state spaces.
Part 2. Rewards in Reinforcement Learning (MCMA)
An agent is learning to maximize total reward in an environment with delayed and sparse rewards.
Which of the following statements about rewards are true? Select all that apply.
A. Shaping rewards can help the agent learn faster by providing intermediate signals.
B. The reward signal directly defines what the agent should optimize for.
C. Adding random noise to the reward always improves exploration.
D. Sparse rewards can make it more difficult for the agent to learn an optimal policy.
1
, E. Discounting future rewards too heavily can cause the agent to ignore long-term consequences.
F. The total return is defined as the sum of immediate rewards only, without any consideration
of future rewards.
2