A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting

Philip Amortila, Nan Jiang, Tengyang Xie

Extensions for general d𝑑d and the controlled setting

The extension to the controlled case is similar. Let aa denote the action of sAs_{A} in Figure 1. We introduce a second action a′a^{\prime} for sAs_{A} that transitions to sCs_{C} with reward, and sCs_{C} is absorbing with reward r′r^{\prime}. Let the 2-dimensional feature map be: φ(sA,a)=[γ,0]⊤\varphi(s_{A},a)=[\gamma,0]^{\top}, φ(sA,a′)=[0,γ]⊤\varphi(s_{A},a^{\prime})=[0,\gamma]^{\top}, φ(sB)=⊤\varphi(s_{B})=^{\top}, φ(sC)=⊤\varphi(s_{C})=^{\top}. It is easy to verify that Q⋆Q^{\star} is realizableIn fact, Q-functions in this MDP do not depend on the policy, since only sAs_{A} has multiple actions., but Q⋆(sA,a)=γ1−γrQ^{\star}(s_{A},a)=\frac{\gamma}{1-\gamma}r and Q⋆(sA,a′)=γ1−γr′Q^{\star}(s_{A},a^{\prime})=\frac{\gamma}{1-\gamma}r^{\prime} can independently take arbitrary values between [0,γ/(1−γ)][0,\gamma/(1-\gamma)] (assuming rewards lie in $$), so the learner cannot choose a near-optimal action even with infinite data.

These observations combine to give us the following result:

For any d≥1,γ∈(0,1)d\geq 1,\gamma\in(0,1), given realizable linear features, the value function learned by any batch RL algorithm must have Ω(1)\Omega(1) worst-case error, even with an infinitely large dataset that has Θ(1/d)\Theta(1/d) feature coverage.

Final Remark

While the discounted setting allows a very simple construction for the lower bound, this does not imply that the construction for the finite-horizon setting can be simplified in a similar manner. In fact, we believe that the careful construction of Wang et al. (2020) that cleverly exponentiates a negligibly small error is necessary for the finite-horizon setting. Such a difference between the finite-horizon setting and the discounted setting, however, does challenge the conventional wisdom that the results in the finite-horizon setting and the discounted setting are often similar and translate to each other with H=O(1/(1−γ))H=O(1/(1-\gamma)) up to minor differences. Are these two lower bounds “essentially the same”, or does their difference imply some fundamental difference between the finite-horizon and the discounted settings? We leave this open question to the readers.

Acknowledgement

NJ thanks Ruosong Wang for helpful discussions.

References