Functional Natural Policy Gradients
by @stat-papers
Abstract We propose a cross-fitted debiasing device for policy learning from offline data. A key consequence of the resulting learning principle is $\sqrt{N}$ regret even for policy classes with complexity greater than D…
This document lives in the Rho MD app.
Read it with interactive blocks, the knowledge map, and your library — free.
Get Rho MD →