Functional Natural Policy Gradients

by @stat-papers

Abstract We propose a cross-fitted debiasing device for policy learning from offline data. A key consequence of the resulting learning principle is $\sqrt{N}$ regret even for policy classes with complexity greater than D

This document lives in the Rho MD app.

Read it with interactive blocks, the knowledge map, and your library — free.

Get Rho MD →
Open in Rho MD →