no code implementations • 8 Mar 2024 • Alex Ayoub, Kaiwen Wang, Vincent Liu, Samuel Robertson, James McInerney, Dawen Liang, Nathan Kallus, Csaba Szepesvári
We propose training fitted Q-iteration with log-loss (FQI-LOG) for batch reinforcement learning (RL).