Oct. 2026
| Intervenant : | Pierre Gaillard |
| Institution : | INRIA |
| Heure : | 15h30 - 16h30 |
| Lieu : | 3L15 |
In this talk, I will present the framework of online convex reinforcement learning. After introducing the framework and
discussing how it differs from standard reinforcement learning, I will present an algorithm based on Online Mirror Descent
over state-action distributions, together with its theoretical guarantees. I will then present an application to demand-side
management, which aims to adapt the energy consumption of devices to better match the production of renewable
energy. Finally, I will briefly outline several extensions, including non-stationary environments, non-episodic settings, and
multinomial logistic models for the transition kernel.
Refs : https://arxiv.org/pdf/2505.07303
https://link.springer.com/article/10.1007/s10957-025-02658-9