The ShareChat Machine Learning Seminar invites world-leading researchers to present recent advances in their work in all areas related to machine learning, decision science, and the problems we solve at ShareChat.
The Offline Contextual Bandit framework is a powerful formulation for aligning offline learning of Recommender Systems with online metrics. This is achieved through the use of off-policy estimators that can capture how well a decision system will do online before deployment. Recent advances in off-policy estimation mitigate the statistical limitations of these methods in large catalogue scenarios, but computational issues linked to optimizing these estimators as learning objectives remain unresolved. Indeed, both the complexity and the variance of common optimization routines scale linearly with the catalog size, making these approaches infeasible in practice. This talk proposes optimization routines with sublinear complexities and improved gradient variance for the prevalent, maximum inner product search (MIPS) decision rule class used for fast online delivery. Our first results are based on deriving new gradient formulas for softmax policies, which we improve a second time by reinterpreting policies as randomized decision rules. We will demonstrate these speed-ups for both single-item and slate recommendation problems and empirically validate that the solutions proposed accelerate optimization while enhancing recommendation quality across different scenarios.
Short Bio:
Otmane Sakhi is a Research Scientist at Criteo, having completed his PhD at Institut Polytechnique de Paris and Criteo AI Lab under the guidance of Nicolas Chopin and David Rohde. His doctoral research focused on understanding recommendation through Statistical Learning theory. Specifically, his work contributes and tailors the offline contextual bandit toolbox to improve large scale recommender systems. Prior to his doctoral studies, he earned a M.Eng. degree in Applied Mathematics from CentraleSupelec and an MVA M.Sc. degree from ENS Paris-Saclay. Presently, his key research interest lies in pushing the boundaries of learning from logged interactions to refine real-world, large-scale interactive systems.