We are in part 2 of the series in Multi-Armed Bandits. This time around the focus is around the improvement upon the greedy algorithms from the previous streams. In this session, we looked at explore-then-commit algorithms. We learned about the concept of Empirical Means and how its it the foundation for UCB Algorithms. We understood and explored the UCB (Upper Confidence Bounds) algorithms and its various improvements, Like UCB-V, UCB-V Tuned, UCB plus. We also discussed some more recent boosting algorithms on UCB Boost, that will be covered in the next session.
This session is suitable for beginners and included hands-on exercises in Python. So I would love it if you Subscribe to my channel: http://bit.ly/aiwyz-yt
If you missed the previous session, you can view it here: • An absolute beginners guide to multi-...
I would love it if you Subscribe to my channel: http://bit.ly/aiwyz-yt
The code that we used in this session can be found here:: http://bit.ly/githubMAB or https://github.com/setuc/multi-arm-ba...
The papers that we referred can be downloaded here: http://bit.ly/MABYT-29
1. Audibert, Munos, Szepesvári - Exploration-exploitation tradeoff using variance estimates in multi-armed bandits - 2009
2. Even-Bar, Mannor, Mansour - Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
3. Garivier, Cappé - The KL-UCB algorithm for bounded stochastic bandits and beyond - 2011
4. Garivier, Kaufmann - Optimal best arm identification with fixed confidence
5. Lai, Robbins - Asymptotically efficient adaptive allocation rules - 1985
6. Liu et al. - UCBoost A boosting approach to tame complexity and optimality for stochastic bandits