Bernoulli multi-armed bandit problem under delayed feedback
Open Access
- 1 January 2021
- journal article
- Published by Taras Shevchenko National University of Kyiv in Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics
- No. 1,p. 20-26
- https://doi.org/10.17721/1812-5409.2021/1.2
Abstract
Online learning under delayed feedback has been recently gaining increasing attention. Learning with delays is more natural in most practical applications since the feedback from the environment is not immediate. For example, the response to a drug in clinical trials could take a while. In this paper, we study the multi-armed bandit problem with Bernoulli distribution in the environment with delays by evaluating the Explore-First algorithm. We obtain the upper bounds of the algorithm, the theoretical results are applied to develop the software framework for conducting numerical experiments.Keywords
This publication has 12 references indexed in Scilit:
- Introduction to Multi-Armed BanditsFoundations and Trends® in Machine Learning, 2019
- Simulation of Cox Random ProcessesPublished by Elsevier BV ,2016
- Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit ProblemsFoundations and Trends® in Machine Learning, 2012
- The International Stroke Trial databaseTrials, 2011
- Metric Characterization of Random Variables and Random ProcessesPublished by American Mathematical Society (AMS) ,2000
- Asymptotically efficient adaptive allocation rulesAdvances in Applied Mathematics, 1985
- Sequential Medical TrialsJournal of the American Statistical Association, 1963
- Probability Inequalities for Sums of Bounded Random VariablesJournal of the American Statistical Association, 1963
- Some aspects of the sequential design of experimentsBulletin of the American Mathematical Society, 1952
- ON THE LIKELIHOOD THAT ONE UNKNOWN PROBABILITY EXCEEDS ANOTHER IN VIEW OF THE EVIDENCE OF TWO SAMPLESBiometrika, 1933