Bernoulli multi-armed bandit problem under delayed feedback

Open Access

journal article
Published by Taras Shevchenko National University of Kyiv in Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics

No. 1,p. 20-26
https://doi.org/10.17721/1812-5409.2021/1.2

Abstract

Online learning under delayed feedback has been recently gaining increasing attention. Learning with delays is more natural in most practical applications since the feedback from the environment is not immediate. For example, the response to a drug in clinical trials could take a while. In this paper, we study the multi-armed bandit problem with Bernoulli distribution in the environment with delays by evaluating the Explore-First algorithm. We obtain the upper bounds of the algorithm, the theoretical results are applied to develop the software framework for conducting numerical experiments.

Keywords

This publication has 12 references indexed in Scilit:

Introduction to Multi-Armed Bandits
Foundations and Trends® in Machine Learning, 2019
Simulation of Cox Random Processes
Published by Elsevier BV ,2016
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Foundations and Trends® in Machine Learning, 2012
The International Stroke Trial database
Trials, 2011
Metric Characterization of Random Variables and Random Processes
Published by American Mathematical Society (AMS) ,2000
Asymptotically efficient adaptive allocation rules
Advances in Applied Mathematics, 1985
Sequential Medical Trials
Journal of the American Statistical Association, 1963
Probability Inequalities for Sums of Bounded Random Variables
Journal of the American Statistical Association, 1963
Some aspects of the sequential design of experiments
Bulletin of the American Mathematical Society, 1952
ON THE LIKELIHOOD THAT ONE UNKNOWN PROBABILITY EXCEEDS ANOTHER IN VIEW OF THE EVIDENCE OF TWO SAMPLES
Biometrika, 1933

Cited by 2 articles