Fetching the paper…
Reading the bibliography…
In many RL applications, once training ends, it is vital to detect any deterioration in the agent performance as soon as possible.
Lecarpentier, E. and Rachelson, E · 1904
Earlier work this paper cites.
The generalization of student’s ratio
Hotelling, H · 1931
Earlier work this paper cites.
On the problem of the most efficient tests of statistical hypotheses
Neyman, J., Pearson, E. S., and Pearson, K · 1933
Earlier work this paper cites.
The large-sample distribution of the likelihood ratio for testing composite hypotheses
Wilks, S. S · 1938
Earlier work this paper cites.
Sequential tests of statistical hypotheses
Wald, A · 1945
Earlier work this paper cites.
Continuous Inspection Schemes
Page, E. S · 1954
Earlier work this paper cites.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
An approach to the probability distribution of cusum run length
Brook, D. et al · 1972
Earlier work this paper cites.
Sums of Independent Random Variables
Petrov, V. V · 1972
Earlier work this paper cites.
Group sequential methods in the design and analysis of clinical trials
Pocock, S. J · 1977
Earlier work this paper cites.
Combined shewhart-cusum control chart for improved quality control in clinical chemistry
Westgard, J., Groth, T., Aronsson, T., and Verdier, C · 1977
Earlier work this paper cites.
Distribution of the estimators for autoregressive time series with a unit root
Dickey, D. A. and Fuller, W. A · 1979
Earlier work this paper cites.
Multivariate analysis
K. V. Mardia, J. T. K. and Bibby, J. M · 1979
Earlier work this paper cites.
A multiple testing procedure for clinical trials
O’Brien, P. C. and Fleming, T. R · 1979
Earlier work this paper cites.
Bandit problems
Berry, D. A. and Fristedt, B · 1985
Earlier work this paper cites.
On the analysis and design of cusum-shewhart control schemes
Yashchin, E · 1985
Earlier work this paper cites.
Quality control: an application of the cusum
Williams, S. M. et al · 1992
Earlier work this paper cites.
Interim analysis: The alpha spending function approach
Lan, D. L. D. K. K. G · 1994
Earlier work this paper cites.
Optimization of conditional value-at-risk
Rockafellar, R. T. and Uryasev, S · 2000
Earlier work this paper cites.
Cusum charts for monitoring an autocorrelated process
Lu, C.-W. and Jr., M. R. R · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., van der Laan, M. J., and Robins, J. M · 2001
Earlier work this paper cites.
Second thoughts on the bootstrap
Efron, B · 2003
Earlier work this paper cites.
Cycle-based signal monitoring using a directionally variant multivariate control chart system
Zhou, S., Jin, N., and Jin, J. J · 2005
Earlier work this paper cites.
Lecture notes: Convergence in distribution and central limit theorem
Irwin, M. E · 2006
Earlier work this paper cites.
Changepoint Detection in Periodic and Autocorrelated Time Series
Lund, R., Wang, X. L., Lu, Q. Q., Reeves, J., Gallagher, C., and Feng, Y · 2007
Cited alongside, same era.
Lecture notes in stat c141: The bonferroni correction
Goldman, M · 2008
Cited alongside, same era.
On upper-confidence bound policies for switching bandit problems
Garivier, A. and Moulines, E · 2011
Cited alongside, same era.
Change detection in streaming multivariate data using likelihood detectors
Kuncheva, L. I · 2011
Cited alongside, same era.
Statistical Methods for Quality Improvement
Ryan, T. P · 2011
Cited alongside, same era.
Weak change-point detection using temporal correlation, 2011
Xie, Y. and Siegmund, D · 2011
Cited alongside, same era.
Time limits in reinforcement learning
Pardo, F., Tavakoli, A., Levdik, V., and Kormushev, P · 2017
Later among the works it cites.
Quanttree: Histograms for change detection in multivariate data streams
Boracchi, G., Carrera, D., Cervellera, C., and Maccio, D · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Chow, Y. et al · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Later among the works it cites.
Adversarial attacks on stochastic bandits
Jun, K.-S. et al · 2018
Later among the works it cites.
Pytorch implementations of reinforcement learning algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Multivariate bernoulli distribution
Dai, B., Ding, S., and Wahba, G · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Gur, Y., and Zeevi, A · 2014
Cited alongside, same era.
Concept drift detection through resampling
Harel, M., Crammer, K., El-Yaniv, R., and Mannor, S · 2014
Cited alongside, same era.
Why the monte carlo method is so important today
Kroese, D. P., Brereton, T., Taimre, T., and Botev, Z · 2014
Cited alongside, same era.
Learning in nonstationary environments: A survey
Ditzler, G., Polikar, R., and Alippi, C · 2015
Cited alongside, same era.
Kostrikov, I · 2018
Later among the works it cites.
Correlation priors for reinforcement learning
Alt, B., Sosic, A., and Koeppl, H · 2019
Later among the works it cites.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Cheng, R. et al · 2019
Later among the works it cites.
A hitchhiker’s guide to statistical comparisons of reinforcement learning algorithms, 2019
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2019
Later among the works it cites.
Challenges of real-world reinforcement learning, 2019
Dulac-Arnold, G., Mankowitz, D., and Hester, T · 2019
Later among the works it cites.
Better algorithms for stochastic bandits with adversarial corruptions
Gupta, A., Koren, T., and Talwar, K · 2019
Later among the works it cites.
A survey of learning in multiagent environments: Dealing with non-stationarity, 2019
Hernandez-Leal, P., Kaisers, M., Baarslag, T., and de Cote, E. M · 2019
Later among the works it cites.
Autoregressive policies for continuous control deep reinforcement learning, 2019
Korenkevych, D., Mahmood, A. R., Vasan, G., and Bergstra, J · 2019
Later among the works it cites.
Distribution-dependent and time-uniform bounds for piecewise i.i.d bandits
Mukherjee, S. and Maillard, O.-A · 2019
Later among the works it cites.
Assessing the safety and reliability of autonomous vehicles from road testing
Zhao, X. et al · 2019
Later among the works it cites.
Multi-player bandits: The adversarial case
Alatur, P., Levy, K. Y., and Krause, A · 2020
Closest in time.
Agent57: Outperforming the atari human benchmark
Badia, A. P. et al · 2020
Closest in time.
Measuring the reliability of reinforcement learning algorithms
Chan, S. C. et al · 2020
Closest in time.
Conditional value at risk (cvar)
Chen, J · 2020
Closest in time.
Context-aware dynamics model for generalization in model-based rl
Lee, K. et al · 2020
Closest in time.
Bandits with adversarial scaling
Lykouris, T., Mirrokni, V., and Leme, R. P · 2020
Closest in time.
Deployment-efficient reinforcement learning via model-based offline optimization
Matsushima, T., Furuta, H., Matsuo, Y., Nachum, O., and Gu, S · 2020
Closest in time.
Multi-agent manipulation via locomotion using hierarchical sim2real
Nachum, O., Ahn, M., Ponte, H., Gu, S. S., and Kumar, V · 2020
Closest in time.
Mopo: Model-based offline policy optimization, 2020
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Closest in time.