Fetching the paper…
Reading the bibliography…
Reinforcement learners are agents that learn to pick actions that lead to high reward.
W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika , vol. 25, no. 3/4, pp. 285–294, 1933
1933
Earlier work this paper cites.
D. Blackwell and L. Dubins, “Merging of opinions with increasing information,” The Annals of Mathematical Statistics , vol. 33, no. 3, pp. 882–886, 1962
1962
Earlier work this paper cites.
J. A. Clouse and P. E. Utgoff, “A teaching method for reinforcement learning,” in Machine Learning Proceedings 1992 . Elsevier, 1992, pp. 92–101
1992
Earlier work this paper cites.
M. Heger, “Consideration of risk in reinforcement learning,” in Machine Learning Proceedings 1994 . Elsevier, 1994, pp. 105–111
1994
Earlier work this paper cites.
J. A. Clouse, “On integrating apprentice learning and reinforcement learning,” Ph.D. dissertation, University of Massachusetts Amherst, 1997
1997
Earlier work this paper cites.
C. G. Atkeson and S. Schaal, “Robot learning from demonstration,” in ICML , vol. 97. Citeseer, 1997, pp. 12–20
1997
Earlier work this paper cites.
R. Maclin and J. W. Shavlik, “Creating advice-taking reinforcement learners,” in Learning to learn . Springer, 1998, pp. 311–347
1998
Earlier work this paper cites.
O. Mihatsch and R. Neuneier, “Risk-sensitive reinforcement learning,” Machine learning , vol. 49, no. 2-3, pp. 267–290, 2002
2002
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1
2004
Earlier work this paper cites.
A. Nilim and L. El Ghaoui, “Robust control of markov decision processes with uncertain transition matrices,” Operations Research , vol. 53, no. 5, pp. 780–798, 2005
2005
Earlier work this paper cites.
M. Hutter, Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability . Berlin, Germany: Springer, 2005
2005
Earlier work this paper cites.
A. L. Thomaz, C. Breazeal et al. , “Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance,” in Aaai , vol. 6. Boston, MA, 2006, pp. 1000–1005
2006
Earlier work this paper cites.
U. Syed and R. E. Schapire, “A game-theoretic approach to apprenticeship learning,” in NIPS , 2008
2008
Earlier work this paper cites.
A. Hans, D. Schneegaß, A. M. Schäfer, and S. Udluft, “Safe exploration for reinforcement learning.” in ESANN , 2008, pp. 143–148
2008
Earlier work this paper cites.
L. Orseau, “Optimality issues of universal greedy agents with static priors,” in International Conference on Algorithmic Learning Theory . Springer, 2010, pp. 345–359
2010
Earlier work this paper cites.
V. S. Borkar, “Learning algorithms for risk-sensitive control,” in Proceedings of the 19th International Symposium on Mathematical Theory of Networks and Systems–MTNS , vol. 5, no. 9, 2010
2010
Cited alongside, same era.
K. Judah, S. Roy, A. Fern, and T. Dietterich, “Reinforcement learning via practice and critique advice,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 24, no. 1, 2010
2010
Cited alongside, same era.
T. Lattimore and M. Hutter, “Asymptotically optimal agents,” in Proc. 22nd International Conf. on Algorithmic Learning Theory (ALT’11) , ser. LNAI, vol. 6925. Espoo, Finland: Springer, 2011, pp. 368–382
2011
Cited alongside, same era.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proc. 14th International Conf. on Artificial Intelligence and Statistics , 2011, pp. 627–635
2011
Cited alongside, same era.
J. Leike, T. Lattimore, L. Orseau, and M. Hutter, “Thompson sampling is asymptotically optimal in general environments,” in Proc. 32nd International Conf. on Uncertainty in Artificial Intelligence (UAI’16) . New Jersey, USA: AUAI Press, 2016, pp. 417–426
2016
Later among the works it cites.
J. Leike, “Nonparametric general reinforcement learning,” arXiv preprint arXiv:1611.08944 , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in neural information processing systems , 2016, pp. 4565–4573
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Veness, K. S. Ng, M. Hutter, W. Uther, and D. Silver, “A Monte-Carlo AIXI approximation,” Journal of Artificial Intelligence Research , vol. 40, pp. 95–142, 2011
2011
Cited alongside, same era.
2012
Cited alongside, same era.
J. García and F. Fernández, “Safe exploration of state and action spaces in reinforcement learning,” Journal of Artificial Intelligence Research , vol. 45, pp. 515–564, 2012
2012
Cited alongside, same era.
2012
Cited alongside, same era.
L. Orseau, T. Lattimore, and M. Hutter, “Universal knowledge-seeking agents for stochastic environments,” in Proc. 24th International Conf. on Algorithmic Learning Theory (ALT’13) , ser. LNAI, vol. 8139. Singapore: Springer, 2013, pp. 158–172
2013
Cited alongside, same era.
J. García, D. Acera, and F. Fernández, “Safe reinforcement learning through probabilistic policy reuse,” RLDM 2013 , p. 14, 2013
2013
Cited alongside, same era.
T. Lattimore and M. Hutter, “Bayesian reinforcement learning with exploration,” in International Conference on Algorithmic Learning Theory . Springer, 2014, pp. 170–184
2014
Cited alongside, same era.
T. Lattimore and M. Hutter, “General time consistent discounting,” Theoretical Computer Science , vol. 519, pp. 140–154, 2014
2014
Cited alongside, same era.
M. Turchetta, F. Berkenkamp, and A. Krause, “Safe exploration in finite markov decision processes with gaussian processes,” in Advances in Neural Information Processing Systems , 2016, pp. 4312–4320
2016
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Aslanides, J. Leike, and M. Hutter, “Universal reinforcement learning algorithms: survey and experiments,” in Proceedings of the 26th International Joint Conference on Artificial Intelligence , 2017, pp. 1403–1410
2017
Later among the works it cites.
S. Lamont, J. Aslanides, J. Leike, and M. Hutter, “Generalised discount functions applied to a monte-carlo ai u implementation,” in Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems , 2017, pp. 1589–1591
2017
Later among the works it cites.
J. Leike and M. Hutter, “On the computability of Solomonoff induction and AIXI,” Theoretical Computer Science , vol. 716, pp. 28–49, 2018
2018
Later among the works it cites.
M. K. Cohen, E. Catt, and M. Hutter, “A strongly asymptotically optimal agent in general environments,” IJCAI , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Cohen, B. Vellambi, and M. Hutter, “Asymptotically unambitious artificial general intelligence,” in Proc. 34rd AAAI Conference on Artificial Intelligence (AAAI’20) , vol. 34. New York, USA: AAAI Press, 2020
2020
Closest in time.
W. Saunders, G. Sastry, A. Stuhlmueller, and O. Evans, “Trial without error: Towards safe reinforcement learning via human intervention,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems , 2018, pp. 2067–2069
2069
Closest in time.