Fetching the paper…
Reading the bibliography…
Artificial behavioral agents are often evaluated based on their consistent behaviors and performance to take sequential actions in an environment to maximize some notion of cumulative reward.
Thompson, W.: On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25
1933
Earlier work this paper cites.
Tversky, A., Kahneman, D.: The Framing of Decisions and the Psychology of Choice. Science 211
1981
Earlier work this paper cites.
Lai, T.L., Robbins, H.: Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics 6
1985
Earlier work this paper cites.
Bechara, A., Damasio, A.R., Damasio, H., Anderson, S.W.: Insensitivity to future consequences following damage to human prefrontal cortex. Cognition 50
1994
Earlier work this paper cites.
Rummery, G.A., Niranjan, M.: On-line Q-learning using connectionist systems, vol. 37. University of Cambridge, Department of Engineering Cambridge, England (1994)
1994
Earlier work this paper cites.
Schultz, W., Dayan, P., Montague, P.R.: A Neural Substrate of Prediction and Reward. Science 275
1997
Earlier work this paper cites.
Auer, P., Cesa-Bianchi, N.: On-line learning with malicious noise and the closure algorithm. Ann. Math. Artif. Intell. 23
1998
Earlier work this paper cites.
Sutton, R.S., Barto, A.G.: Introduction to Reinforcement Learning. MIT Press, Cambridge, MA, USA, 1st edn. (1998)
1998
Earlier work this paper cites.
Sutton, R.S., Barto, A.G., et al.: Introduction to reinforcement learning, vol. 135. MIT press Cambridge (1998)
1998
Earlier work this paper cites.
Auer, P., Cesa-Bianchi, N., Fischer, P.: Finite-time analysis of the multiarmed bandit problem. Machine Learning 47
2002
Earlier work this paper cites.
Auer, P., Cesa-Bianchi, N., Freund, Y., Schapire, R.E.: The nonstochastic multiarmed bandit problem. SIAM J. Comput. 32
2002
Earlier work this paper cites.
Auer, P., Cesa-Bianchi, N., Freund, Y., Schapire, R.E.: The nonstochastic multiarmed bandit problem. SIAM Journal on Computing 32
2002
Earlier work this paper cites.
Even-Dar, E., Mansour, Y.: Learning rates for q-learning. Journal of Machine Learning Research 5
2003
Earlier work this paper cites.
Frank, M.J., Seeberger, L.C., O’reilly, R.C.: By carrot or by stick: cognitive reinforcement learning in parkinsonism. Science 306
2004
Earlier work this paper cites.
O’Doherty, J., Dayan, P., Schultz, J., Deichmann, R., Friston, K., Dolan, R.J.: Dissociable Roles of Ventral and Dorsal Striatum in Instrumental. Science 304
2004
Earlier work this paper cites.
Bayer, H.M., Glimcher, P.W.: Midbrain Dopamine Neurons Encode a Quantitative Reward Prediction Error Signal. Neuron 47
2005
Earlier work this paper cites.
Frank, M.J., O’Reilly, R.C.: A Mechanistic Account of Striatal Dopamine Function in Human Cognition: Psychopharmacological Studies With Cabergoline and Haloperidol. Behavioral Neuroscience 120
2006
Cited alongside, same era.
Seymour, B., Singer, T., Dolan, R.: The neurobiology of punishment. Nature Reviews Neuroscience 8
2007
Cited alongside, same era.
Dayan, P., Niv, Y.: Reinforcement learning: the good, the bad and the ugly. Current opinion in neurobiology 18
2008
Cited alongside, same era.
Langford, J., Zhang, T.: The epoch-greedy algorithm for multi-armed bandits with side information. In: Advances in neural information processing systems. pp. 817–824 (2008)
2008
Cited alongside, same era.
Fridberg, D.J., Queller, S., Ahn, W.Y., Kim, W., Bishara, A.J., Busemeyer, J.R., Porrino, L., Stout, J.C.: Cognitive mechanisms underlying risky decision-making in chronic cannabis users. Journal of mathematical psychology 54
Bouneffouf, D., Féraud, R.: Multi-armed bandit problem with known trend. Neurocomputing 205
2016
Later among the works it cites.
Bouneffouf, D., Rish, I., Cecchi, G.A.: Bandit models of human behavior: Reward processing in mental disorders. In: International Conference on Artificial General Intelligence. pp. 237–248. Springer (2017)
2017
Later among the works it cites.
Bouneffouf, D., Rish, I., Cecchi, G.A., Féraud, R.: Context attentive bandits: contextual bandit with restricted context. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. pp. 1468–1475 (2017)
2017
Later among the works it cites.
Elfwing, S., Seymour, B.: Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the maxpain algorithm. In: 2017 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob). pp. 140–147. IEEE (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
Hasselt, H.V.: Double q-learning. In: Advances in Neural Information Processing Systems. pp. 2613–2621 (2010)
2010
Cited alongside, same era.
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., Schapire, R.: Contextual bandit algorithms with supervised learning guarantees. In: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. pp. 19–26 (2011)
2011
Cited alongside, same era.
Chapelle, O., Li, L.: An empirical evaluation of thompson sampling. In: Advances in neural information processing systems. pp. 2249–2257 (2011)
2011
Cited alongside, same era.
Li, L., Chu, W., Langford, J., Wang, X.: Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms. In: King, I., Nejdl, W., Li, H. (eds.) WSDM. pp. 297–306. ACM (2011), http://dblp.uni-trier.de/db/conf/wsdm/wsdm2011.html#LiCLW11
2011
Cited alongside, same era.
Maia, T.V., Frank, M.J.: From reinforcement learning models to psychiatric and neurological disorders. Nature Neuroscience 14
2011
Cited alongside, same era.
Agrawal, S., Goyal, N.: Analysis of thompson sampling for the multi-armed bandit problem. In: COLT 2012 - The 25th Annual Conference on Learning Theory, June 25-27, 2012, Edinburgh, Scotland. pp. 39.1–39.26 (2012), http://www.jmlr.org/proceedings/papers/v23/agrawal12/agrawal12.pdf
2012
Cited alongside, same era.
Horstmann, A., Villringer, A., Neumann, J.: Iowa gambling task: There is more to consider than long-term outcome. using a linear equation model to disentangle the impact of outcome and frequency of gains and losses. Frontiers in Neuroscience 6
2012
Cited alongside, same era.
Holmes, A.J., Patrick, L.M.: The Myth of Optimality in Clinical Neuroscience. Trends in Cognitive Sciences 22
2017
Later among the works it cites.
Lin, B., Bouneffouf, D., Cecchi, G.A., Rish, I.: Contextual bandit with adaptive feature extraction. In: 2018 IEEE International Conference on Data Mining Workshops (ICDMW). pp. 937–944. IEEE (2018)
2018
Later among the works it cites.
Lin, B., Bouneffouf, D., Cecchi, G.: Split q learning: reinforcement learning with two-stream rewards. In: Proceedings of the 28th International Joint Conference on Artificial Intelligence. pp. 6448–6449. AAAI Press (2019)
2019
Later among the works it cites.
Lin, B.: Diabolical games: Reinforcement learning environments for lifelong learning (2020)
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Lin, B., Bouneffouf, D., Reinen, J., Rish, I., Cecchi, G.: A story of two streams: Reinforcement learning models from human behavior and neuropsychiatry. In: Proceedings of the Nineteenth International Conference on Autonomous Agents and Multi-Agent Systems, AAMAS-20. pp. 744–752. International Foundation for Autonomous Agents and Multiagent Systems (5 2020)
2020
Closest in time.
2020
Closest in time.
Lin, B., Zhang, X.: VoiceID on the fly: A speaker recognition system that learns from scratch. In: INTERSPEECH (2020)
2020
Closest in time.
Lin, B.: Offline reinforcement learning in bandits. arXiv preprint (2021)
2021
Closest in time.