Fetching the paper…
Reading the bibliography…
We consider a two-agent MDP framework where agents repeatedly solve a task in a collaborative setting.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Boutilier, C · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E · 1997
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Blum, A., Kalai, A., and Wasserman, H · 2003
Earlier work this paper cites.
Experts in a markov decision process
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2005
Earlier work this paper cites.
Regret minimization and the price of total anarchy
Blum, A., Hajiaghayi, M., Ligett, K., and Roth, A · 2008
Earlier work this paper cites.
On agnostic boosting and parity learning
Kalai, A. T., Mansour, Y., and Verbin, E · 2008
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2009
Earlier work this paper cites.
Intrinsic robustness of the price of anarchy
Roughgarden, T · 2009
Earlier work this paper cites.
Arbitrarily modulated markov decision processes
Yu, J. Y. and Mannor, S · 2009
Earlier work this paper cites.
Markov decision processes with arbitrary reward processes
Yu, J. Y., Mannor, S., and Shimkin, N · 2009
Earlier work this paper cites.
Policy teaching through reward function learning
Zhang, H., Parkes, D. C., and Chen, Y · 2009
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Neu, G., Antos, A., György, A., and Szepesvári, C · 2010
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Daskalakis, C., Deckelbaum, A., and Kim, A · 2011
Cited alongside, same era.
Algorithmic and human teaching of sequential decision tasks
Cakmak, M. and Lopes, M · 2012
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G., Gyorgy, A., and Szepesvári, C · 2012
Cited alongside, same era.
Cryptography from learning parity with noise
Pietrzak, K · 2012
Cited alongside, same era.
Online learning in markov decision processes with adversarially chosen transition probability distributions
Abbasi, Y., Bartlett, P. L., Kanade, V., Seldin, Y., and Szepesvári, C · 2013
Cited alongside, same era.
Better rates for any adversarial deterministic mdp
Dekel, O. and Hazan, E · 2013
Cited alongside, same era.
Interactive teaching strategies for agent training
Amir, O., Kamar, E., Kolobov, A., and Grosz, B · 2016
Later among the works it cites.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Later among the works it cites.
Corralling a band of bandit algorithms
Agarwal, A., Luo, H., Neyshabur, B., and Schapire, R. E · 2017
Later among the works it cites.
Multi-view decision processes: The helper-ai problem
Dimitrakakis, C., Parkes, D. C., Radanovic, G., and Tylkin, P · 2017
Later among the works it cites.
Game-theoretic modeling of human adaptation in human-robot collaboration
Nikolaidis, S., Nath, S., Procaccia, A. D., and Srinivasa, S · 2017
Later among the works it cites.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization, learning, and games with predictable sequences
Rakhlin, S. and Sridharan, K · 2013
Cited alongside, same era.
On actively teaching the crowd to classify
Singla, A., Bogunovic, I., Bartók, G., Karbasi, A., and Krause, A · 2013
Cited alongside, same era.
Online learning in markov decision processes with changing cost sequences
Dick, T., Gyorgy, A., and Szepesvari, C · 2014
Cited alongside, same era.
Learning hurdles for sleeping experts
Kanade, V. and Steinke, T · 2014
Cited alongside, same era.
Fast convergence of regularized learning in games
Syrgkanis, V., Agarwal, A., Luo, H., and Schapire, R. E · 2015
Cited alongside, same era.
Machine teaching: An inverse problem to machine learning and an approach toward optimal education
Zhu, X · 2015
Cited alongside, same era.
Learning with opponent-learning awareness
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2018
Later among the works it cites.
Modeling others using oneself in multi-agent reinforcement learning
Raileanu, R., Denton, E., Szlam, A., and Fergus, R · 2018
Later among the works it cites.
Prediction with a short memory
Sharan, V., Kakade, S., Liang, P., and Valiant, G · 2018
Later among the works it cites.
Learning to interact with learning agents
Singla, A., Hassani, S. H., and Krause, A · 2018
Later among the works it cites.
An overview of machine teaching
Zhu, X., Singla, A., Zilles, S., and Rafferty, A. N · 2018
Later among the works it cites.
Learning to collaborate in markov decision processes
Radanovic, G., Devidze, R., Parkes, D., and Singla, A · 2019
Closest in time.