Emergent complexity via multi-agent competition
Original
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I. (2017) · 2017
Later among the works it cites.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I. (2017) · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., and Bolton, A. (2017) · 2017
Later among the works it cites.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J. (2017) · 2017
Later among the works it cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Wen, Z. and Van Roy, B. (2017) · 2017
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T. (2018) · 2018
Later among the works it cites.
On oracle-efficient PAC rl with rich observations
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2018) · 2018
Later among the works it cites.
Balancing two-player stochastic games with soft q-learning
Original
Grau-Moya, J., Leibfried, F., and Bou-Ammar, H. (2018) · 2018
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Original
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P. A., Strouse, D., Leibo, J. Z., and De Freitas, N. (2018) · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2018) · 2018
Later among the works it cites.
OpenAI Five
OpenAI (2018) · 2018
Later among the works it cites.
Actor-critic fictitious play in simultaneous move multistage games
Perolat, J., Piot, B., and Pietquin, O. (2018) · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T. (2019) · 2019
Later among the works it cites.
Online stochastic shortest path with bandit feedback and unknown transition function
Rosenberg, A. and Mansour, Y. (2019) · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Russo, D. (2019) · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Simchowitz, M. and Jamieson, K. G. (2019) · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D. (2019) · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E. (2019) · 2019
Later among the works it cites.
The leave-one-out approach for matrix completion: Primal and dual analysis
Original
Ding, L. and Chen, Y. (2020) · 2020
Closest in time.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F. (2020) · 2020
Closest in time.