A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, R. Hadsell, Policy Distillation, in: International Conference on Learning Representations, 2016
2016
Later among the works it cites.
P. Hernandez-Leal, M. E. Taylor, B. Rosman, L. E. Sucar, E. Munoz de Cote, Identifying and Tracking Switching, Non-stationary Opponents: a Bayesian Approach, in: Multiagent Interaction without Prior Coordination Workshop at AAAI, Phoenix, AZ, USA, 2016
2016
Later among the works it cites.
B. Rosman, M. Hawasly, S. Ramamoorthy, Bayesian Policy Reuse, Machine Learning 104 (1) (2016) 99–127
2016
Later among the works it cites.
F. A. Oliehoek, C. Amato, et al., A concise introduction to decentralized POMDPs, Springer, 2016
2016
Later among the works it cites.
M. Johnson, K. Hofmann, T. Hutton, D. Bignell, The Malmo platform for artificial intelligence experimentation., in: IJCAI, 2016, pp. 4246–4247
2016
Later among the works it cites.
Deep Reinforcement Learning: Pong from Pixels, https://karpathy.github.io/2016/05/31/rl/ , [Online; accessed 7-May-2019] (2016)
2016
Later among the works it cites.
T. D. Kulkarni, K. Narasimhan, A. Saeedi, J. Tenenbaum, Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation, in: Advances in neural information processing systems, 2016, pp. 3675–3683
2016
Later among the works it cites.
B. Kartal, E. Nunes, J. Godoy, M. Gini, Monte Carlo tree search with branch and bound for multi-robot task allocation, in: The IJCAI-16 Workshop on Autonomous Mobile Service Robots, 2016
2016
Later among the works it cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al., Mastering the game of Go without human knowledge, Nature 550 (7676) (2017) 354
2017
Later among the works it cites.
M. Moravčík, M. Schmid, N. Burch, V. Lisý, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, M. Bowling, DeepStack: Expert-level artificial intelligence in heads-up no-limit poker, Science 356 (6337) (2017) 508–513
2017
Later among the works it cites.
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, S. Levine, Q-prop: Sample-efficient policy gradient with an off-policy critic, in: International Conference on Learning Representations, 2017
2017
Later among the works it cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, I. Mordatch, Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments., in: Advances in Neural Information Processing Systems, 2017, pp. 6379–6390
2017
Later among the works it cites.
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, K. Kavukcuoglu, Reinforcement Learning with Unsupervised Auxiliary Tasks., in: International Conference on Learning Representations, 2017
2017
Later among the works it cites.
OpenAI Baselines: ACKTR & A2C, https://openai.com/blog/baselines-acktr-a2c/ , [Online; accessed 29-April-2019] (2017)
2017
Later among the works it cites.
S. S. Gu, T. Lillicrap, R. E. Turner, Z. Ghahramani, B. Schölkopf, S. Levine, Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning, in: Advances in Neural Information Processing Systems, 2017, pp. 3846–3855
2017
Later among the works it cites.
T. Haarnoja, H. Tang, P. Abbeel, S. Levine, Reinforcement learning with deep energy-based policies, in: Proceedings of the 34th International Conference on Machine Learning-Volume 70, 2017, pp. 1352–1361
2017
Later among the works it cites.
Multiagent Learning, Foundations and Recent Trends, https://www.cs.utexas.edu/~larg/ijcai17_tutorial/multiagent_learning.pdf , [Online; accessed 7-September-2018] (2017)
2017
Later among the works it cites.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, R. Vicente, Multiagent cooperation and competition with deep reinforcement learning, PLOS ONE 12 (4) (2017) e0172395
2017
Later among the works it cites.
J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, Multi-agent Reinforcement Learning in Sequential Social Dilemmas, in: Proceedings of the 16th Conference on Autonomous Agents and Multiagent Systems, Sao Paulo, 2017
2017
Later among the works it cites.
A. Lazaridou, A. Peysakhovich, M. Baroni, Multi-Agent Cooperation and the Emergence of (Natural) Language, in: International Conference on Learning Representations, 2017
2017
Later among the works it cites.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, J. Vian, Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability, in: Proceedings of the 34th International Conference on Machine Learning, Sydney, 2017
2017
Later among the works it cites.
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, S. Whiteson, Counterfactual Multi-Agent Policy Gradients., in: 32nd AAAI Conference on Artificial Intelligence, 2017
2017
Later among the works it cites.
J. N. Foerster, N. Nardelli, G. Farquhar, T. Afouras, P. H. S. Torr, P. Kohli, S. Whiteson, Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning., in: International Conference on Machine Learning, 2017
2017
Later among the works it cites.
M. Lanctot, V. F. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Pérolat, D. Silver, T. Graepel, A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning., in: Advances in Neural Information Processing Systems, 2017
2017
Later among the works it cites.
J. K. Gupta, M. Egorov, M. Kochenderfer, Cooperative multi-agent control using deep reinforcement learning, in: G. Sukthankar, J. A. Rodriguez-Aguilar (Eds.), Autonomous Agents and Multiagent Systems, Springer International Publishing, Cham, 2017, pp. 66–83
2017
Later among the works it cites.
K. A. Ciosek, S. Whiteson, Offer: Off-environment reinforcement learning, in: Thirty-First AAAI Conference on Artificial Intelligence, 2017
2017
Later among the works it cites.
L. Pinto, J. Davidson, R. Sukthankar, A. Gupta, Robust adversarial reinforcement learning, in: Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 2817–2826
2017
Later among the works it cites.
S. Damer, M. Gini, Safely using predictions in general-sum normal form games, in: Proceedings of the 16th Conference on Autonomous Agents and Multiagent Systems, Sao Paulo, 2017
2017
Later among the works it cites.
P. Hernandez-Leal, M. Kaisers, Towards a Fast Detection of Opponents in Repeated Stochastic Games, in: G. Sukthankar, J. A. Rodriguez-Aguilar (Eds.), Autonomous Agents and Multiagent Systems: AAMAS 2017 Workshops, Best Papers, Sao Paulo, Brazil, May 8-12, 2017, Revised Selected Papers, 2017, pp. 239–257
2017
Later among the works it cites.
P. Hernandez-Leal, Y. Zhan, M. E. Taylor, L. E. Sucar, E. Munoz de Cote, Efficiently detecting switches against non-stationary opponents, Autonomous Agents and Multi-Agent Systems 31 (4) (2017) 767–789
2017
Later among the works it cites.
P. Hernandez-Leal, M. Kaisers, Learning against sequential opponents in repeated stochastic games, in: The 3rd Multi-disciplinary Conference on Reinforcement Learning and Decision Making, Ann Arbor, 2017
2017
Later among the works it cites.
Do I really have to cite an arXiv paper?, http://approximatelycorrect.com/2017/08/01/do-i-have-to-cite-arxiv-paper/ , [Online; accessed 21-May-2019] (2017)
2017
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, W. Zaremba, Hindsight experience replay, in: Advances in Neural Information Processing Systems, 2017
2017
Later among the works it cites.
J. K. Gupta, M. Egorov, M. J. Kochenderfer, Cooperative Multi-agent Control using deep reinforcement learning, in: Adaptive Learning Agents at AAMAS, Sao Paulo, 2017
2017
Later among the works it cites.
K. Greff, R. K. Srivastava, J. Koutnik, B. R. Steunebrink, J. Schmidhuber, LSTM: A Search Space Odyssey, IEEE Transactions on Neural Networks and Learning Systems 28 (10) (2017) 2222–2232
2017
Later among the works it cites.
M. Babaeizadeh, I. Frosio, S. Tyree, J. Clemons, J. Kautz, Reinforcement learning through asynchronous advantage actor-critic on a GPU, in: International Conference on Learning Representations, 2017
2017
Later among the works it cites.
T. Vodopivec, S. Samothrakis, B. Ster, On Monte Carlo tree search and reinforcement learning, Journal of Artificial Intelligence Research 60 (2017) 881–936
2017
Later among the works it cites.
S. V. Albrecht, P. Stone, Autonomous agents modelling other agents: A comprehensive survey and open problems, Artificial Intelligence 258 (2018) 66–95
2018
Closest in time.
N. Brown, T. Sandholm, Superhuman AI for heads-up no-limit poker: Libratus beats top professionals, Science 359 (6374) (2018) 418–424
2018
Closest in time.
Open AI Five, https://blog.openai.com/openai-five , [Online; accessed 7-September-2018] (2018)
2018
Closest in time.
R. S. Sutton, A. G. Barto, Reinforcement learning: An introduction, 2nd Edition, MIT Press, 2018
2018
Closest in time.
V. François-Lavet, P. Henderson, R. Islam, M. G. Bellemare, J. Pineau, et al., An introduction to deep reinforcement learning, Foundations and Trends® in Machine Learning 11 (3-4) (2018) 219–354
2018
Closest in time.
Y. Yang, J. Hao, M. Sun, Z. Wang, C. Fan, G. Strbac, Recurrent Deep Multiagent Q-Learning for Autonomous Brokers in Smart Grid, in: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 2018
2018
Closest in time.
J. Zhao, G. Qiu, Z. Guan, W. Zhao, X. He, Deep reinforcement learning for sponsored search real-time bidding, in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ACM, 2018, pp. 1021–1030
2018
Closest in time.
G. Palmer, K. Tuyls, D. Bloembergen, R. Savani, Lenient Multi-Agent Deep Reinforcement Learning., in: International Conference on Autonomous Agents and Multiagent Systems, 2018
2018
Closest in time.
A. Darwiche, Human-level intelligence or animal-like abilities?, Commun. ACM 61 (10) (2018) 56–67
2018
Closest in time.
H. Liu, Y. Feng, Y. Mao, D. Zhou, J. Peng, Q. Liu, Action-depedent control variates for policy optimization via stein’s identity, in: International Conference on Learning Representations, 2018
2018
Closest in time.
G. Tucker, S. Bhupatiraju, S. Gu, R. E. Turner, Z. Ghahramani, S. Levine, The mirage of action-dependent baselines in reinforcement learning, in: International Conference on Machine Learning, 2018
2018
Closest in time.
J. N. Foerster, R. Y. Chen, M. Al-Shedivat, S. Whiteson, P. Abbeel, I. Mordatch, Learning with Opponent-Learning Awareness., in: Proceedings of 17th International Conference on Autonomous Agents and Multiagent Systems, Stockholm, Sweden, 2018
2018
Closest in time.
S. Fujimoto, H. van Hoof, D. Meger, Addressing function approximation error in actor-critic methods, in: International Conference on Machine Learning, 2018
2018
Closest in time.
T. Lu, D. Schuurmans, C. Boutilier, Non-delusional Q-learning and value-iteration, in: Advances in Neural Information Processing Systems, 2018, pp. 9949–9959
2018
Closest in time.
D. Isele, A. Cosgun, Selective experience replay for lifelong learning, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
D. Steckelmacher, D. M. Roijers, A. Harutyunyan, P. Vrancx, H. Plisnier, A. Nowé, Reinforcement learning in pomdps with memoryless options and option-observation initiation sets, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al., IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures, in: International Conference on Machine Learning, 2018
2018
Closest in time.
T. Haarnoja, A. Zhou, P. Abbeel, S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, in: International Conference on Machine Learning, 2018
2018
Closest in time.
D. Balduzzi, S. Racaniere, J. Martens, J. Foerster, K. Tuyls, T. Graepel, The mechanics of n-player differentiable games, in: Proceedings of the 35th International Conference on Machine Learning, Proceedings of Machine Learning Research, Stockholm, Sweden, 2018, pp. 354–363
2018
Closest in time.
J. Pérolat, B. Piot, O. Pietquin, Actor-critic fictitious play in simultaneous move multistage games, in: 21st International Conference on Artificial Intelligence and Statistics, 2018
2018
Closest in time.
G. Bono, J. S. Dibangoye, L. Matignon, F. Pereyron, O. Simonin, Cooperative multi-agent policy gradient, in: European Conference on Machine Learning, 2018
2018
Closest in time.
S. Srinivasan, M. Lanctot, V. Zambaldi, J. Pérolat, K. Tuyls, R. Munos, M. Bowling, Actor-critic policy optimization in partially observable multiagent environments, in: Advances in Neural Information Processing Systems, 2018, pp. 3422–3435
2018
Closest in time.
M. Raghu, A. Irpan, J. Andreas, R. Kleinberg, Q. Le, J. Kleinberg, Can Deep Reinforcement Learning solve Erdos-Selfridge-Spencer Games?, in: Proceedings of the 35th International Conference on Machine Learning, 2018
2018
Closest in time.
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, I. Mordatch, Emergent Complexity via Multi-Agent Competition., in: International Conference on Machine Learning, 2018
2018
Closest in time.
I. Mordatch, P. Abbeel, Emergence of grounded compositional language in multi-agent populations, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
R. Raileanu, E. Denton, A. Szlam, R. Fergus, Modeling Others using Oneself in Multi-Agent Reinforcement Learning., in: International Conference on Machine Learning, 2018
2018
Closest in time.
Z.-W. Hong, S.-Y. Su, T.-Y. Shann, Y.-H. Chang, C.-Y. Lee, A Deep Policy Inference Q-Network for Multi-Agent Systems, in: International Conference on Autonomous Agents and Multiagent Systems, 2018
2018
Closest in time.
N. C. Rabinowitz, F. Perbet, H. F. Song, C. Zhang, S. M. A. Eslami, M. Botvinick, Machine Theory of Mind., in: International Conference on Machine Learning, Stockholm, Sweden, 2018
2018
Closest in time.
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. F. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, T. Graepel, Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward., in: Proceedings of 17th International Conference on Autonomous Agents and Multiagent Systems, Stockholm, Sweden, 2018
2018
Closest in time.
T. Rashid, M. Samvelyan, C. S. de Witt, G. Farquhar, J. N. Foerster, S. Whiteson, QMIX - Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning., in: International Conference on Machine Learning, 2018
2018
Closest in time.
Y. Zheng, Z. Meng, J. Hao, Z. Zhang, T. Yang, C. Fan, A Deep Bayesian Policy Reuse Approach Against Non-Stationary Agents, in: Advances in Neural Information Processing Systems, 2018, pp. 962–972
2018
Closest in time.
Capture the Flag: the emergence of complex cooperative agents, https://deepmind.com/blog/capture-the-flag/ , [Online; accessed 7-September-2018] (2018)
2018
Closest in time.
T. De Bruin, J. Kober, K. Tuyls, R. Babuška, Experience selection in deep reinforcement learning for control, The Journal of Machine Learning Research 19 (1) (2018) 347–402
2018
Closest in time.
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, M. Bowling, Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents, Journal of Artificial Intelligence Research 61 (2018) 523–562
2018
Closest in time.
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, D. Silver, Rainbow: Combining improvements in deep reinforcement learning, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, D. Meger, Deep Reinforcement Learning That Matters., in: 32nd AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
K. Clary, E. Tosch, J. Foley, D. Jensen, Let’s play again: Variability of deep reinforcement learning agents in Atari environments, in: NeurIPS Critiquing and Correcting Trends Workshop, 2018
2018
Closest in time.
Z. C. Lipton, J. Steinhardt, Troubling trends in machine learning scholarship, in: ICML Machine Learning Debates workshop, 2018
2018
Closest in time.
D. Sculley, J. Snoek, A. Wiltschko, A. Rahimi, Winner’s curse? on pace, progress, and empirical rigor, in: ICLR Workshop, 2018
2018
Closest in time.
K. Azizzadenesheli, B. Yang, W. Liu, E. Brunskill, Z. Lipton, A. Anandkumar, Surprising negative results for generative adversarial tree search, in: Critiquing and Correcting Trends in Machine Learning Workshop, 2018
2018
Closest in time.
G. Melis, C. Dyer, P. Blunsom, On the state of the art of evaluation in neural language models, in: International Conference on Learning Representations, 2018
2018
Closest in time.
D. Amodei, D. Hernandez, AI and Compute (2018). URL https://blog.openai.com/ai-and-compute
2018
Closest in time.
Y. Yu, Towards sample efficient reinforcement learning., in: IJCAI, 2018, pp. 5739–5743
2018
Closest in time.
Y. Yang, R. Luo, M. Li, M. Zhou, W. Zhang, J. Wang, Mean field multi-agent reinforcement learning, in: Proceedings of the 35th International Conference on Machine Learning, Stockholm Sweden, 2018
2018
Closest in time.
A. Grover, M. Al-Shedivat, J. K. Gupta, Y. Burda, H. Edwards, Learning Policy Representations in Multiagent Systems., in: International Conference on Machine Learning, 2018
2018
Closest in time.
C. K. Ling, F. Fang, J. Z. Kolter, What game are we playing? end-to-end learning in normal and extensive form games, in: Twenty-Seventh International Joint Conference on Artificial Intelligence, 2018
2018
Closest in time.
F. L. Silva, A. H. R. Costa, A survey on transfer learning for multiagent reinforcement learning systems, Journal of Artificial Intelligence Research 64 (2019) 645–703
2019
Closest in time.
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. M. Czarnecki, A. Dudzik, A. Huang, P. Georgiev, R. Powell, T. Ewalds, D. Horgan, M. Kroiss, I. Danihelka, J. Agapiou, J. Oh, V. Dalibard, D. Choi, L. Sifre, Y. Sulsky, S. Vezhnevets, J. Molloy, T. Cai, D. Budden, T. Paine, C. Gulcehre, Z. Wang, T. Pfaff, T. Pohlen, Y. Wu, D. Yogatama, J. Cohen, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, C. Apps, K. Kavukcuoglu, D. Hassabis, D. Silver, AlphaStar: Mastering the Real-Time Strategy Game StarCraft II, https://deepmind.com/blog/alphastar-mastering-real-time-strategy-game-starcraft-ii/ (2019)
2019
Closest in time.
G. Palmer, R. Savani, K. Tuyls, Negative update intervals in deep multi-agent reinforcement learning, in: 18th International Conference on Autonomous Agents and Multiagent Systems, 2019
2019
Closest in time.
G. Bacchiani, D. Molinari, M. Patander, Microscopic traffic simulation by cooperative multi-agent deep reinforcement learning, in: AAMAS, 2019
2019
Closest in time.
X. Song, T. Wang, C. Zhang, Convergence of multi-agent learning with a finite step size in general-sum games, in: 18th International Conference on Autonomous Agents and Multiagent Systems, 2019
2019
Closest in time.
J. Z. Leibo, J. Perolat, E. Hughes, S. Wheelwright, A. H. Marblestone, E. Duéñez-Guzmán, P. Sunehag, I. Dunning, T. Graepel, Malthusian reinforcement learning, in: 18th International Conference on Autonomous Agents and Multiagent Systems, 2019
2019
Closest in time.
T. Yang, J. Hao, Z. Meng, C. Zhang, Y. Z. Z. Zheng, Towards Efficient Detection and Optimal Response against Sophisticated Opponents, in: IJCAI, 2019
2019
Closest in time.
W. Kim, M. Cho, Y. Sung, Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning, in: 33rd AAAI Conference on Artificial Intelligence, 2019
2019
Closest in time.
arXiv:https://science.sciencemag.org/content/364/6443/859.full.pdf
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castañeda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, N. Sonnerat, T. Green, L. Deason, J. Z. Leibo, D. Silver, D. Hassabis, K. Kavukcuoglu, T. Graepel, Human-level performance in 3d multiplayer games with population-based reinforcement learning , Science 364 (6443) (2019) 859–865 · 2019
Closest in time.
S. Li, Y. Wu, X. Cui, H. Dong, F. Fang, S. Russell, Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient, in: AAAI Conference on Artificial Intelligence, 2019
2019
Closest in time.
R. Lowe, J. Foerster, Y.-L. Boureau, J. Pineau, Y. Dauphin, On the pitfalls of measuring emergent communication, in: 18th International Conference on Autonomous Agents and Multiagent Systems, 2019
2019
Closest in time.
J. Castellini, F. A. Oliehoek, R. Savani, S. Whiteson, The Representational Capacity of Action-Value Networks for Multi-Agent Reinforcement Learning, in: 18th International Conference on Autonomous Agents and Multiagent Systems, 2019
2019
Closest in time.
Collaboration & Credit Principles, How can we be good stewards of collaborative trust?, http://colah.github.io/posts/2019-05-Collaboration/index.html , [Online; accessed 31-May-2019] (2019)
2019
Closest in time.
P. Hernandez-Leal, B. Kartal, M. E. Taylor, Agent Modeling as Auxiliary Task for Deep Reinforcement Learning, in: AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 2019
2019
Closest in time.
C. Gao, B. Kartal, P. Hernandez-Leal, M. E. Taylor, On Hard Exploration for Reinforcement Learning: a Case Study in Pommerman, in: AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 2019
2019
Closest in time.
S. Liu, G. Lever, J. Merel, S. Tunyasuvunakool, N. Heess, T. Graepel, Emergent coordination through competition, in: International Conference on Learning Representations, 2019
2019
Closest in time.
J. Z. Forde, M. Paganini, The scientific method in the science of machine learning, in: ICLR Debugging Machine Learning Models workshop, 2019
2019
Closest in time.
K. Azizzadenesheli, Maybe a few considerations in reinforcement learning research?, in: Reinforcement Learning for Real Life Workshop, 2019
2019
Closest in time.
C. Lyle, P. S. Castro, M. G. Bellemare, A comparative analysis of expected and distributional reinforcement learning, in: Thirty-Third AAAI Conference on Artificial Intelligence, 2019
2019
Closest in time.
B. Kartal, P. Hernandez-Leal, M. E. Taylor, Using Monte Carlo tree search as a demonstrator within asynchronous deep RL, in: AAAI Workshop on Reinforcement Learning in Games, 2019
2019
Closest in time.
C. Gao, P. Hernandez-Leal, B. Kartal, M. E. Taylor, Skynet: A Top Deep RL Agent in the Inaugural Pommerman Team Competition, in: 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making, 2019
2019
Closest in time.
G. Cuccu, J. Togelius, P. Cudré-Mauroux, Playing Atari with six neurons, in: Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 998–1006
2019
Closest in time.
G. Best, O. M. Cliff, T. Patten, R. R. Mettu, R. Fitch, Dec-MCTS: Decentralized planning for multi-robot active perception, The International Journal of Robotics Research 38 (2-3) (2019) 316–337
2019
Closest in time.
M. Suau de Castro, E. Congeduti, R. A. Starre, A. Czechowski, F. A. Oliehoek, Influence-based abstraction in deep reinforcement learning, in: Adaptive, Learning Agents workshop, 2019
2019
Closest in time.
C. Szepesvári, M. L. Littman, A unified analysis of value-function-based reinforcement-learning algorithms, Neural computation 11 (8) (1999) 2017–2060
2060
Closest in time.