Fetching the paper…
Reading the bibliography…
In partially observable reinforcement learning, offline training gives access to latent information which is not available during online training and/or execution, such as the system state.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning . PMLR, 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J. Williams and Jing Peng. 1991 · 1991
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
Satinder P. Singh, Tommi Jaakkola, and Michael I. Jordan. 1994 · 1994
Earlier work this paper cites.
Solving large POMDPs using real time dynamic programming. In AAAI Fall Symposium on POMDPs
Blai Bonet. 1998 · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. 1998 · 1998
Earlier work this paper cites.
Actor-critic algorithms. In Advances in Neural Information Processing Systems . 1008–1014
Vijay R. Konda and John N. Tsitsiklis. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems . 1057–1063
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
R-maddpg for partially observable environments and limited communication
Rose E. Wang, Michael Everett, and Jonathan P. How. 2020 · 2002
Earlier work this paper cites.
Deep Multi-Agent Reinforcement Learning for Decentralized Continuous Cooperative Control
Christian Schroeder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. 2021 · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter. 2004 · 2004
Earlier work this paper cites.
Weighted QMIX: Expanding Monotonic Value Function Factorisation
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson. 2020 · 2006
Earlier work this paper cites.
Belief-Grounded Networks for AcceleratedRobot Learning under Partial Observability
Hai Nguyen, Brett Daley, Xinchao Song, Chistopher Amato, and Robert Platt. 2020 · 2010
Earlier work this paper cites.
Thomas Degris, Martha White, and Richard S. Sutton. 2012 · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms. In International conference on machine learning . PMLR, 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Later among the works it cites.
Cm3: Cooperative multi-goal multi-stage multi-agent reinforcement learning
Jiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura, and Hongyuan Zha. 2018 · 2018
Later among the works it cites.
gym-pomdps: Gym environments from POMDP files
Andrea Baisero. 2019 · 2019
Later among the works it cites.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 4213–4220
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell. 2019 · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration. In Advances in Neural Information Processing Systems . 7613–7624
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with deep energy-based policies. In International Conference on Machine Learning , Vol. 70. PMLR
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems . 6379–6390
Ryan Lowe, Yi I. Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Asymmetric actor critic for image-based robot learning
Lerrel Pinto, Marcin Andrychowicz, Peter Welinder, Wojciech Zaremba, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas. 2017 · 2017
Cited alongside, same era.
On the Study of Cooperative Multi-Agent Policy Gradient
Guillaume Bono, Jilles Dibangoye, Laëtitia Matignon, Florian Pereyron, and Olivier Simonin. 2018 · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Differentiable particle filters: End-to-end learning with algorithmic priors
Rico Jonschkowski, Divyam Rastogi, and Oliver Brock. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Learning by cheating. In Conference on Robot Learning . PMLR, 66–75
Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl. 2020 · 2020
Later among the works it cites.
gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning
Andrea Baisero and Sammie Katt. 2021 · 2021
Closest in time.
Multi-agent reinforcement learning with directed exploration and selective memory reuse. In Proceedings of the ACM Symposium on Applied Computing . 777–784
Shuo Jiang and Christopher Amato. 2021 · 2021
Closest in time.
Pomdp Robot Domains
Hai Nguyen. 2021 · 2021
Closest in time.
Robust Asymmetric Learning in POMDPs
Andrew Warrington, J. Wilder Lavington, Adam Scibior, Mark Schmidt, and Frank Wood. 2021 · 2021
Closest in time.
Local Advantage Actor-Critic for Robust Multi-Agent Deep Reinforcement Learning. In International Symposium on Multi-Robot and Multi-Agent Systems . IEEE, 155–163
Yuchen Xiao, Xueguang Lyu, and Christopher Amato. 2021 · 2021
Closest in time.
Unbiased Asymmetric Reinforcement Learning under Partial Observability
Andrea Baisero and Christopher Amato. 2022 · 2022
Closest in time.