Fetching the paper…
Reading the bibliography…
Learning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios.
Calculation of the wasserstein distance between probability distributions on the line
SS Vallender · 1974
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Experiments with reinforcement learning in problems with continuous state and action spaces
Juan C Santamaria, Richard S Sutton, and Ashwin Ram · 1997
Earlier work this paper cites.
Annealed importance sampling
Radford M Neal · 2001
Earlier work this paper cites.
A cooperative multi-agent transportation management and route guidance system
Jeffrey L Adler and Victor J Blue · 2002
Earlier work this paper cites.
Distributed optimization in sensor networks
Michael Rabbat and Robert Nowak · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Achieving controllability of electric loads
Duncan S Callaway and Ian A Hiskens · 2010
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
On the solution of the KKT conditions of generalized nash equilibrium problems
Axel Dreves, Francisco Facchinei, Christian Kanzow, and Simone Sagratella · 2011
Earlier work this paper cites.
On the empirical estimation of integral probability metrics
Bharath K Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, Gert RG Lanckriet, et al · 2012
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Q( λ \lambda ) with off-policy corrections
Rémi Munos · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Frans A Oliehoek, Christopher Amato, et al · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Earlier work this paper cites.
An alternative softmax operator for reinforcement learning
Kavosh Asadi and Michael L Littman · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P Schoellig, and Andreas Krause · 2017
Earlier work this paper cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Earlier work this paper cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
Gregory Kahn, Adam Villaflor, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, et al · 2019
Later among the works it cites.
Revisiting the softmax bellman operator: New benefits and new perspective
Zhao Song, Ron Parr, and Lawrence Carin · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2018
Cited alongside, same era.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Supervised policy update for deep reinforcement learning
Quan Vuong, Yiming Zhang, and Keith W Ross · 2018
Cited alongside, same era.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 2019
Cited alongside, same era.
OPAL: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2020
Later among the works it cites.
Adversarial attacks and detection on reinforcement learning-based interactive recommender systems
Yuanjiang Cao, Xiaocong Chen, Lina Yao, Xianzhi Wang, and Wei Emma Zhang · 2020
Later among the works it cites.
P3O: Policy-on policy-off policy optimization
Rasool Fakoor, Pratik Chaudhari, and Alexander J Smola · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Softmax deep double deterministic policy gradients
Ling Pan, Qingpeng Cai, and Longbo Huang · 2020
Later among the works it cites.
Expert-supervised reinforcement learning for offline policy learning and evaluation
Aaron Sonabend-W, Junwei Lu, Leo A Celi, Tianxi Cai, and Peter Szolovits · 2020
Later among the works it cites.
Randomized value functions via multiplicative normalizing flows
Ahmed Touati, Harsh Satija, Joshua Romoff, Joelle Pineau, and Pascal Vincent · 2020
Later among the works it cites.
Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2020
Later among the works it cites.
Mastering complex control in MOBA games with deep reinforcement learning
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, et al · 2020
Later among the works it cites.
Celebrating diversity in shared multi-agent reinforcement learning
Chenghao Li, Chengjie Wu, Tonghan Wang, Jun Yang, Qianchuan Zhao, and Chongjie Zhang · 2021
Closest in time.
Modeling the interaction between agents in cooperative multi-agent reinforcement learning
Xiaoteng Ma, Yiqin Yang, Chenghao Li, Yiwen Lu, Qianchuan Zhao, and Jun Yang · 2021
Closest in time.