Fetching the paper…
Reading the bibliography…
Existing Maximum-Entropy (MaxEnt) Reinforcement Learning (RL) methods for continuous action spaces are typically formulated based on actor-critic frameworks and optimized through alternating steps of policy evaluation and policy improvement.
Monte Carlo methods. Vol. 1: basics
Malvin H. Kalos and Paula A. Whitlock · 1986
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Exponential convergence of Langevin distributions and their discrete approximations
Gareth O. Roberts and Richard L. Tweedie · 1996
Earlier work this paper cites.
Optimal scaling of discrete approximations to Langevin diffusions
Gareth O. Roberts and Jeffrey S. Rosenthal · 1998
Earlier work this paper cites.
Learning to Drive a Bicycle Using Reinforcement Learning and Shaping
Jette Randløv and Preben Alstrøm · 1998
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
A. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Independent Component Analysis: Algorithms and Applications
A. Hyvärinen and E. Oja · 2000
Earlier work this paper cites.
Energy-based models for sparse overcomplete representations
Yee Whye Teh, Max Welling, Simon Osindero, and Geoffrey E. Hinton · 2003
Earlier work this paper cites.
Theory and Application of Reward Shaping in Reinforcement Learning
Adam Laud · 2004
Earlier work this paper cites.
Path Integrals and Symmetry Breaking for Optimal Control Theory
H J Kappen · 2005
Earlier work this paper cites.
Estimation of Non-Normalized Statistical Models by Score Matching
A. Hyvärinen · 2005
Earlier work this paper cites.
A Tutorial on Energy-Based Learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Aurelio Ranzato, and Fu Jie Huang · 2006
Earlier work this paper cites.
A Modern Introduction to Probability and Statistics
Michel Dekking · 2007
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian Ziebart, Andrew Maas, J. Bagnell, and Anind Dey · 2008
Earlier work this paper cites.
Robot Trajectory Optimization using Approximate Inference
Marc Toussaint · 2009
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
Importance Sampling: A Review
Surya T. Tokdar and Robert E. Kass · 2010
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
On Stochastic Optimal Control and Reinforcement Learning by Approximate Inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Multilevel Monte Carlo Methods
Michael B. Giles · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
MADE: Masked Autoencoder for Distribution Estimation
Mathieu Germain, Karol Gregor, Iain Murray, and H. Larochelle · 2015
Earlier work this paper cites.
NICE: Non-linear Independent Components Estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2015
Cited alongside, same era.
Human-level Control through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Kirkeby Fidjeland, Georg Ostrovski, Stig Petersen, Charlie Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Taming the Noise in Reinforcement Learning via Soft Updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm
Qiang Liu and Dilin Wang · 2016
Cited alongside, same era.
Learning to Draw Samples: With Application to Amortized MLE for Generative Adversarial Learning
Emerging Convolutions for Generative Normalizing Flows
Emiel Hoogeboom, Rianne van den Berg, and Max Welling · 2019
Later among the works it cites.
MaCow: Masked Convolutional Generative Flow
Xuezhe Ma and Eduard H. Hovy · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Woodbury Transformations for Deep Generative Flows
You Lu and Bert Huang · 2020
Later among the works it cites.
Relative Gradient Optimization of the Jacobian Term in Unsupervised Deep Learning
L. Gresele, G. Fissore, A. Javaloy, B. Schölkopf, and A. Hyvärinen · 2020
Later among the works it cites.
Self Normalizing Flows
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dilin Wang and Qiang Liu · 2016
Cited alongside, same era.
Improved Variational Inference with Inverse Autoregressive Flow
Diederik P. Kingma, Tim Salimans, and Max Welling · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
PGQ: Combining Policy Gradient and Q-learning
Brendan O’Donoghue, Rémi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Cited alongside, same era.
Masked Autoregressive Flow for Density Estimation
George Papamakarios, Iain Murray, and Theo Pavlakou · 2017
Cited alongside, same era.
Density Estimation using Real NVP
Laurent Dinh, Jascha Narain Sohl-Dickstein, and Samy Bengio · 2017
Cited alongside, same era.
T. Anderson Keller, Jorn W. T. Peters, Priyank Jaini, Emiel Hoogeboom, Patrick Forr’e, and Max Welling · 2020
Later among the works it cites.
Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without Sampling
Will Grathwohl, Kuan-Chieh Jackson Wang, Jörn-Henrik Jacobsen, David Kristjanson Duvenaud, and Richard S. Zemel · 2020
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2021
Later among the works it cites.
Soft Actor-Critic for Navigation of Mobile Robots
Junior Costa de Jesus, Victor Augusto Kich, Alisson Henrique Kolling, Ricardo Bedin Grando, Marco Antonio de Souza Leite Cuadros, and Daniel Fernando Tello Gamarra · 2021
Later among the works it cites.
Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, N. Rudin, Arthur Allshire, Ankur Handa, and Gavriel State · 2021
Later among the works it cites.
A Statistical Analysis of Polyak-Ruppert Averaged Q-Learning
Xiang Li, Wenhao Yang, Jiadong Liang, Zhihua Zhang, and Michael I. Jordan · 2021
Later among the works it cites.
Stable-Baselines3: Reliable Reinforcement Learning Implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Maximum Entropy RL Provably Solves Some Robust RL Problems
Benjamin Eysenbach and Sergey Levine · 2022
Later among the works it cites.
Path Planning for Multi-Arm Manipulators Using Soft Actor-Critic Algorithm with Position Prediction of Moving Obstacles via LSTM
Kwan-Woo Park, MyeongSeop Kim, Jung-Su Kim, and Jae-Han Park · 2022
Later among the works it cites.
Robot Skill Adaptation via Soft Actor-Critic Gaussian Mixture Models
Iman Nematollahi, Erick Rosete-Beas, Adrian Roefer, Tim Welschehold, Abhinav Valada, and Wolfram Burgard · 2022
Later among the works it cites.
ButterflyFlow: Building Invertible Layers with Butterfly Matrices
Chenlin Meng, Linqi Zhou, Kristy Choi, Tri Dao, and Stefano Ermon · 2022
Later among the works it cites.
Optimistic Curiosity Exploration and Conservative Exploitation with Linear Reward Shaping
Hao Sun, Lei Han, Rui Yang, Xiaoteng Ma, Jian Guo, and Bolei Zhou · 2022
Later among the works it cites.
CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms
Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and João G.M. Araújo · 2022
Later among the works it cites.
Latent State Marginalization as a Low-cost Approach to Improving Exploration
Dinghuai Zhang, Aaron Courville, Yoshua Bengio, Qinqing Zheng, Amy Zhang, and Ricky T. Q. Chen · 2023
Later among the works it cites.
Training Energy-Based Normalizing Flow with Score-Matching Objectives
Chen-Hao Chao, Wei-Fang Sun, Yen-Chang Hsu, Zsolt Kira, and Chun-Yi Lee · 2023
Later among the works it cites.
Gymnasium, 2023
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis · 2023
Later among the works it cites.
Adaptive reward shifting based on behavior proximity for offline reinforcement learning
Zhe Zhang and Xiaoyang Tan · 2023
Later among the works it cites.
normflows: A PyTorch Package for Normalizing Flows
Vincent Stimper, David Liu, Andrew Campbell, Vincent Berenz, Lukas Ryll, Bernhard Schölkopf, and José Miguel Hernández-Lobato · 2023
Later among the works it cites.
skrl: Modular and Flexible Library for Reinforcement Learning
Antonio Serrano-Muñoz, Dimitrios Chrysostomou, Simon Bøgh, and Nestor Arana-Arexolaleiba · 2023
Later among the works it cites.
S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor Critic
Safa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang, Bo An, Haipeng Chen, and Sanjay Chawla · 2024
Closest in time.