Noisy networks for exploration
Original
Fortunato, Meire, Azar, Mohammad Gheshlaghi, Piot, Bilal, Menick, Jacob, Osband, Ian, Graves, Alex, Mnih, Vlad, Munos, Remi, Hassabis, Demis, Pietquin, Olivier, et al · 2017
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Original
Grathwohl, Will, Choi, Dami, Wu, Yuhuai, Roeder, Geoff, and Duvenaud, David · 2017
Later among the works it cites.
Controlled sequential monte carlo
Original
Heng, Jeremy, Bishop, Adrian N, Deligiannidis, George, and Doucet, Arnaud · 2017
Later among the works it cites.
Efficient structured inference for stochastic recurrent neural networks
Liu, Hao, He, Lirong, Bai, Haoli, and Xu, Zenglin · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, Ryan, Wu, Yi, Tamar, Aviv, Harb, Jean, Abbeel, Pieter, and Mordatch, Igor · 2017
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
Original
Peng, Xue Bin, Andrychowicz, Marcin, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Parameter space noise for exploration
Original
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y, Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, George, Mnih, Andriy, Maddison, Chris J, Lawson, John, and Sohl-Dickstein, Jascha · 2017
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Original
Arjona-Medina, Jose A, Gillhofer, Michael, Widrich, Michael, Unterthiner, Thomas, and Hochreiter, Sepp · 2018
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Original
Buesing, Lars, Weber, Theophane, Racaniere, Sebastien, Eslami, SM, Rezende, Danilo, Reichert, David P, Viola, Fabio, Besse, Frederic, Gregor, Karol, Hassabis, Demis, et al · 2018
Later among the works it cites.
Implicit reparameterization gradients
Original
Figurnov, Michael, Mohamed, Shakir, and Mnih, Andriy · 2018
Later among the works it cites.
Temporal difference variational auto-encoder
Original
Gregor, Karol, Papamakarios, George, Besse, Frederic, Buesing, Lars, and Weber, Theophane · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Original
Haarnoja, Tuomas, Zhou, Aurick, Abbeel, Pieter, and Levine, Sergey · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Original
Igl, Maximilian, Zintgraf, Luisa, Le, Tuan Anh, Wood, Frank, and Whiteson, Shimon · 2018
Later among the works it cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Original
Kosiorek, Adam R, Kim, Hyunjik, Posner, Ingmar, and Teh, Yee Whye · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Original
Levine, Sergey · 2018
Later among the works it cites.
Neural belief states for partially observed domains
Moreno, Pol, Humplik, Jan, Papamakarios, George, Avila Pires, Bernardo, Buesing, Lars, Heess, Nicolas, and Weber, Theophane · 2018
Later among the works it cites.
Variational bayes with synthetic likelihood
Ong, Victor MH, Nott, David J, Tran, Minh-Ngoc, Sisson, Scott A, and Drovandi, Christopher C · 2018
Later among the works it cites.
Total stochastic gradient algorithms and applications in reinforcement learning
Parmas, Paavo · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Original
Tucker, George, Bhupatiraju, Surya, Gu, Shixiang, Turner, Richard E, Ghahramani, Zoubin, and Levine, Sergey · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Original
Wu, Cathy, Rajeswaran, Aravind, Duan, Yan, Kumar, Vikash, Bayen, Alexandre M., Kakade, Sham, Mordatch, Igor, and Abbeel, Pieter · 2018
Later among the works it cites.
Backprop-q: Generalized backpropagation for stochastic computation graphs
Original
Xu, Xiaoran, Zu, Songpeng, and Zhou, Hanning · 2018
Later among the works it cites.
Probabilistic planning with sequential monte carlo methods
Piché, Alexandre, Thomas, Valentin, Ibrahim, Cyril, Bengio, Yoshua, and Pal, Chris · 2019
Closest in time.