Sequence level training with recurrent neural networks
Ranzato, Marc’Aurelio, Chopra, Sumit, Auli, Michael, and Zaremba, Wojciech · 2015
Later among the works it cites.
Visual chunking: A list prediction framework for region-based object detection
Rhinehart, Nicholas, Zhou, Jiaji, Hebert, Martial, and Bagnell, J Andrew · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael I, and Moritz, Philipp · 2015
Later among the works it cites.
Improving multi-step prediction of learned time series models
Venkatraman, Arun, Hebert, Martial, and Bagnell, J Andrew · 2015
Later among the works it cites.
An actor-critic algorithm for sequence prediction
Original
Bahdanau, Dzmitry, Brakel, Philemon, Xu, Kelvin, Goyal, Anirudh, Lowe, Ryan, Pineau, Joelle, Courville, Aaron, and Bengio, Yoshua · 2016
Later among the works it cites.
Openai gym
Original
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, Chelsea, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
Generative adversarial imitation learning
Ho, Jonathan and Ermon, Stefano · 2016
Later among the works it cites.
Plato: Policy learning using adaptive trajectory optimization
Original
Kahn, Gregory, Zhang, Tianhao, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
Deep reinforcement learning for dialogue generation
Original
Li, Jiwei, Monroe, Will, Ritter, Alan, Galley, Michel, Gao, Jianfeng, and Jurafsky, Dan · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, David et al · 2016
Later among the works it cites.