Fetching the paper…
Reading the bibliography…
We introduce Wasserstein Adversarial Proximal Policy Optimization (WAPPO), a novel algorithm for visual transfer in Reinforcement Learning that explicitly learns to align the distributions of extracted features between a source and target task.
State of the art—a survey of partially observable markov decision processes: Theory, models, and algorithms
George E Monahan · 1982
Earlier work this paper cites.
Principal component analysis
Svante Wold, Kim Esbensen, and Paul Geladi · 1987
Earlier work this paper cites.
The earth mover’s distance as a metric for image retrieval
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas · 2000
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
Patrice Y Simard, David Steinkraus, John C Platt, et al · 2003
Earlier work this paper cites.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
Andrew Y Ng · 2004
Earlier work this paper cites.
Optimal Transport: Old and New
Cédric Villani · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
Transfer in reinforcement learning: a framework and a survey
Alessandro Lazaric · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Safety and control in collaborative robotics
Tanya M Anandan · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Simultaneous deep transfer across domains and tasks
Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
A neural algorithm of artistic style
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2015
Earlier work this paper cites.
Multivariate Density Estimation: Theory, Practice, and Visualization
David W Scott · 2015
Earlier work this paper cites.
Cad2rl: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Finale Doshi-Velez and George Konidaris · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio · 2018
Later among the works it cites.
Illuminating generalization in deep reinforcement learning through procedural level generation
Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager, Ahmed Khalifa, Julian Togelius, and Sebastian Risi · 2018
Later among the works it cites.
Investigating human priors for playing video games
Rachit Dubey, Pulkit Agrawal, Deepak Pathak, Tom Griffiths, and Alexei Efros · 2018
Later among the works it cites.
Reptile: a scalable metalearning algorithm
Alex Nichol and John Schulman · 2018
Later among the works it cites.
Direct policy transfer via hidden parameter markov decision processes
Jiayu Yao, Taylor Killian, George Konidaris, and Finale Doshi-Velez · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Cited alongside, same era.
Adversarial discriminative domain adaptation
Eric Tzeng, Judy Hoffman, Kate Saenko, and Trevor Darrell · 2017
Cited alongside, same era.
Learning invariant feature spaces to transfer skills with reinforcement learning
Abhishek Gupta, Coline Devin, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Robust and efficient transfer learning with hidden parameter markov decision processes
Taylor W Killian, Samuel Daulton, George Konidaris, and Finale Doshi-Velez · 2017
Cited alongside, same era.
Later among the works it cites.
Cycada: Cycle-consistent adversarial domain adaptation
Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2019
Later among the works it cites.
Obstacle tower: A generalization challenge in vision, control, and planning
Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper, Ervin Teng, Hunter Henry, Adam Crespi, Julian Togelius, and Danny Lange · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Later among the works it cites.
Robust domain randomization for reinforcement learning
Reda Bahi Slaoui, William R Clements, Jakob N Foerster, and Sébastien Toth · 2019
Later among the works it cites.
Vr-goggles for robots: Real-to-sim domain adaptation for visual control
Jingwei Zhang, Lei Tai, Peng Yun, Yufeng Xiong, Ming Liu, Joschka Boedecker, and Wolfram Burgard · 2019
Later among the works it cites.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Later among the works it cites.
Neural style transfer: A review
Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2020
Closest in time.
Adapting deep visuomotor representations with weak pairwise constraints
Eric Tzeng, Coline Devin, Judy Hoffman, Chelsea Finn, Pieter Abbeel, Sergey Levine, Kate Saenko, and Trevor Darrell · 2020
Closest in time.
Invariant causal prediction for block mdps
Amy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos, Marta Kwiatkowska, Joelle Pineau, Yarin Gal, and Doina Precup · 2020
Closest in time.