Fetching the paper…
Reading the bibliography…
This paper tackles the problem of learning value functions from undirected state-only experience (state transitions without action labels i.e.
Imitation learning from video by leveraging proprioception
Faraz Torabi, Garrett Warnell, and Peter Stone · 1905
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in markov decision processes
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Bounding performance loss in approximate mdp homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2008
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin V Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
Tianfan Xue, Jiajun Wu, Katherine L Bouman, and William T Freeman · 2016
Cited alongside, same era.
Unsupervised perceptual rewards for imitation learning
Pierre Sermanet, Kelvin Xu, and Sergey Levine · 2017
Cited alongside, same era.
On evaluation of embodied navigation agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, and Amir Zamir · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Few-shot goal inference for visuomotor learning and planning
Annie Xie, Avi Singh, Sergey Levine, and Chelsea Finn · 2018
Cited alongside, same era.
Estimating q ( s , s ′ ) q(s,s^{\prime}) with deep deterministic dynamics gradients
Ashley D. Edwards, Himanshu Sahni, Rosanne Liu, Jane Hung, Ankit Jain, Rui Wang, Adrien Ecoffet, Thomas Miconi, Charles Isbell, and Jason Yosinski · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning, 2020
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li, Mohammad Norouzi, Matt Hoffman, Ofir Nachum, George Tucker, Nicolas Heess, and Nando deFreitas · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved conditional vrnns for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Cited alongside, same era.
Perceptual values from observation
Ashley D. Edwards and Charles L. Isbell · 2019
Cited alongside, same era.
Imitating latent policies from observation
Ashley D Edwards, Himanshu Sahni, Yannick Schroecker, and Charles L Isbell · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham Kakade, and Sergey Levine · 2019
Cited alongside, same era.
Habitat: A platform for embodied AI research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
End-to-end robotic reinforcement learning without reward engineering
Avi Singh, Larry Yang, Chelsea Finn, and Sergey Levine · 2019
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
IRIS: implicit reinforcement without interaction at scale for learning control from offline robot manipulation data
Ajay Mandlekar, Fabio Ramos, Byron Boots, Silvio Savarese, Fei-Fei Li, Animesh Garg, and Dieter Fox · 2020
Later among the works it cites.
State-only imitation learning for dexterous manipulation
Ilija Radosavovic, Xiaolong Wang, Lerrel Pinto, and Jitendra Malik · 2020
Later among the works it cites.
Offline reinforcement learning from images with latent space models
Rafael Rafailov, Tianhe Yu, Aravind Rajeswaran, and Chelsea Finn · 2020
Later among the works it cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2020
Later among the works it cites.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
Shuran Song, Andy Zeng, Johnny Lee, and Thomas Funkhouser · 2020
Later among the works it cites.
Plannable approximations to MDP homomorphisms: Equivariance under actions
Elise van der Pol, Thomas Kipf, Frans A. Oliehoek, and Max Welling · 2020
Later among the works it cites.
Model-based offline planning, 2021
Arthur Argenson and Gabriel Dulac-Arnold · 2021
Later among the works it cites.
A workflow for offline model-free robotic reinforcement learning
Aviral Kumar, Anikait Singh, Stephen Tian, Chelsea Finn, and Sergey Levine · 2021
Later among the works it cites.