Fetching the paper…
Reading the bibliography…
Offline reinforcement learning can enable policy learning from pre-collected, sub-optimal datasets without online interactions.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,” Biometrika
1952
Earlier work this paper cites.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,” IEEE Transactions on Systems, Man, and Cybernetics
1983
Earlier work this paper cites.
D. A. Pomerleau, “Efficient Training of Artificial Neural Networks for Autonomous Navigation,” Neural Computation
1991
Earlier work this paper cites.
S. Russell, “Learning agents for uncertain environments (extended abstract),” in Proceedings of the Eleventh Annual Conference on Computational Learning Theory
1998
Earlier work this paper cites.
A. Y. Ng and S. J. Russell, “Algorithms for inverse reinforcement learning,” in Proceedings of the Seventeenth International Conference on Machine Learning
2000
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al
2008
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” Advances in neural information processing systems
2016
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in neural information processing systems
2017
Earlier work this paper cites.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proceedings of the AAAI conference on artificial intelligence
2018
Earlier work this paper cites.
The MIT Press, second ed., 2018
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction · 2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” 2021
2021
Cited alongside, same era.
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” Advances in neural information processing systems
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. J. Ma, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman, “Liv: Language-image representations and rewards for robotic control,” in International Conference on Machine Learning
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
T. Ni, H. Sikchi, Y. Wang, T. Gupta, L. Lee, and B. Eysenbach, “f-irl: Inverse reinforcement learning via state marginal matching,” in Conference on Robot Learning
2021
Cited alongside, same era.
K. Lee, L. Smith, and P. Abbeel, “Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training,” 2021
2021
Cited alongside, same era.
T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” 2021
2021
Cited alongside, same era.
X. Lin, Y. Wang, J. Olkin, and D. Held, “Softgym: Benchmarking deep reinforcement learning for deformable object manipulation,” 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al
2021
Cited alongside, same era.
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson, “Implicit behavioral cloning,” in Conference on Robot Learning
2022
Cited alongside, same era.
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto, “Behavior transformers: Cloning k k modes with one stone,” Advances in neural information processing systems
2022
Cited alongside, same era.
D. Seita, Y. Wang, S. J. Shetty, E. Y. Li, Z. Erickson, and D. Held, “Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds,” in Conference on Robot Learning
2023
Later among the works it cites.
Y. Wang, Z. Sun, Z. Erickson, and D. Held, “One policy to dress them all: Learning to dress people with diverse poses and garments,” in Robotics: Science and Systems (RSS)
2023
Later among the works it cites.
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Wang, Z. Sun, J. Zhang, Z. Xian, E. Biyik, D. Held, and Z. Erickson, “Rl-vlm-f: Reinforcement learning from vision language foundation model feedback,” 2024
2024
Closest in time.
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu, “Text2reward: Reward shaping with language models for reinforcement learning,” 2024
2024
Closest in time.
Y. Wang, Z. Xian, F. Chen, T.-H. Wang, Y. Wang, K. Fragkiadaki, Z. Erickson, D. Held, and C. Gan, “Robogen: Towards unleashing infinite data for automated robot learning via generative simulation,” in International conference on machine learning
2024
Closest in time.
S. Sontakke, J. Zhang, S. Arnold, K. Pertsch, E. Bıyık, D. Sadigh, C. Finn, and L. Itti, “Roboclip: One demonstration is enough to learn robot policies,” Advances in Neural Information Processing Systems
2024
Closest in time.
2024
Closest in time.
Z. Sun, Y. Wang, D. Held, and Z. Erickson, “Force-constrained visual policy: Safe robot-assisted dressing via multi-modal sensing,” IEEE Robotics and Automation Letters
2024
Closest in time.