Fetching the paper…
Reading the bibliography…
Safe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints.
E. Altman, “Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program,” Mathematical methods of operations research
1998
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International conference on machine learning
2017
Earlier work this paper cites.
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne, “Imitation learning: A survey of learning methods,” ACM Computing Surveys (CSUR)
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems
2017
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017
2017
Earlier work this paper cites.
F. Codevilla, M. Müller, A. López, V. Koltun, and A. Dosovitskiy, “End-to-end driving via conditional imitation learning,” in 2018 IEEE international conference on robotics and automation (ICRA)
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al
2018
Earlier work this paper cites.
M.-F. Chang, J. Lambert, P. Sangkloy, J. Singh, S. Bak, A. Hartnett, D. Wang, P. Carr, S. Lucey, D. Ramanan, et al
2019
Earlier work this paper cites.
W. M. Czarnecki, R. Pascanu, S. Osindero, S. Jayakumar, G. Swirszcz, and M. Jaderberg, “Distilling policy distillation,” in The 22nd international conference on artificial intelligence and statistics
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Ray, J. Achiam, and D. Amodei, “Benchmarking Safe Exploration in Deep Reinforcement Learning,” 2019
2019
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al
2020
Cited alongside, same era.
W. Ding, M. Xu, and D. Zhao, “Cmts: A conditional multiple trajectory synthesizer for generating safety-critical driving scenarios,” in 2020 IEEE International Conference on Robotics and Automation (ICRA)
2020
Cited alongside, same era.
J. Li, L. Sun, W. Zhan, and M. Tomizuka, “Interaction-aware behavior planning for autonomous vehicles validated with real traffic data,” in Dynamic Systems and Control Conference
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Li, C. Tang, M. Tomizuka, and W. Zhan, “Hierarchical planning through goal-conditioned offline reinforcement learning,” IEEE Robotics and Automation Letters
2022
Later among the works it cites.
J. Li, C. Tang, M. Tomizuka, and W. Zhan, “Dealing with the unknown: Pessimistic offline reinforcement learning,” in Conference on Robot Learning
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Wang, Y. Zhang, X. Zhang, Z. Wu, X. Zhu, S. Jin, T. Tang, and M. Tomizuka, “Offline-online learning of deformation model for cable manipulation with graph neural networks,” IEEE Robotics and Automation Letters
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, and D. Anguelov, “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
Cited alongside, same era.
J. Li, L. Sun, J. Chen, M. Tomizuka, and W. Zhan, “A safe hierarchical planning framework for complex driving scenarios based on reinforcement learning,” in 2021 IEEE International Conference on Robotics and Automation (ICRA)
2021
Cited alongside, same era.
J. Li, H. Ma, Z. Zhang, J. Li, and M. Tomizuka, “Spatio-temporal graph dual-attention network for multi-agent prediction and tracking,” IEEE Transactions on Intelligent Transportation Systems
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” Advances in neural information processing systems
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning
2021
Cited alongside, same era.
Y. Sun, W. L. Ubellacker, W.-L. Ma, X. Zhang, C. Wang, N. V. Csomay-Shanklin, M. Tomizuka, K. Sreenath, and A. D. Ames, “Online learning of unknown dynamics for model-based controllers in legged locomotion,” IEEE Robotics and Automation Letters
2021
Cited alongside, same era.
2022
Later among the works it cites.
P. Ladosz, L. Weng, M. Kim, and H. Oh, “Exploration in deep reinforcement learning: A survey,” Information Fusion
2022
Later among the works it cites.
S. Gronauer, “Bullet-safety-gym: A framework for constrained reinforcement learning,” 2022
2022
Later among the works it cites.
Q. Li, Z. Peng, L. Feng, Q. Zhang, Z. Xue, and B. Zhou, “Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence
2022
Later among the works it cites.
S. S. Shperberg, B. Liu, and P. Stone, “Relaxed exploration constrained reinforcement learning,” in Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems
2023
Closest in time.
I. Uchendu, T. Xiao, Y. Lu, B. Zhu, M. Yan, J. Simon, M. Bennice, C. Fu, C. Ma, J. Jiao, et al
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. Zhang, C. Wang, L. Sun, Z. Wu, X. Zhu, and M. Tomizuka, “Efficient sim-to-real transfer of contact-rich manipulation skills with online admittance residual learning,” in 7th Annual Conference on Robot Learning
2023
Closest in time.