Fetching the paper…
Reading the bibliography…
Reward design is a key component of deep reinforcement learning, yet some tasks and designer's objectives may be unnatural to define as a scalar cost function.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Monitoring temporal properties of continuous signals
Oded Maler and Dejan Nickovic · 2004
Earlier work this paper cites.
Robustness of temporal logic specifications for continuous-time signals
Georgios E Fainekos and George J Pappas · 2009
Earlier work this paper cites.
Modeling difference rewards for multiagent learning
Scott Proper and Kagan Tumer · 2012
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Autonomous driving using model predictive control and a kinematic bicycle vehicle model
J Kong, M Pfeiffer, G Schildbach, and F Borrelli · 2015
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Control barrier function based quadratic programs for safety critical systems
Aaron D Ames, Xiangru Xu, Jessy W Grizzle, and Paulo Tabuada · 2016
Earlier work this paper cites.
Reinforcement learning with temporal logic rewards
Xiao Li, Cristian-Ioan Vasile, and Calin Belta · 2017
Earlier work this paper cites.
Robust online monitoring of signal temporal logic
Jyotirmoy V Deshmukh, Alexandre Donzé, Shromona Ghosh, Xiaoqing Jin, Garvit Juniwal, and Sanjit A Seshia · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Scenario model predictive control for lane change assistance and autonomous driving on highways
Gianluca Cesari, Georg Schildbach, Ashwin Carvalho, and Francesco Borrelli · 2017
Cited alongside, same era.
Carla: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Cited alongside, same era.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith · 2018
Cited alongside, same era.
Control barrier functions for signal temporal logic tasks
Lars Lindemann and Dimos V Dimarogonas · 2018
Cited alongside, same era.
Trajectory planning and tracking for autonomous overtaking: State-of-the-art and future prospects
Shilp Dixit, Saber Fallah, Umberto Montanaro, Mehrdad Dianati, Alan Stevens, Francis Mccullough, and Alexandros Mouzakitis · 2018
Cited alongside, same era.
Multi-agent reinforcement learning with temporal logic specifications
Lewis Hammond, Alessandro Abate, Julian Gutierrez, and Michael Wooldridge · 2021
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2021
Later among the works it cites.
Safe multi-agent reinforcement learning via shielding
Ingy ElSayed-Aly, Suda Bharadwaj, Christopher Amato, Rüdiger Ehlers, Ufuk Topcu, and Lu Feng · 2021
Later among the works it cites.
Safety-critical model predictive control with discrete-time control barrier function
Jun Zeng, Bike Zhang, and Koushil Sreenath · 2021
Later among the works it cites.
Deep multi-agent reinforcement learning for highway on-ramp merging in mixed traffic
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structured reward shaping using signal temporal logic specifications
Anand Balakrishnan and Jyotirmoy V Deshmukh · 2019
Cited alongside, same era.
Multi-agent reinforcement learning-based resource allocation for uav networks
Jingjing Cui, Yuanwei Liu, and Arumugam Nallanathan · 2019
Cited alongside, same era.
Joint optimization of multi-uav target assignment and path planning based on multi-agent reinforcement learning
Han Qie, Dianxi Shi, Tianlong Shen, Xinhai Xu, Yuan Li, and Liujing Wang · 2019
Cited alongside, same era.
Multi-agent deep reinforcement learning for large-scale traffic signal control
Tianshu Chu, Jie Wang, Lara Codecà, and Zhaojian Li · 2019
Cited alongside, same era.
A formal methods approach to interpretable reinforcement learning for robotic planning
Xiao Li, Zachary Serlin, Guang Yang, and Calin Belta · 2019
Cited alongside, same era.
Extended markov games to learn multiple tasks in multi-agent reinforcement learning
Borja G León and Francesco Belardinelli · 2020
Cited alongside, same era.
Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V Albrecht · 2020
Cited alongside, same era.
Dong Chen, Mohammad Hajidavalloo, Zhaojian Li, Kaian Chen, Yongqiang Wang, Longsheng Jiang, and Yue Wang · 2021
Later among the works it cites.
Zhili Zhang, Songyang Han, Jiangwei Wang, and Fei Miao · 2022
Later among the works it cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig · 2022
Later among the works it cites.
Reinforcement learning with stochastic reward machines
Jan Corazza, Ivan Gavran, and Daniel Neider · 2022
Later among the works it cites.
Accelerated reinforcement learning for temporal logic control objectives
Yiannis Kantaros · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
A multi-agent deep reinforcement learning coordination framework for connected and automated vehicles at merging roadways
Sai Krishna Sumanth Nakka, Behdad Chalaki, and Andreas A Malikopoulos · 2022
Later among the works it cites.
State-wise safe reinforcement learning: A survey
Weiye Zhao, Tairan He, Rui Chen, Tianhao Wei, and Changliu Liu · 2023
Closest in time.
Safe reinforcement learning under temporal logic with reward design and quantum action selection
Mingyu Cai, Shaoping Xiao, Junchao Li, and Zhen Kan · 2023
Closest in time.