Fetching the paper…
Reading the bibliography…
Multi-agent control problems constitute an interesting area of application for deep reinforcement learning models with continuous action spaces.
A numerically stable dual method for solving strictly convex quadratic programs
Goldfarb, D. and Idnani, A · 1983
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Altman, E · 1998
Earlier work this paper cites.
Soft constraints and exact penalty functions in model predictive control
Kerrigan, E. C. and Maciejowski, J. M · 2000
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Convex analysis
Rockafellar, R. T · 2015
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E., Abdou, M., Perot, E., and Yogamani, S · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., and Topcu, U · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Cited alongside, same era.
Data center cooling using model-predictive control
Lazic, N., Boutilier, C., Lu, T., Wong, E., Roy, B., Ryu, M., and Imwalle, G · 2018
Emergence of grounded compositional language in multi-agent populations
Mordatch, I. and Abbeel, P · 2018
Later among the works it cites.
Linear model predictive safety certification for learning-based control
Wabersich, K. P. and Zeilinger, M. N · 2018
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control, 2019
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Later among the works it cites.
Learning safe unlabeled multi-robot planning with motion constraints
Khan, A., Zhang, C., Li, S., Wu, J., Schlotfeldt, B., Tang, S. Y., Ribeiro, A., Bastani, O., and Kumar, V · 2019
Later among the works it cites.
Mamps: Safe multi-agent reinforcement learning via model predictive shielding
Zhang, W., Bastani, O., and Kumar, V · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Combating reinforcement learning’s sisyphean curse with intrinsic fear, 2018
Lipton, Z. C., Azizzadenesheli, K., Kumar, A., Li, L., Gao, J., and Deng, L · 2018
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I
Cited in the paper.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I
Cited in the paper.
Exploration-exploitation in constrained mdps, 2020
Efroni, Y., Mannor, S., and Pirotta, M · 2020
Later among the works it cites.
Voronoi-based multi-robot autonomous exploration in unknown environments via deep reinforcement learning
Hu, J., Niu, H., Carrasco, J., Lennox, B., and Arvin, F · 2020
Later among the works it cites.