Fetching the paper…
Reading the bibliography…
The standard formulation of Reinforcement Learning lacks a practical way of specifying what are admissible and forbidden behaviors.
Studies in linear and nonlinear programming, 1958
Uzawa, H., Anow, K., and Hurwicz, L · 1958
Earlier work this paper cites.
Iterative methods using lagrange multipliers for solving extremal problems with constraints of the equation type
Polyak, B · 1970
Earlier work this paper cites.
An extragradient method for finding saddle points and for other problems
Korpelevich, G · 1976
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Borkar, V. S · 2005
Earlier work this paper cites.
Where do rewards come from
Singh, S., Lewis, R. L., and Barto, A. G · 2009
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Reinforcement learning and the reward engineering principle
Dewey, D · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Grounding english commands to reward functions
MacGlashan, J., Babes-Vroman, M., desJardins, M., Littman, M. L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L · 2015
Earlier work this paper cites.
Developing a predictive approach to knowledge
White, A · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A · 2017
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Later among the works it cites.
Augmenting automated game testing with deep reinforcement learning
Bergdahl, J., Gordillo, C., Tollmar, K., and Gisslén, L · 2020
Later among the works it cites.
Balancing constraints and rewards with meta-gradient d4pg
Calian, D. A., Mankowitz, D. J., Zahavy, T., Xu, Z., Oh, J., Levine, N., and Mann, T · 2020
Later among the works it cites.
“it’s unwieldy and it takes a lot of time”—challenges and opportunities for creating agents in commercial games
Jacob, M., Devlin, S., and Hofmann, K · 2020
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
Lin, T., Jin, C., and Jordan, M · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J · 2018
Cited alongside, same era.
Mindermann, S., Shah, R., Gleave, A., and Hadfield-Menell, D · 2018
Cited alongside, same era.
Simplifying reward design through divide-and-conquer
Ratner, E., Hadfield-Menell, D., and Dragan, A. D · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Turchetta, M., Kolobov, A., Shah, S., Krause, A., and Agarwal, A · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2020
Later among the works it cites.
First order constrained optimization in policy space
Zhang, Y., Vuong, Q., and Ross, K. W · 2020
Later among the works it cites.
On the expressivity of markov reward
Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M., Precup, D., and Singh, S · 2021
Closest in time.
Graph augmented deep reinforcement learning in the gamerland3d environment
Beeching, E., Peter, M., Marcotte, P., Debangoye, J., Simonin, O., Romoff, J., and Wolf, C · 2021
Closest in time.
Configurable agent with reward as input: A play-style continuum generation
de Woillemont, P. L. P., Labory, R., and Corruble, V · 2021
Closest in time.
Navigation turing test (ntt): Learning to evaluate human-like navigation
Devlin, S., Georgescu, R., Momennejad, I., Rzepecki, J., Zuniga, E., Costello, G., Leroy, G., Shaw, A., and Hofmann, K · 2021
Closest in time.
Adversarial reinforcement learning for procedural content generation
Gisslén, L., Eakins, A., Gordillo, C., Bergdahl, J., and Tollmar, K · 2021
Closest in time.
Improving playtesting coverage via curiosity driven reinforcement learning agents
Gordillo, C., Bergdahl, J., Tollmar, K., and Gisslén, L · 2021
Closest in time.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S · 2021
Closest in time.
Imitation learning: Progress, taxonomies and opportunities
Zheng, B., Verma, S., Zhou, J., Tsang, I., and Chen, F · 2021
Closest in time.
Exploring safer behaviors for deep reinforcement learning
Marchesini, E., Corsi, D., and Farinelli, A · 2022
Closest in time.