Fetching the paper…
Reading the bibliography…
When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A framework for behavioural cloning
Bain, M. and Sammut, C · 1995
Earlier work this paper cites.
Constrained Markov decision processes: stochastic modeling
Altman, E · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
Preference learning with gaussian processes
Chu, W. and Ghahramani, Z · 2005
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Borkar, V. S · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Bhatnagar, S. and Lakshmanan, K · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Van Moffaert, K. and Nowé, A · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., Ondruska, P., and Posner, I · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Later among the works it cites.
Incorporating behavioral constraints in online ai systems
Balakrishnan, A., Bouneffouf, D., Mattei, N., and Rossi, F · 2019
Later among the works it cites.
Human compatible: Artificial intelligence and the problem of control
Russell, S · 2019
Later among the works it cites.
Maximum likelihood constraint inference for inverse reinforcement learning
Scobee, D. R. and Sastry, S. S · 2019
Later among the works it cites.
The alignment problem: Machine learning and human values
Christian, B · 2020
Later among the works it cites.
B-pref: Benchmarking preference-based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
The highd dataset: A drone dataset of naturalistic vehicle trajectories on german highways for validation of highly automated driving systems
Krajewski, R., Bock, J., Kloeker, L., and Eckstein, L · 2018
Cited alongside, same era.
Lee, K., Smith, L., Dragan, A., and Abbeel, P · 2021
Later among the works it cites.
Inverse constrained reinforcement learning
Malik, S., Anwar, U., Aghasi, A., and Ahmed, A · 2021
Later among the works it cites.
Maximum likelihood constraint inference from stochastic demonstrations
McPherson, D. L., Stocking, K. C., and Sastry, S. S · 2021
Later among the works it cites.
Bayesian inverse constrained reinforcement learning
Papadimitriou, D., Anwar, U., and Brown, D. S · 2021
Later among the works it cites.
Discretizing dynamics for maximum likelihood constraint inference, 2021
Stocking, K. C., McPherson, D. L., Matthew, R. P., and Tomlin, C. J · 2021
Later among the works it cites.
Commonroad-rl: A configurable reinforcement learning environment for motion planning of autonomous vehicles
Wang, X., Krasowski, H., and Althoff, M · 2021
Later among the works it cites.
Learning behavioral soft constraints from demonstrations, 2022
Glazier, A., Loreggia, A., Mattei, N., Rahgooy, T., Rossi, F., and Venable, B · 2022
Later among the works it cites.
A primer on maximum causal entropy inverse reinforcement learning
Gleave, A. and Toyer, S · 2022
Later among the works it cites.
Benchmarking constraint inference in inverse reinforcement learning
Liu, G., Luo, Y., Gaurav, A., Rezaee, K., and Poupart, P · 2022
Later among the works it cites.