Efficient Probabilistic Performance Bounds for Inverse Reinforcement Learning
Brown, D. S. and Niekum, S · 2018
Later among the works it cites.
Risk-aware active inverse reinforcement learning
Brown, D. S., Cui, Y., and Niekum, S · 2018
Later among the works it cites.
Active reward learning from critiques
Cui, Y. and Niekum, S · 2018
Later among the works it cites.
Impossibility and uncertainty theorems in ai value alignment (or why your agi should not have a utility function)
Original
Eckersley, P · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Later among the works it cites.
Variational bayesian dropout: Pitfalls and fixes
Hron, J., Matthews, D. G., Ghahramani, Z., et al · 2018
Later among the works it cites.
Learning safe policies with expert guidance
Huang, J., Wu, F., Precup, D., and Cai, Y · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Original
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Later among the works it cites.
Projected bnns: Avoiding weight-space pathologies by learning latent representations of neural network weights
Original
Pradier, M. F., Pan, W., Yao, J., Ghosh, S., and Doshi-Velez, F · 2018
Later among the works it cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Later among the works it cites.
Asking easy questions: A user-friendly approach to active reward learning
Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D. S., Goo, W., Prabhat, N., and Niekum, S · 2019
Later among the works it cites.
Causal confusion in imitation learning
de Haan, P., Jayaraman, D., and Levine, S · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Original
Ghasemipour, S. K. S., Zemel, R., and Gu, S · 2019
Later among the works it cites.
Importance sampling policy evaluation with an estimated behavior policy
Hanna, J., Niekum, S., and Stone, P · 2019
Later among the works it cites.
Learning from a learner
Jacq, A., Geist, M., Paiva, A., and Pietquin, O · 2019
Later among the works it cites.
Risk-sensitive generative adversarial imitation learning
Lacotte, J., Ghavamzadeh, M., Chow, Y., and Pavone, M · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Later among the works it cites.
Ensembledagger: A bayesian approach to safe imitation learning
Menda, K., Driggs-Campbell, K., and Kochenderfer, M. J · 2019
Later among the works it cites.
Learning reward functions by integrating human demonstrations and preferences
Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D · 2019
Later among the works it cites.
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Original
Petrik, M. and Russell, R. H · 2019
Later among the works it cites.
Functional variational bayesian neural networks
Original
Sun, S., Zhang, G., Shi, J., and Grosse, R · 2019
Later among the works it cites.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks
Original
Thananjeyan, B., Balakrishna, A., Rosolia, U., Li, F., McAllister, R., Gonzalez, J. E., Levine, S., Borrelli, F., and Goldberg, K · 2019
Later among the works it cites.
Pragmatic-pedagogic value alignment
Fisac, J. F., Gates, M. A., Hamrick, J. B., Liu, C., Hadfield-Menell, D., Palaniappan, M., Malik, D., Sastry, S. S., Griffiths, T. L., and Dragan, A. D · 2020
Closest in time.