Fetching the paper…
Reading the bibliography…
We study goal misgeneralization, a type of out-of-distribution generalization failure in reinforcement learning (RL).
On the shortest spanning subtree of a graph and the traveling salesman problem
Kruskal, J. B · 1956
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R., Barto, R., Barto, A., Barto, C., Bach, F., and Press, M · 1998
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K · 2005
Earlier work this paper cites.
The basic ai drives
Omohundro, S · 2008
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
Krueger, D., Maharaj, T., and Leike, J · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Quiñonero-Candela, J., Sugiyama, M., Lawrence, N. D., and Schwaighofer, A · 2009
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Singh, S., Lewis, R. L., Barto, A. G., and Sorg, J · 2010
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Corrigibility
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E · 2015
Earlier work this paper cites.
Towards resolving unidentifiability in inverse reinforcement learning
Amin, K. and Singh, S. P · 2016
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P. F., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
What does the universal prior actually look like?, Nov 2016
Christiano, P. F · 2016
Earlier work this paper cites.
Faulty reward functions in the wild
Clark, J. and Amodei, D · 2016
Earlier work this paper cites.
Optimization daemons, Mar 2016
Yudkowsky, E · 2016
Cited alongside, same era.
Good and safe uses of AI oracles
Armstrong, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
On motivations for miri’s highly reliable agent design research, Jan 2017
Taylor, J · 2017
Cited alongside, same era.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
2d robustness
Mikulik, V · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Pérez, G., Camargo, C. Q., and Louis, A. A · 2019
Later among the works it cites.
Invariant risk minimization, 2020
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari, 2018
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Agents and devices: A relative definition of agency
Orseau, L., McGill, S. M., and Legg, S · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
Mesa-optimization, Feb 2018
Rice, I. and many authors · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., Hormozdiari, F., Houlsby, N., Hou, S., Jerfel, G., Karthikesalingam, A., Lucic, M., Ma, Y., McLean, C., Mincu, D., Mitani, A., Montanari, A., Nado, Z., Natarajan, V., Nielson, C., Osborne, T. F., Raman, R., Ramasamy, K., Sayres, R., Schrouff, J., Seneviratne, M., Sequeira, S., Suresh, H., Veitch, V., Vladymyrov, M., Wang, X., Webster, K., Yadlowsky, S., Yun, T., Zhai, X., and Sculley, D · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Later among the works it cites.
Understanding rl vision
Hilton, J., Cammarata, N., Carter, S., Goh, G., and Olah, C · 2020
Later among the works it cites.
Training procgen environment with pytorch
Lee, H · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Out of Distribution Generalization in Machine Learning
Arjovsky, M · 2021
Closest in time.
Unsolved problems in ml safety, 2021
Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J · 2021
Closest in time.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2021
Closest in time.
Out-of-distribution generalization via risk extrapolation (rex), 2021
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A · 2021
Closest in time.
Generalization ¿ utility
Shah, R · 2021
Closest in time.
Optimal policies tend to seek power
Turner, A., Smith, L., Shah, R., Critch, A., and Tadepalli, P · 2021
Closest in time.
Zhuang, S. and Hadfield-Menell, D · 2021
Closest in time.
The effects of reward misspecification: Mapping and mitigating misaligned models, 2022
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Closest in time.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Shah, R., Varma, V., Kumar, R., Phuong, M., Krakovna, V., Uesato, J., and Kenton, Z · 2022
Closest in time.