Fetching the paper…
Reading the bibliography…
The field of AI alignment is concerned with AI systems that pursue unintended goals.
The intentional stance
Dennett, D. C · 1987
Earlier work this paper cites.
On the foundations of noise-free selective classification
El-Yaniv, R. et al · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Li, L., Littman, M. L., Walsh, T. J., and Strehl, A. L · 2011
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude. coursera: Neural networks for machine learning
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Earlier work this paper cites.
“Why should I trust you?" Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Automated inference on criminality using face images
Wu, X. and Zhang, X · 2016
Earlier work this paper cites.
Physiognomy’s new clothes, 2017
Aguera y Arcas, B., Mitchell, M., and Todorov, A · 2017
Earlier work this paper cites.
Gender shades: intersectional phenotypic and demographic evaluation of face datasets and gender classifiers
Buolamwini, J. A · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A. D · 2017
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Cited alongside, same era.
Women also snowboard: Overcoming bias in captioning models
Hendricks, L. A., Burns, K., Saenko, K., Darrell, T., and Rohrbach, A · 2018
Cited alongside, same era.
Irving, G., Christiano, P., and Amodei, D · 2018
Cited alongside, same era.
The surprising creativity of digital evolution
Lehman, J., Clune, J., and Misevic, D · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition
Winkler, J. K., Fink, C., Toberer, F., Enk, A., Deinlein, T., Hofmann-Wellenhof, R., Thomas, L., Lallas, A., Blum, A., Stolz, W., et al · 2019
Later among the works it cites.
Writeup: Progress on AI safety via debate, 2020
Barnes, E., Christiano, P., Ouyang, L., and Irving, G · 2020
Later among the works it cites.
Thread: circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automated classification of skin lesions: From pixels to practice
Narla, A., Kuprel, B., Sarin, K., Novoa, R., and Ko, J · 2018
Cited alongside, same era.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Cited alongside, same era.
Agents and devices: A relative definition of agency
Orseau, L., McGill, S. M., and Legg, S · 2018
Cited alongside, same era.
Building safe artificial intelligence: specification, robustness, and assurance, 2018
Ortega, P. A., Maini, V., and Team, D. S · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study
Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., and Oermann, E. K · 2018
Cited alongside, same era.
Solving Rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Cited alongside, same era.
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al · 2020
Later among the works it cites.
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D · 2020
Later among the works it cites.
Specification gaming: the flip side of AI ingenuity
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., and Legg, S · 2020
Later among the works it cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Later among the works it cites.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Rieger, L., Singh, C., Murdoch, W., and Yu, B · 2020
Later among the works it cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Later among the works it cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Later among the works it cites.
Exploring the limits of out-of-distribution detection
Fort, S., Ren, J., and Lakshminarayanan, B · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Optimal policies tend to seek power
Turner, A. M., Smith, L., Shah, R., Critch, A., and Tadepalli, P · 2021
Later among the works it cites.
Learning robust real-time cultural transmission without human data
CGI, T., Bhoopchand, A., Brownfield, B., Collister, A., Lago, A. D., Edwards, A., Everett, R., Frechette, A., Oliveira, Y. G., Hughes, E., et al · 2022
Closest in time.
Goal misgeneralization in deep reinforcement learning
Di Langosco, L. L., Koch, J., Sharkey, L. D., Pfau, J., and Krueger, D · 2022
Closest in time.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Closest in time.