Fetching the paper…
Reading the bibliography…
In large deep neural networks that seem to perform surprisingly well on many tasks, we also observe a few failures related to accuracy, social biases, and alignment with human values, among others.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Safety verification of deep neural networks
Huang, X., Kwiatkowska, M., Wang, S., and Wu, M · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Uber’s self-driving car didn’t malfunction, it was just bad, 2018
Madrigal, A. C · 2018
Earlier work this paper cites.
Deep reinforcement learning for de novo drug design
Popova, M., Isayev, O., and Tropsha, A · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Chip placement with deep reinforcement learning
Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J., Songhori, E., Wang, S., Lee, Y.-J., Johnson, E., Pathak, O., Bae, S., et al · 2020
Earlier work this paper cites.
Opportunities and challenges in deep learning adversarial robustness: A survey
Silva, S. H. and Najafirad, P · 2020
Cited alongside, same era.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Cited alongside, same era.
Patchattack: A black-box texture-based attack with reinforcement learning
Yang, C., Kortylewski, A., Xie, C., Cao, Y., and Yuille, A · 2020
Cited alongside, same era.
Deep learning: a statistical viewpoint
Bartlett, P. L., Montanari, A., and Rakhlin, A · 2021
Cited alongside, same era.
Disentangling epistemic and aleatoric uncertainty in reinforcement learning
Charpentier, B., Senanayake, R., Kochenderfer, M., and Günnemann, S · 2022
Later among the works it cites.
How do we fail? stress testing perception in autonomous vehicles
Delecki, H., Itkina, M., Lange, B., Senanayake, R., and Kochenderfer, M · 2022
Later among the works it cites.
Domino: Discovering systematic errors with cross-modal embeddings
Eyuboglu, S., Varma, M., Saab, K., Delbrouck, J.-B., Lee-Messer, C., Dunnmon, J., Zou, J., and Ré, C · 2022
Later among the works it cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, P., Itkina, M., Senanayake, R., and Kochenderfer, M. J · 2021
Cited alongside, same era.
A survey of algorithms for black-box safety validation of cyber-physical systems
Corso, A., Moss, R. J., Koren, M., Lee, R., and Kochenderfer, M. J · 2021
Cited alongside, same era.
Exploring the limits of out-of-distribution detection
Fort, S., Ren, J., and Lakshminarayanan, B · 2021
Cited alongside, same era.
How to train your robot with deep reinforcement learning: lessons we have learned
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., and Levine, S · 2021
Cited alongside, same era.
Human-in-the-Loop Machine Learning: Active learning and annotation for human-centered AI
Monarch, R. M · 2021
Cited alongside, same era.
Out-of-distribution detection for automotive perception
Nitsch, J., Itkina, M., Senanayake, R., Nieto, J., Schmidt, M., Siegwart, R., Kochenderfer, M. J., and Cadena, C · 2021
Cited alongside, same era.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q · 2021
Cited alongside, same era.
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Later among the works it cites.
Distilling model failures as directions in latent space
Jain, S., Lawrence, H., Moitra, A., and Madry, A · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2023
Later among the works it cites.
Curiosity-driven red-teaming for large language models, 2024
Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J., Srivastava, A., and Agrawal, P · 2024
Closest in time.
Lance: Stress-testing visual models by generating language-guided counterfactual images
Prabhu, V., Yenamandra, S., Chattopadhyay, P., and Hoffman, J · 2024
Closest in time.
The role of predictive uncertainty and diversity in embodied ai and robot learning
Senanayake, R · 2024
Closest in time.