Fetching the paper…
Reading the bibliography…
Pretrained large language models (LLMs) are becoming increasingly powerful and ubiquitous in mainstream applications such as being a personal assistant, a dialogue model, etc.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 1905
Earlier work this paper cites.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank · 1906
Earlier work this paper cites.
What does BERT look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 1906
Earlier work this paper cites.
Visualizing and measuring the geometry of BERT
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda B. Viégas, and Martin Wattenberg · 1906
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 1908
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky · 1908
Earlier work this paper cites.
On the validity of self-attention as explanation in transformer models
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, and Roger Wattenhofer · 1908
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 1909
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding, 2019
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 1910
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with checklist
Marco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2005
Earlier work this paper cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2006
Earlier work this paper cites.
Attention flows: Analyzing and comparing attention mechanisms in language models
Joseph F. DeRose, Jiayao Wang, and Matthew Berger · 2009
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2012
Earlier work this paper cites.
Do recommender systems manipulate consumer preferences? a study of anchoring effects
Gediminas Adomavicius, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang · 2013
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2017
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems, 2017
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James R. Glass · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Senteval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela · 2018
Cited alongside, same era.
Avoiding reasoning shortcuts: Adversarial evaluation, training, and model development for multi-hop QA
Yichen Jiang and Mohit Bansal · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Cited alongside, same era.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Jonas Wallat, Jaspreet Singh, and Avishek Anand · 2020
Cited alongside, same era.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models, 2021
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2021
Later among the works it cites.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
An overview of early vision in inceptionv1
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Curve detectors
Nick Cammarata, Gabriel Goh, Shan Carter, Ludwig Schubert, Michael Petrov, and Chris Olah · 2020
Cited alongside, same era.
Naturally occurring equivariance in neural networks
Chris Olah, Nick Cammarata, Chelsea Voss, Ludwig Schubert, and Gabriel Goh · 2020
Cited alongside, same era.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2020
Cited alongside, same era.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2020
Cited alongside, same era.
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Later among the works it cites.
Matija Franklin, Hal Ashton, Rebecca Gorman, and Stuart Armstrong · 2022
Later among the works it cites.
Godel: Large-scale pre-training for goal-directed dialog
Baolin Peng, Michel Galley, Pengcheng He, Chris Brockett, Lars Liden, Elnaz Nouri, Zhou Yu, Bill Dolan, and Jianfeng Gao · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability, 2022
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, Buck Shlegeris, and Nate Thomas · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Later among the works it cites.
Attention understands semantic relations
Anastasia Chizhikova, Sanzhar Murzakhmetov, Oleg Serikov, Tatiana Shavrina, and Mikhail Burtsev · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Locating and editing factual associations in gpt, 2022
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Later among the works it cites.
A mechanistic interpretability analysis of grokking
Neel Nanda and Tom Lieberum · 2022
Later among the works it cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit, 2022
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Later among the works it cites.
Emergent abilities of large language models, 2022
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2022
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2022
Later among the works it cites.