Fetching the paper…
Reading the bibliography…
We analyze a family of large language models in such a lightweight manner that can be done on a single GPU.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón F. Astudillo. 2016 · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie. 2016 · 2016
Earlier work this paper cites.
A regularized framework for sparse and structured neural attention
Vlad Niculae and Mathieu Blondel. 2017 · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
Understanding convolutional neural networks for text classification
Alon Jacovi, Oren Sar Shalom, and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Earlier work this paper cites.
Understanding multi-head attention in abstractive summarization
Joris Baan, Maartje ter Hoeve, Marlies van der Wees, Anne Schuth, and Maarten de Rijke. 2019 · 2019
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen. 2019 · 2019
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Of non-linearity and commutativity in bert
Sumu Zhao, Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2021 · 2021
Later among the works it cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Later among the works it cites.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Jones, , Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah. 2022 · 2022
Later among the works it cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022 · 2022
Later among the works it cites.
Transformer language models without positional encodings still learn positional information
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019a · 2019
Cited alongside, same era.
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2020
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. 2022 · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023 · 2023
Closest in time.
Explaining how transformers use context to build predictions
Javier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, and Marta R. Costa-jussà. 2023 · 2023
Closest in time.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Closest in time.
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023 · 2023
Closest in time.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.