“Sparse coding with an overcomplete basis set: A strategy employed by V1?”
Bruno Olshausen and David Field · 1997
Earlier work this paper cites.
“The Berkeley framenet project”
Collin Baker, Charles Fillmore and John Lowe · 1998
Earlier work this paper cites.
“Efficient sparse coding algorithms”
Honglak Lee, Alexis Battle, Rajat Raina and Andrew Ng · 2006
Earlier work this paper cites.
“Visualizing data using t-SNE.”
Laurens van Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
“Neural machine translation of rare words with subword units”
Original
Rico Sennrich, Barry Haddow and Alexandra Birch · 2015
Earlier work this paper cites.
“Neural discrete representation learning”
Aaron van Oord and Oriol Vinyals · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“The building blocks of interpretability”
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye and Alexander Mordvintsev · 2018
Earlier work this paper cites.
“Are sixteen heads really better than one?”
Paul Michel, Omer Levy and Graham Neubig · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Scaling laws for neural language models”
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Earlier work this paper cites.
“Zoom in: An introduction to circuits”
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov and Shan Carter · 2020
Earlier work this paper cites.
“A simple but tough-to-beat data augmentation approach for natural language understanding and generation”
Original
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu and Weizhu Chen · 2020
Earlier work this paper cites.
“Investigating gender bias in language models using causal mediation analysis”
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer and Stuart Shieber · 2020
Earlier work this paper cites.
“Low-complexity probing via finding subnetworks”
Original
Steven Cao, Victor Sanh and Alexander Rush · 2021
Earlier work this paper cites.
“A mathematical framework for transformer circuits”
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen and Tom Conerly · 2021
Earlier work this paper cites.
“Causal analysis of syntactic agreement mechanisms in neural language models”
Original
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen and Yonatan Belinkov · 2021
Earlier work this paper cites.