Fetching the paper…
Reading the bibliography…
Recently, emergence has received widespread attention from the research community along with the success of large-scale models.
A logical calculus of the ideas immanent in nervous activity
McCulloch Warren S and Pitts Walter. 1943 · 1943
Earlier work this paper cites.
Backpropagation Applied to Handwritten Zip Code Recognition
Yann LeCun, Bernhard E. Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, and Lawrence D. Jackel. 1989 · 1989
Earlier work this paper cites.
Neural networks: a comprehensive foundation
Simon Haykin. 1994 · 1994
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun Yann, Bengio Yoshua, et al · 1995
Earlier work this paper cites.
Stable Hebbian learning from spike timing-dependent plasticity
Van Rossum Mark CW, Bi Guo Qiang, and Turrigiano Gina G. 2000 · 2000
Earlier work this paper cites.
The self-tuning neuron: synaptic scaling of excitatory synapses
Turrigiano Gina G. 2008 · 2008
Earlier work this paper cites.
What are artificial neural networks?
Anders Krogh. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In CVPR . IEEE Computer Society, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Neuropil distribution in the cerebral cortex differs between humans and chimpanzees
Spocter Muhammad A, Hopkins William D, Barks Sarah K, Bianchi Serena, Hehmeyer Abigail E, Anderson Sarah M, Stimpson Cheryl D, Fobbs Archibald J, Hof Patrick R, and Sherwood Chet C. 2012 · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks. In NIPS . 1106–1114
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
Neuron-based heredity and human evolution
Gash Don M and Deane Andrew S. 2015 · 2015
Earlier work this paper cites.
Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting. In NIPS . 802–810
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions. In CVPR . IEEE Computer Society, 1–9
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle. In ITW . IEEE, 1–5
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In CVPR . IEEE Computer Society, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Automatic differentiation in PyTorch
Paszke Adam, Gross Sam, Chintala Soumith, Chanan Gregory, Yang Edward, DeVito Zachary, Lin Zeming, Desmaison Alban, Antiga Luca, and Lerer Adam. 2017 · 2017
Earlier work this paper cites.
Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model. In NIPS . 5617–5627
Xingjian Shi, Zhihan Gao, Leonard Lausen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. 2017 · 2017
Cited alongside, same era.
Transcellular nanoalignment of synaptic function
Biederer Thomas, Kaeser Pascal S, and Blanpied Thomas A. 2017 · 2017
Cited alongside, same era.
Attention is All you Need. In NIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
PredRNN: Recurrent Neural Networks for Predictive Learning using Spatiotemporal LSTMs. In NIPS . 879–888
Yunbo Wang, Mingsheng Long, Jianmin Wang, Zhifeng Gao, and Philip S. Yu. 2017 · 2017
Cited alongside, same era.
Complexity in estimating past and future extreme short-duration rainfall
Zhang Xuebin, Zwiers Francis W, Li Guilong, Wan Hui, and Cannon Alex J. 2017 · 2017
Cited alongside, same era.
Transformer Feed-Forward Layers Are Key-Value Memories. In EMNLP (1) . Association for Computational Linguistics, 5484–5495
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Later among the works it cites.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In ICCV . IEEE, 9992–10002
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021 · 2021
Later among the works it cites.
Softmax Linear Units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer El Showk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah. 2022 · 2022
Later among the works it cites.
Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
New theory cracks open the black box of deep learning
Wolchover Natalie. 2018 · 2018
Cited alongside, same era.
On the Information Bottleneck Theory of Deep Learning. In ICLR (Poster) . OpenReview.net
Andrew M. Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D. Tracey, and David D. Cox. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1) . Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In ICLR (Poster) . OpenReview.net
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Àgata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020 · 2020
Cited alongside, same era.
A solution to the learning dilemma for recurrent networks of spiking neurons
Guillaume Bellec, Franz Scherr, Anand Subramoney, Elias Hajek, Darjan Salaj, Robert Legenstein, and Wolfgang Maass. 2020 · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah Chris, Cammarata Nick, Schubert Ludwig, Goh Gabriel, Petrov Michael, and Carter Shan. 2020 · 2020
Cited alongside, same era.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Meta AI. 2023 · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning . PMLR, 2397–2430
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Analyzing Transformers in Embedding Space. In ACL (1) . Association for Computational Linguistics, 16124–16170
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023 · 2023
Closest in time.
Finding Neurons in a Haystack: Case Studies with Sparse Probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Brain-inspired learning in artificial neural networks: a review
Schmidgall Samuel, Achterberg Jascha, Miconi Thomas, Kirsch Louis, Ziaei Rojin, Hajiseyedrazi S, and Eshraghian Jason. 2023 · 2023
Closest in time.
Are Emergent Abilities of Large Language Models a Mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2023 · 2023
Closest in time.
Language models can explain neurons in language models
Bills Steven, Cammarata Nick, Mossing Dan, Tillman Henk, Gao Leo, Goh Gabriel, Sutskever Ilya, Leike Jan, Wu Jeff, and Saunders William. 2023a · 2023
Closest in time.
Language models can explain neurons in language models
Bills Steven, Cammarata Nick, Mossing Dan, Tillman Henk, Gao Leo, Goh Gabriel, Sutskever Ilya, Leike Jan, Wu Jeff, and Saunders William. 2023b · 2023
Closest in time.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken Trenton, Templeton Adly, Batson Joshua, Chen Brian, Jermyn Adam, Conerly Tom, Turner Nick, Anil Cem, Denison Carson, Askell Amanda, et al · 2023
Closest in time.
Yuxiang Zhou, Jiazheng Li, Yanzheng Xiang, Hanqi Yan, Lin Gui, and Yulan He. 2023 · 2023
Closest in time.