Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential.
To filter prune, or to layer prune, that is the question, 2020
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh, Hong Zhang, and Nilanjan Ray · 2007
Earlier work this paper cites.
Robust recovery of signals from a structured union of subspaces
Yonina C. Eldar and Moshe Mishali · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention, 2021
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2012
Earlier work this paper cites.
Highway and residual networks learn unrolled iterative estimation
Klaus Greff, Rupesh K Srivastava, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Recovery guarantee of weighted low-rank approximation via alternating minimization, 2016
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks, 2016
Andreas Veit, Michael Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
A proposal on machine learning via dynamical systems
Weinan Ee · 2017
Earlier work this paper cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Earlier work this paper cites.
Convolutional neural networks analyzed via convolutional sparse coding
Vardan Papyan, Yaniv Romano, and Michael Elad · 2017
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice, 2017
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Residual connections encourage iterative inference
Stanisław Jastrz Ebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2018
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study, 2018
Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Residual networks as nonlinear systems: Stability analysis using linearization, 2019
Kai Rothauge, Zhewei Yao, Zixi Hu, and Michael W. Mahoney · 2019
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale, 2019
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Traces of class/cross-class structure pervade deep learning spectra, 2020
Vardan Papyan · 2020
Earlier work this paper cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, X. Y. Han, and David L. Donoho · 2020
Earlier work this paper cites.
Why should we add early exits to neural networks?
Simone Scardapane, Michele Scarpiniti, Enzo Baccarelli, and Aurelio Uncini · 2020
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Scaling properties of deep residual networks, 2021
Alain-Sam Cohen, Rama Cont, Alain Rossier, and Renyuan Xu · 2021
Cited alongside, same era.
Layer folding: Neural network depth reduction using activation linearization, 2021
Amir Ben Dror, Niv Zehngut, Avraham Raviv, Evgeny Artyomov, Ran Vitek, and Roy Jevnisek · 2021
Cited alongside, same era.
A mathematical principle of deep learning: Learn the geodesic curve in the wasserstein space, 2021
Kuo Gai and Shihua Zhang · 2021
Cited alongside, same era.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Probing neural networks with t-sne, class-specific projections and a guided tour, 2021
Christopher R. Hoyt and Art B. Owen · 2021
Cited alongside, same era.
Discrete, compositional, and symbolic representations through attractor dynamics, 2023
Andrew Nam, Eric Elmoznino, Nikolay Malkin, Chen Sun, Yoshua Bengio, and Guillaume Lajoie · 2023
Later among the works it cites.
Neural collapse in the intermediate hidden layers of classification neural networks, 2023
Liam Parker, Emre Onal, Anton Stengel, and Jake Intrater · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Deep neural collapse is provably optimal for the deep unconstrained features model, 2023
Peter Súkeník, Marco Mondelli, and Christoph Lampert · 2023
Later among the works it cites.
Max-margin token selection in attention mechanism, 2023
Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, and Samet Oymak · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Ode transformer: An ordinary differential equation-inspired model for neural machine translation, 2021
Bei Li, Quan Du, Tao Zhou, Shuhan Zhou, Xin Zeng, Tong Xiao, and Jingbo Zhu · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers, 2021
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Separation and concentration in deep networks, 2021
John Zarka, Florentin Guth, and Stéphane Mallat · 2021
Cited alongside, same era.
Nearest class-center simplification through intermediate layers
Ido Ben-Shaul and Shai Dekel · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model, 2022
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach · 2022
Cited alongside, same era.
Introducing mpt-30b: Raising the bar for open-source foundation models, 2023a
MosaicML NLP Team · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023b
MosaicML NLP Team · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
The geometry of hidden representations of large transformer models, 2023
Lucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio, Alessio Ansuini, and Alberto Cazzaniga · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent, 2023
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
A neural ode interpretation of transformer layers, 2022
Yaofeng Desmond Zhong, Tongtao Zhang, Amit Chakraborty, and Biswadip Dey · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Slicegpt: Compress large language models by deleting rows and columns, 2024
Saleh Ashkboos, Maximilian L. Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman · 2024
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2024
Closest in time.
Unifying low dimensional observations in deep learning through the deep linear unconstrained feature model, 2024
Connall Garrod and Jonathan P. Keating · 2024
Closest in time.
Olmo: Accelerating the science of language models, 2024
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Closest in time.
The unreasonable ineffectiveness of the deeper layers, 2024
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts · 2024
Closest in time.
Early-exit neural networks with nested prediction sets, 2024
Metod Jazbec, Patrick Forré, Stephan Mandt, Dan Zhang, and Eric Nalisnick · 2024
Closest in time.
Layermerge: Neural network depth compression through layer pruning and merging, 2024
Jinuk Kim, Marwa El Halabi, Mingi Ji, and Hyun Oh Song · 2024
Closest in time.
The remarkable robustness of llms: Stages of inference?, 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Falcon2-11b technical report, 2024
Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru, Mugariya Farooq, Giulia Campesan, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Maksim Velikanov, Basma El Amel Boussaha, Mohammed Al-Yafeai, Hamza Alobeidli, Leen Al Qadi, Mohamed El Amine Seddik, Kirill Fedyanin, Reda Alami, and Hakim Hacid · 2024
Closest in time.
Code llama: Open foundation models for code, 2024
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2024
Closest in time.
Lines of thought in large language models, 2024
Raphaël Sarfati, Toni J. B. Liu, Nicolas Boullé, and Christopher J. Earls · 2024
Closest in time.
Gemma: Open models based on gemini research and technology, 2024
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy · 2024
Closest in time.
Tycho F. A. van der Ouderaa, Markus Nagel, Mart van Baalen, Yuki M. Asano, and Tijmen Blankevoort · 2024
Closest in time.
Linguistic collapse: Neural collapse in (large) language models, 2024
Robert Wu and Vardan Papyan · 2024
Closest in time.
Benefits of transformer: In-context learning in linear regression tasks with unstructured data
Yue Xing, Xiaofeng Lin, Namjoon Suh, Qifan Song, and Guang Cheng · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan · 2024
Closest in time.
Neural rank collapse: Weight decay and small within-class variability yield low-rank bias, 2024
Emanuele Zangrando, Piero Deidda, Simone Brugiapaglia, Nicola Guglielmi, and Francesco Tudisco · 2024
Closest in time.