Fetching the paper…
Reading the bibliography…
In this paper, we comprehensively study the inductive biases of two major approaches to augmenting Transformers with a recurrent mechanism: (1) the approach of incorporating a depth-wise recurrence similar to Universal Transformers; and (2) the approach of incorporating a chunk-wise temporal recurrence like Temporal Latent Bottleneck.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The ACL Anthology network corpus
Dragomir R. Radev, Pradeep Muthukrishnan, and Vahed Qazvinian · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Self-delimiting neural networks
Jürgen Schmidhuber · 2012
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Tree-structured composition in neural networks without tree-structured architectures
Samuel R. Bowman, Christopher D. Manning, and Christopher Potts · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Lukasz Kaiser and Ilya Sutskever · 2016
Earlier work this paper cites.
Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognition
Shawn Tan and Khe Chai Sim · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre · 2018
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
ListOps: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel Bowman · 2018
Earlier work this paper cites.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz · 2018
Earlier work this paper cites.
Deep equilibrium models
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2019
Earlier work this paper cites.
Recursive routing networks: Learning to compose modules for language understanding
Ignacio Cases, Clemens Rosenbaum, Matthew Riemer, Atticus Geiger, Tim Klinger, Alex Tamkin, Olivia Li, Sandhini Agarwal, Joshua D. Greene, Dan Jurafsky, Christopher Potts, and Lauri Karttunen · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Earlier work this paper cites.
Cooperative learning of disjoint syntax and semantics
Serhii Havrylov, Germán Kruszewski, and Armand Joulin · 2019
Earlier work this paper cites.
Ordered memory
Yikang Shen, Shawn Tan, Arian Hosseini, Zhouhan Lin, Alessandro Sordoni, and Aaron C Courville · 2019
Earlier work this paper cites.
Highway transformer: Self-gating enhanced self-attentive networks
Yekun Chai, Shuo Jin, and Xinwen Hou · 2020
Earlier work this paper cites.
Depth-adaptive transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli · 2020
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Earlier work this paper cites.
Neurocoder: Learning general-purpose computation using stored neural programs
Hung Le and Svetha Venkatesh · 2020
Earlier work this paper cites.
MART: Memory-augmented recurrent transformer for coherent video paragraph captioning
Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara Berg, and Mohit Bansal · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Bert loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei · 2020
Cited alongside, same era.
Pondernet: Learning to ponder
Andrea Banino, Jan Balaguer, and Charles Blundell · 2021
Cited alongside, same era.
Skyformer: Remodel self-attention with gaussian kernel and nyström method
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang · 2021
Cited alongside, same era.
Transformer in transformer
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Cited alongside, same era.
Dynamic neural networks: A survey
Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang · 2021
Cited alongside, same era.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Early exit with disentangled representation and equiangular tight frame
Yixin Ji, Jikai Wang, Juntao Li, Qiang Chen, Wenliang Chen, and Min Zhang · 2023
Later among the works it cites.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Later among the works it cites.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Cited alongside, same era.
Dynamic inference with neural interpreters
Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter Vincent Gehler, Yoshua Bengio, Francesco Locatello, and Bernhard Schölkopf · 2021
Cited alongside, same era.
Modeling hierarchical structures with continuous recursive neural networks
Jishnu Ray Chowdhury and Cornelia Caragea · 2021
Cited alongside, same era.
Consistent accelerated inference via confident adaptive transformers
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay · 2021
Cited alongside, same era.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Cited alongside, same era.
BERxiT: Early exiting for BERT with better fine-tuning and extension to regression
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin · 2021
Cited alongside, same era.
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Later among the works it cites.
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić, and Edoardo Maria Ponti · 2023
Later among the works it cites.
Block-state transformers
Jonathan Pilault, Mahan Fathi, Orhan Firat, Christopher Pal, Pierre-Luc Bacon, and Ross Goroshin · 2023
Later among the works it cites.
Sparse modular activation for efficient sequence modeling
Liliang Ren, Yang Liu, Shuohang Wang, Yichong Xu, Chenguang Zhu, and ChengXiang Zhai · 2023
Later among the works it cites.
Moduleformer: Learning modular large language models from uncurated data
Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, and Chuang Gan · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman · 2023
Later among the works it cites.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2023
Later among the works it cites.
Sparse universal transformer
Shawn Tan, Yikang Shen, Zhenfang Chen, Aaron Courville, and Chuang Gan · 2023
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Tran, Dani Yogatama, and Donald Metzler · 2023
Later among the works it cites.
Adaptive computation with elastic input sequence
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
On the long range abilities of transformers
Itamar Zimerman and Lior Wolf · 2023
Later among the works it cites.
xlstm: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Closest in time.
Beyond attention: Breaking the limits of transformer context length with recurrent memory
Aydar Bulatov, Yuri Kuratov, Yermek Kapushev, and Mikhail Burtsev · 2024
Closest in time.
Moeut: Mixture-of-experts universal transformers
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, and Christopher D Manning · 2024
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Closest in time.
Stack attention: Improving the ability of transformers to model hierarchical patterns
Brian DuSell and David Chiang · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Albert Gu and Tri Dao · 2024
Closest in time.
Transformerfam: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar · 2024
Closest in time.
In search of needles in a 10m haystack: Recurrent memory finds what llms miss
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Closest in time.
A transformer with stack attention
Jiaoda Li, Jennifer C White, Mrinmaya Sachan, and Ryan Cotterell · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al · 2024
Closest in time.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Closest in time.
Hgrn2: Gated linear rnns with state expansion
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong · 2024
Closest in time.