Fetching the paper…
Reading the bibliography…
The Universal Transformer (UT) is a variant of the Transformer that shares parameters across its layers.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Compositional generalization in a deep seq2seq model by separating syntax and semantics
Jake Russin, Jason Jo, Randall C O’Reilly, and Yoshua Bengio. 2019 · 1904
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Compositional generalization for primitive substitutions
Yuanpeng Li, Liang Zhao, Jianyu Wang, and Joel Hestness. 2019 · 1910
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, et al. 2019 · 1912
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Church and Patrick Hanks. 1990 · 1990
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Compositional generalization in semantic parsing: Pre-training vs. specialized architectures
Daniel Furrer, Marc van Zee, Nathan Scales, and Nathanael Schärli. 2020 · 2007
Earlier work this paper cites.
Very deep transformers for neural machine translation
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao. 2020 · 2008
Earlier work this paper cites.
Recursive top-down production for sentence generation with latent trees
Shawn Tan, Yikang Shen, Timothy J O’Donnell, Alessandro Sordoni, and Aaron Courville. 2020 · 2010
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2020 · 2011
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. 2014 · 2014
Earlier work this paper cites.
Tree-structured composition in neural networks without tree-structured architectures
Samuel R Bowman, Christopher D Manning, and Christopher Potts. 2015 · 2015
Earlier work this paper cites.
Łukasz Kaiser and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognition
Shawn Tan and Khe Chai Sim. 2016 · 2016
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Later among the works it cites.
Sequence-to-sequence learning with latent neural grammars
Yoon Kim. 2021 · 2021
Later among the works it cites.
Lessons on parameter sharing across layers in transformers
Sho Takase and Shun Kiyono. 2021 · 2021
Later among the works it cites.
Disentangled sequence to sequence learning for compositional generalization
Hao Zheng and Mirella Lapata. 2021 · 2021
Later among the works it cites.
Mod-squad: Designing mixture of experts as modular multi-task learners
Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik Learned-Miller, and Chuang Gan. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. 2018 · 2018
Cited alongside, same era.
Scaling neural machine translation
Ott Myle, Edunov Sergey, Grangier David, Auli Michael, et al. 2018 · 2018
Cited alongside, same era.
Igloo: Slicing the features space to represent sequences
Vsevolod Sourkov. 2018 · 2018
Cited alongside, same era.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Ordered memory
Yikang Shen, Shawn Tan, Arian Hosseini, Zhouhan Lin, Alessandro Sordoni, and Aaron C Courville. 2019 · 2019
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A Ortega. 2022 · 2022
Later among the works it cites.
Compositional semantic parsing with large language models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2022 · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith. 2022 · 2022
Later among the works it cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q Tran, Yi Tay, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Mixture of attention heads: Selecting attention heads per token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2022 · 2022
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al. 2023 · 2023
Closest in time.
Moduleformer: Learning modular large language models from uncurated data
Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, and Chuang Gan. 2023 · 2023
Closest in time.