Fetching the paper…
Reading the bibliography…
Transformers achieve unrivalled performance in modelling language, but remain inefficient in terms of memory and time complexity.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Human behavior and the principle of least effort: An introduction to human ecology
George Kingsley Zipf. 1949 · 1949
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah Hihi and Yoshua Bengio. 1995 · 1995
Earlier work this paper cites.
Finding structure via compression
Jason L. Hutchens and Michael D. Alder. 1998 · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Japanese and Korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
A clockwork RNN
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel. 2015 · 2015
Earlier work this paper cites.
Zipf’s law of abbreviation as a language universal
Chris Bentz and Ramon Ferrer-i Cancho. 2016 · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Kamil Rocki, Tomasz Kornuta, and Tegan Maharaj. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
David Krueger, Tegan Maharaj, Janos Kramar, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Christopher Pal. 2017 · 2017
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Cited alongside, same era.
Fast-slow recurrent neural networks
Asier Mujika, Florian Meier, and Angelika Steger. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Skip RNN: Learning to skip state updates in recurrent neural networks
Víctor Campos, Brendan Jou, Xavier Giró i Nieto, Jordi Torres, and Shih-Fu Chang. 2018 · 2018
Cited alongside, same era.
Segmental contrastive predictive coding for unsupervised word segmentation
Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velazquez, and Najim Dehak. 2021 · 2021
Later among the works it cites.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. 2021 · 2021
Later among the works it cites.
Revisiting the Uniform Information Density hypothesis
Clara Meister, Tiago Pimentel, Patrick Haller, Lena Jäger, Ryan Cotterell, and Roger Levy. 2021 · 2021
Later among the works it cites.
Combiner: Full attention Transformer with sparse computation cost
Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai. 2021 · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taku Kudo. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Neural speed reading via skim-RNN
Minjoon Seo, Sewon Min, Ali Farhadi, and Hannaneh Hajishirzi. 2018 · 2018
Cited alongside, same era.
Preserving activations in recurrent neural networks based on surprisal
Tayfun Alpay, Fares Abawi, and Stefan Wermter. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019 · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Closest in time.
Variable-rate hierarchical CPC leads to acoustic unit discovery in speech
Santiago Cuervo, Adrian Łańcucki, Ricard Marxer, Paweł Rychlikowski, and Jan Chorowski. 2022 · 2022
Closest in time.
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, et al. 2022 · 2022
Closest in time.
FNet: Mixing tokens with Fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. 2022 · 2022
Closest in time.
Large text compression benchmark
Matt Mahoney. 2006 · 2022
Closest in time.
Hierarchical transformers are more efficient language models
Piotr Nawrot, Szymon Tworkowski, Michał Tyrolski, Lukasz Kaiser, Yuhuai Wu, Christian Szegedy, and Henryk Michalewski. 2022 · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Closest in time.
Scrolls: Standardized comparison over long language sequences
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, et al. 2022 · 2022
Closest in time.
The bitter lesson
Richard Sutton. 2019 · 2022
Closest in time.
Charformer: Fast character transformers via gradient-based subword tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2022 · 2022
Closest in time.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022 · 2022
Closest in time.
Memorizing transformers
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy. 2022 · 2022
Closest in time.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Closest in time.