Fetching the paper…
Reading the bibliography…
Large pre-trained language models based on transformer architecture have drastically changed the natural language processing (NLP) landscape.
Language models are few-shot learners
Tom et al Brown. 2020 · 1901
Earlier work this paper cites.
Bridging the domain gap in cross-lingual document classification
Guokun Lai, Barlas Oguz, Yiming Yang, and Veselin Stoyanov. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Evaluation of spoken language systems: The atis domain
Patti Price. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality
Piotr Indyk and Rajeev Motwani. 1998 · 1998
Earlier work this paper cites.
Identifying and filtering near-duplicate documents
Andrei Z. Broder. 2000 · 2000
Earlier work this paper cites.
Summary cache: a scalable wide-area web cache sharing protocol
Li Fan, Pei Cao, Jussara Almeida, and Andrei Z Broder. 2000 · 2000
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S Charikar. 2002 · 2002
Earlier work this paper cites.
Detecting near-duplicates for web crawling
Gurmeet Singh Manku, Arvind Jain, and Anish Das Sarma. 2007 · 2007
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
In defense of minhash over simhash
Anshumali Shrivastava and Ping Li. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. 2015 · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Projectionnet: Learning efficient on-device deep networks using neural projections
Efficient on-device models using neural projections
Sujith Ravi. 2019 · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
End-to-end slot alignment and recognition for cross-lingual nlu
Weijia Xu, Batool Haider, and Saab Mansour. 2020 · 2020
Later among the works it cites.
Larger-scale transformers for multilingual masked language modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sujith Ravi. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018 · 2018
Cited alongside, same era.
Self-governing neural networks for on-device short text classification
Sujith Ravi and Zornitsa Kozareva. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
PRADO: Projection attention networks for document classification on-device
Prabhu Kaliamoorthi, Sujith Ravi, and Zornitsa Kozareva. 2019 · 2019
Cited alongside, same era.
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021 · 2021
Later among the works it cites.
Distilling large language models into tiny and effective students using pqrnn
Prabhu Kaliamoorthi, Aditya Siddhant, Edward Li, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
Mtop: A comprehensive multilingual task-oriented semantic parsing benchmark
Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021 · 2021
Later among the works it cites.
Pay attention to mlps
Hanxiao Liu, Zihang Dai, David So, and Quoc V Le. 2021 · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. 2021 · 2021
Later among the works it cites.
Extremely small BERT models from mixed-vocabulary training
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2021 · 2021
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Closest in time.
Efficient language modeling with sparse all-mlp
Ping Yu, Mikel Artetxe, Myle Ott, Sam Shleifer, Hongyu Gong, Ves Stoyanov, and Xian Li. 2022 · 2022
Closest in time.