Fetching the paper…
Reading the bibliography…
Inspired by the notion that ``{\it to copy is easier than to memorize}``, in this work, we introduce GNN-LM, which extends the vanilla neural language model (LM) by allowing to reference similar contexts in the entire training corpus.
A maximum likelihood approach to continuous speech recognition
Lalit R Bahl, Frederick Jelinek, and Robert L Mercer · 1983
Earlier work this paper cites.
Estimation of probabilities in the language model of the ibm speech recognition system
Arthur Nadas · 1984
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman · 1999
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon · 2001
Earlier work this paper cites.
Sac: Accelerating and structuring self-attention via sparse adaptive connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li · 2003
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2005
Earlier work this paper cites.
Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer · 2006
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Statistical language models based on neural networks
Tomáš Mikolov et al · 2012
Earlier work this paper cites.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A Smith, and Mike Lewis · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Optimized product quantization
Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun · 2013
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2016
Earlier work this paper cites.
Graph convolutional encoders for syntax-aware neural machine translation
Jasmijn Bastings, Ivan Titov, Wilker Aziz, Diego Marcheggiani, and Khalil Sima’an · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
Will Hamilton, Zhitao Ying, and Jure Leskovec · 2017
Earlier work this paper cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Data noising as smoothing in neural network language models
Ziang Xie, Sida I Wang, Jiwei Li, Daniel Lévy, Aiming Nie, Dan Jurafsky, and Andrew Y Ng · 2017
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Cited alongside, same era.
Question answering by reasoning across documents with graph convolutional networks
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2018
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Édouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Later among the works it cites.
Session-based recommendation with graph neural networks
Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Graph convolutional networks for text classification
Liang Yao, Chengsheng Mao, and Yuan Luo · 2019
Later among the works it cites.
Bp-transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Fast parametric learning with activation memorization
Jack Rae, Chris Dyer, Peter Dayan, and Timothy Lillicrap · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Cited alongside, same era.
Linfeng Song, Zhiguo Wang, Mo Yu, Yue Zhang, Radu Florian, and Daniel Gildea · 2018
Cited alongside, same era.
Learning to remember translation history with a continuous cache
Zhaopeng Tu, Yang Liu, Shuming Shi, and Tong Zhang · 2018
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli · 2018
Cited alongside, same era.
Guiding neural machine translation with retrieved translation pieces
Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura · 2018
Cited alongside, same era.
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Later among the works it cites.
Augmenting transformers with knn-based composite memory for dialogue
Angela Fan, Claire Gardent, Chloe Braud, and Antoine Bordes · 2020
Later among the works it cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Later among the works it cites.
Heterogeneous graph transformer
Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun · 2020
Later among the works it cites.
Boosting neural machine translation with similar translations
XU Jitao, Josep M Crego, and Jean Senellart · 2020
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Later among the works it cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Later among the works it cites.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Later among the works it cites.
Learning dense representations of phrases at scale
Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, and Danqi Chen · 2020
Later among the works it cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk · 2020
Later among the works it cites.
Segatron: Segment-aware transformer for language modeling and understanding
He Bai, Peng Shi, Jimmy Lin, Yuqing Xie, Luchen Tan, Kun Xiong, Wen Gao, and Ming Li · 2021
Closest in time.
Bertgcn: Transductive text classification by combining gcn and bert
Yuxiao Lin, Yuxian Meng, Xiaofei Sun, Qinghong Han, Kun Kuang, Jiwei Li, and Fei Wu · 2021
Closest in time.
Fast nearest neighbor machine translation
Yuxian Meng, Xiaoya Li, Xiayu Zheng, Fei Wu, Xiaofei Sun, Tianwei Zhang, and Jiwei Li · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Closest in time.
Chinesebert: Chinese pretraining enhanced by glyph and pinyin information
Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, and Jiwei Li · 2021
Closest in time.
Efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog
David Thulke, Nico Daheim, Christian Dugast, and Hermann Ney · 2021
Closest in time.