Fetching the paper…
Reading the bibliography…
Structural locality is a ubiquitous feature of real-world datasets, wherein data points are organized into local hierarchies.
Maximum entropy techniques for exploiting syntactic, semantic and collocational dependencies in language modeling
Sanjeev Khudanpur and Jun Wu · 2000
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
The sequence memoizer
Frank Wood, Jan Gasthaus, Cédric Archambeau, Lancelot James, and Yee Whye Teh · 2011
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling
Miltiadis Allamanis and Charles Sutton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Code completion with statistical language models
Veselin Raychev, Martin Vechev, and Eran Yahav · 2014
Earlier work this paper cites.
On the localness of software
Zhaopeng Tu, Zhendong Su, and Premkumar Devanbu · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
A neural network approach to context-sensitive generation of conversational responses
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan · 2015
Earlier work this paper cites.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Earlier work this paper cites.
On the naturalness of software
Abram Hindle, Earl T Barr, Mark Gabel, Zhendong Su, and Premkumar Devanbu · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Controlling politeness in neural machine translation via side constraints
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Larger-context language modelling with recurrent neural network
Tian Wang and Kyunghyun Cho · 2016
Cited alongside, same era.
Effective domain mixing for neural machine translation
Denny Britz, Quoc Le, and Reid Pryzant · 2017
Cited alongside, same era.
An empirical comparison of domain adaptation methods for neural machine translation
Chenhui Chu, Raj Dabre, and Sadao Kurohashi · 2017
Cited alongside, same era.
Unbounded cache model for online language modeling with open vocabulary
Edouard Grave, Moustapha Cissé, and Armand Joulin · 2017
Cited alongside, same era.
Are deep neural networks the best choice for modeling source code?
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vincent J Hellendoorn and Premkumar Devanbu · 2017
Cited alongside, same era.
Compressed nonparametric language modelling
Ehsan Shareghi, Gholamreza Haffari, and Trevor Cohn · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Generating sentences by editing prototypes
Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang · 2018
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Later among the works it cites.
Structural language models of code
Uri Alon, Roy Sadaka, Omer Levy, and Eran Yahav · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Latent relation language models
Hiroaki Hayashi, Zecong Hu, Chenyan Xiong, and Graham Neubig · 2020
Later among the works it cites.
Learning sparse prototypes for text generation
Junxian He, Taylor Berg-Kirkpatrick, and Graham Neubig · 2020
Later among the works it cites.
Big code!= big vocabulary: Open-vocabulary models for source code
Rafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, and Andrea Janes · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Closest in time.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Closest in time.
Pretrained language models for text generation: A survey
Junyi Li, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen · 2021
Closest in time.