Fetching the paper…
Reading the bibliography…
In this work, we develop a novel regularizer to improve the learning of long-range dependency of sequence data.
What comes next? extractive summarization by next-sentence prediction
Jingyun Liu, Jackie CK Cheung, and Annie Louis. 2019a · 1901
Earlier work this paper cites.
Ernie: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 1905
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019 · 1907
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies , volume 1
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber. 2001 · 2001
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Introduction to Nonparametric Estimation , 1st edition
Alexandre B. Tsybakov. 2008 · 2008
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever. 2011 · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas. 2012 · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber. 2014 · 2014
Cited alongside, same era.
Learning longer memory in recurrent neural networks
Tomas Mikolov, Armand Joulin, Sumit Chopra, Michael Mathieu, and Marc’Aurelio Ranzato. 2014 · 2014
Cited alongside, same era.
Describing multimedia content using attention-based encoder-decoder networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio. 2015 · 2015
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Topic compositional neural language model
Wenlin Wang, Zhe Gan, Wenqi Wang, Dinghan Shen, Jiaji Huang, Wei Ping, Sanjeev Satheesh, and Lawrence Carin. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov. 2015 · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton. 2015 · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. 2015 · 2015
Cited alongside, same era.
Larger-context language modelling
Tian Wang and Kyunghyun Cho. 2015 · 2015
Cited alongside, same era.
Topicrnn: A recurrent neural network with long-range semantic dependency
Adji B Dieng, Chong Wang, Jianfeng Gao, and John Paisley. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al. 2016 · 2016
Cited alongside, same era.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2017 · 2017
Later among the works it cites.
Mine: mutual information neural estimation
Ishmael Belghazi, Sai Rajeswar, Aristide Baratin, R Devon Hjelm, and Aaron Courville. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Adam Trischler, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Independently recurrent neural network (indrnn): Building a longer and deeper rnn
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao. 2018 · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu H Trinh, Andrew M Dai, Thang Luong, and Quoc V Le. 2018 · 2018
Later among the works it cites.
Adversarially regularized autoencoders
Jake Zhao, Yoon Kim, Kelly Zhang, Alexander Rush, and Yann LeCun. 2018 · 2018
Later among the works it cites.
A cross-domain transferable neural coherence model
Peng Xu, Hamidreza Saghir, Jin Sung Kang, Teng Long, Avishek Joey Bose, Yanshuai Cao, and Jackie Chi Kit Cheung. 2019 · 2019
Closest in time.