Fetching the paper…
Reading the bibliography…
Although Transformer has achieved great successes on many NLP tasks, its heavy structure with fully-connected attention connections leads to dependencies on large training data.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
A joint many-task model: Growing a neural network for multiple NLP tasks
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, and Richard Socher. 2017 · 1933
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa. 2011 · 2011
Earlier work this paper cites.
Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Bidirectional LSTM-CRF models for sequence tagging
Zhiheng Huang, Wei Xu, and Kai Yu. 2015 · 2015
Earlier work this paper cites.
Molding cnns for text: non-linear, non-consecutive convolutions
Tao Lei, Regina Barzilay, and Tommi S. Jaakkola. 2015 · 2015
Earlier work this paper cites.
When are tree structures necessary for deep learning of representations?
Jiwei Li, Thang Luong, Dan Jurafsky, and Eduard H. Hovy. 2015 · 2015
Cited alongside, same era.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, and Tiago Luís. 2015 · 2015
Cited alongside, same era.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Long short-term memory over recursive structures
Xiao-Dan Zhu, Parinaz Sobhani, and Hongyu Guo. 2015 · 2015
Cited alongside, same era.
Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. 2017 · 2017
Later among the works it cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Mo Feng, Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Later among the works it cites.
Adversarial multi-task learning for text classification
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017 · 2017
Later among the works it cites.
Shortcut-stacked sentence encoders for multi-domain inference
Yixin Nie and Mohit Bansal. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A fast unified model for parsing and sentence understanding
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, and Christopher Potts. 2016 · 2016
Cited alongside, same era.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Cited alongside, same era.
Learning natural language inference using bidirectional LSTM model and inner-attention
Yang Liu, Chengjie Sun, Lei Lin, and Xiaolong Wang. 2016 · 2016
Cited alongside, same era.
End-to-end sequence labeling via bi-directional lstm-cnns-crf
Xuezhe Ma and Eduard H. Hovy. 2016 · 2016
Cited alongside, same era.
Natural language inference by tree-based convolution and heuristic matching
Lili Mou, Rui Men, Ge Li, Yan Xu, Lu Zhang, Rui Yan, and Zhi Jin. 2016 · 2016
Cited alongside, same era.
J-NERD: joint named entity recognition and disambiguation with rich linguistic features
Dat Ba Nguyen, Martin Theobald, and Gerhard Weikum. 2016 · 2016
Cited alongside, same era.
Adnan Akhundov, Dietrich Trautmann, and Georg Groh. 2018 · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. 2018 · 2018
Later among the works it cites.
Learning to compose task-specific tree structures
Jihun Choi, Kang Min Yoo, and Sang-goo Lee. 2018 · 2018
Later among the works it cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Dynamic self-attention : Computing attention over words dynamically for sentence embedding
Deunsol Yoon, Dongbok Lee, and SangKeun Lee. 2018 · 2018
Later among the works it cites.
Sentence-state LSTM for text representation
Yue Zhang, Qi Liu, and Linfeng Song. 2018 · 2018
Later among the works it cites.