Fetching the paper…
Reading the bibliography…
Deep learning models usually require a huge amount of data.
Learning represen-tations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
J. Schmidhuber · 1992
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
F. A. Gers, J. Schmidhuber, and F. Cummins · 1999
Earlier work this paper cites.
Recurrent neural networks for time series classification
Michael Hüsken and Peter Stagge · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Regularized multi–task learning
Theodoros Evgeniou and Massimiliano Pontil · 2004
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn et al · 2007
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2007
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Using domain similarity for performance estimation
Vincent Van Asch and Walter Daelemans · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua (14 June Bengio · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
A survey of transfer and multitask learning in bioinformatics
Qian Xu and Qiang Yang · 2011
Earlier work this paper cites.
Domain adaptation using domain similarity-and domain complexity-based instance selection for cross-domain sentiment analysis
Robert Remus · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
What is language policy?
David Cassels Johnson · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches. arxiv
K. Cho, B. Van Merrienboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Tiago Luís, Luís Marujo, Ramón Fernandez Astudillo, Silvio Amir, Chris Dyer, Alan W Black, and Isabel Trancoso · 2015
Cited alongside, same era.
Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/uthealth shared task track 1
Amber Stubbs, Christopher Kotfila, and Özlem Uzuner · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Adversarial deep averaging networks for cross-lingual sentiment classification
Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin et al · 2018
Later among the works it cites.
Fine-tuned language models for text classification
Jeremy Howard and Sebastian Ruder · 2018
Later among the works it cites.
Transfer learning in multilingual neural machine translation with dynamic vocabulary
Surafel M. Lakew et al · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters et al · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kelvin Xu et al · 2015
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Cited alongside, same era.
How transferable are neural networks in nlp applications?
Lili Mou et al · 2016
Cited alongside, same era.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Cited alongside, same era.
Machine comprehension using match-lstm and answer pointer
Shuohang Wang and Jing Jiang · 2016
Cited alongside, same era.
A survey of transfer learning
Karl Weiss, Taghi M. Khoshgoftaar, and DingDing Wang · 2016
Cited alongside, same era.
Dynamic coattention networks for question answering
Caiming Xiong, Victor Zhong, and Richard Socher · 2016
Cited alongside, same era.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Later among the works it cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2018
Later among the works it cites.
Cross-lingual transfer learning for multilingual task oriented dialog
Sebastian Schuster et al · 2018
Later among the works it cites.
Adversarial domain adaptation for duplicate question detection
Darsh J Shah, Tao Lei, Alessandro Moschitti, Salvatore Romeo, and Preslav Nakov · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai et al · 2019
Later among the works it cites.
Unified language model pre-training for natural language understanding and generation
Li Dong et al · 2019
Later among the works it cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby et al · 2019
Later among the works it cites.
An introductory survey on attention mechanisms in nlp problems
Dichao Hu · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Later among the works it cites.
Domain adaptation with bert-based domain classification and data selection
Xiaofei Ma, Peng Xu, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang · 2019
Later among the works it cites.
Evolution of transfer learning in natural language processing
Aditya Malte and Pratik Ratadiya · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford et al · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel et al · 2019
Later among the works it cites.
Bert-a: Finetuning bert with adapters and data augmentation, 2019
Sina Semnani and Kaushik Sadagopan · 2019
Later among the works it cites.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang et al · 2019
Later among the works it cites.
Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach
Wenpeng Yin, Jamaal Hay, and Dan Roth · 2019
Later among the works it cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Closest in time.