Fetching the paper…
Reading the bibliography…
Models that perform well on a training domain often fail to generalize to out-of-domain (OOD) examples.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Domain robustness in neural machine translation
Mathias Müller, Annette Rios Gonzales, and Rico Sennrich. 2019 · 1911
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Transformation Invariance in Pattern Recognition — Tangent Distance and Tangent Propagation , pages 239–274. Springer Berlin Heidelberg, Berlin, Heidelberg
Patrice Y. Simard, Yann A. LeCun, John S. Denker, and Bernard Victorri. 1998 · 1998
Earlier work this paper cites.
Vicinal risk minimization
Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik. 2000 · 2000
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Data Augmentation using Pre-trained Transformer Models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020 · 2003
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
G-DAUG: Generative Data Augmentation for Commonsense Reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020 · 2004
Earlier work this paper cites.
Semi-Supervised Learning (Adaptive Computation and Machine Learning)
Olivier Chapelle, Bernhard Schölkopf, and Alexander Zien. 2006 · 2006
Earlier work this paper cites.
Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
Frustratingly easy domain adaptation
Hal Daumé III. 2007 · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008 · 2008
Earlier work this paper cites.
Dataset Shift in Machine Learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence. 2009 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros. 2011 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
The trilingual ALLEGRA corpus: Presentation and possible use for lexicon induction
Yves Scherrer and Bruno Cartoni. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann. 2012 · 2012
Cited alongside, same era.
Generalized denoising auto-encoders as generative models
Yoshua Bengio, Li Yao, Guillaume Alain, and Pascal Vincent. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Learning with pseudo-ensembles
Philip Bachman, Ouais Alsharif, and Doina Precup. 2014 · 2014
Cited alongside, same era.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Cited alongside, same era.
Data noising as smoothing in neural network language models
Ziang Xie, Sida I. Wang, Jiwei Li, Daniel Levy, Aiming Nie, Dan Jurafksy, and Andrew Y. Ng. 2017 · 2017
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Classical structured prediction losses for sequence to sequence learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
Geometric robustness of deep networks: Analysis and improvement
Can Kanbak, Moosavi-Dezfooli Seyed-Mohsen, and Pascal Frossard. 2018 · 2018
Later among the works it cites.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets
William Yang Wang and Diyi Yang. 2015 · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans. 2016 · 2016
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
SwitchOut: an efficient data augmentation algorithm for neural machine translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Fast and accurate reading comprehension by combining self-attention and convolution
Adams Wei Yu, David Dohan, Quoc Le, Thang Luong, Rui Zhao, and Kai Chen. 2018 · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Later among the works it cites.
Justifying recommendations using distantly-labeled reviews and fined-grained aspects
Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019 · 2019
Later among the works it cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Data augmentation with manifold exploring geometric transformations for increased performance and robustness
Magdalini Paschali, Walter Simson, Abhijit Guha Roy, Muhammad Ferjad Naeem, Rüdiger Göbl, Christian Wachinger, and Nassir Navab. 2019 · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 2019
Later among the works it cites.
Do not have enough data? deep learning to the rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. 2020 · 2020
Closest in time.
Open sourcing german bert
Branden Chan, Timo Möller, Malte Pietsch, Tanay Soni, and Chin Man Yeung. 2020 · 2020
Closest in time.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020 · 2020
Closest in time.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le. 2020 · 2020
Closest in time.