Fetching the paper…
Reading the bibliography…
Domain adaptation of neural networks commonly relies on three training phases: pretraining, selected data training and then fine tuning.
Adaptation of deep bidirectional multilingual transformers for russian language
Yuri Kuratov and Mikhail Arkhipov. 2019 · 1905
Earlier work this paper cites.
Statistical Learning Theory
V.N. Vapnik. 1998 · 1998
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Hierarchical Bayesian domain adaptation
Jenny Rose Finkel and Christopher D. Manning. 2009 · 2009
Earlier work this paper cites.
Adapting naive bayes to domain adaptation for sentiment analysis
Songbo Tan, Xueqi Cheng, Yuefen Wang, and Hongbo Xu. 2009 · 2009
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011 · 2011
Earlier work this paper cites.
Domain adaptation for machine translation by mining unseen words
Hal Daumé III and Jagadeesh Jagarlamudi. 2011 · 2011
Earlier work this paper cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
Adaptation data selection using neural language models: Experiments in machine translation
Kevin Duh, Graham Neubig, Katsuhito Sudoh, and Hajime Tsukada. 2013 · 2013
Earlier work this paper cites.
Semi-supervised learning and domain adaptation in natural language processing
Anders Søgaard. 2013 · 2013
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Neural network methods for natural language processing
Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
Automated curriculum learning for neural networks
Alex Graves, Marc G. Bellemare, Jacob Menick, Rémi Munos, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Dynamic data selection for neural machine translation
Marlies van der Wees, Arianna Bisazza, and Christof Monz. 2017a · 2017
Cited alongside, same era.
Dynamic data selection for neural machine translation
Marlies van der Wees, Arianna Bisazza, and Christof Monz. 2017b · 2017
Cited alongside, same era.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori Hashimoto, and Percy Liang. 2019 · 2019
Later among the works it cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Later among the works it cites.
Dynamic data selection and weighting for iterative back-translation
Zi-Yi Dou, Antonios Anastasopoulos, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Denoising neural machine translation training with trusted data and online data selection
Wei Wang, Taro Watanabe, Macduff Hughes, Tetsuji Nakagawa, and Ciprian Chelba. 2018 · 2018
Cited alongside, same era.
Reinforced co-training
Jiawei Wu, Lei Li, and William Yang Wang. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee. 2020 · 2020
Later among the works it cites.
Findings of the WMT 2020 shared task on parallel corpus filtering and alignment
Philipp Koehn, Vishrav Chaudhary, Ahmed El-Kishky, Naman Goyal, Peng-Jen Chen, and Francisco Guzmán. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Which tasks should be learned together in multi-task learning?
Trevor Standley, Amir Roshan Zamir, Dawn Chen, Leonidas J. Guibas, Jitendra Malik, and Silvio Savarese. 2020 · 2020
Later among the works it cites.
Understanding and improving information transfer in multi-task learning
Sen Wu, Hongyang R. Zhang, and Christopher Ré. 2020 · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. 2020 · 2020
Later among the works it cites.
Auxiliary task update decomposition: The good, the bad and the neutral
Lucio Dery, Yann Dauphin, and David Grangier. 2021 · 2021
Closest in time.
Scalable evaluation and improvement of document set expansion via neural positive-unlabeled learning
Alon Jacovi, Gang Niu, Yoav Goldberg, and Masashi Sugiyama. 2021 · 2021
Closest in time.
Gradient-guided loss masking for neural machine translation
Xinyi Wang, Ankur Bapna, Melvin Johnson, and Orhan Firat. 2021 · 2021
Closest in time.
Reinforcement learning based curriculum optimization for neural machine translation
Gaurav Kumar, George Foster, Colin Cherry, and Maxim Krikun. 2019 · 2061
Closest in time.